Commit graph

3 commits

Author SHA1 Message Date
Gergő Magyar
010a7d806a
fix(schema): declare the full scope-resolution relation cross product (#2792) (#2793)
* fix(schema): declare the full scope-resolution relation cross product (#2792)

`RELATION_SCHEMA` was hand-listed, and every prior fix added only the
FROM/TO pair named in a crash report — `Const→Method` in #2769, the
Swift/Rust member pairs before it. So `analyze` kept aborting at
`assertDeclaredPair` on the next codebase whose edges happened to land on
a different pair; #2792 reports `Class→Variable` on Java.

Audit the surface instead of the symptom. `buildGraphNodeLookup` skips
any node whose label is not in `isLinkableLabel`, so the lookup holds
only linkable-labelled nodes — and both endpoints of every graph-bridge
edge resolve through that lookup. The emittable surface is therefore
exactly:

  FROM  LINKABLE_LABELS + File   (the module-level caller fallback)
  TO    LINKABLE_LABELS + CALL_TARGET_TYPES

`isCallerAnchorLabel` is a strict subset of linkable and contributes
nothing on top. `CALL_TARGET_TYPES` contributes `Delegate`, which
`tryEmitEdgeWithExplicitTargetId` can emit without going through the
lookup at all.

Generate that 14x14 block into the DDL rather than listing it: 223 -> 322
declared pairs, and no future pair from these sets can be missing by
construction. The containment/inheritance/DI/route/cluster/PDG pairs stay
hand-declared — no single predicate describes them.

Both label sets live in the ingestion layer, which `core/lbug` must not
import, so schema.ts carries twin lists. test/unit/schema-pair-coverage.ts
derives the requirement from the originals and fails CI when either set
grows without the pairs landing here — the piecemeal loop this fix ends.

Measured before widening: at 322 pairs the cost is inside noise
(1.09s vs 1.12s per 300 anchored queries on a 32-table DB), but the full
32x32 cross product is ~1.8x on untyped-endpoint anchored queries. The
audited subset is the right scope, not "declare everything".

INCREMENTAL_SCHEMA_VERSION 34 -> 35: LadybugDB fixes endpoint pairs when
the rel table is created, so a pre-v35 database physically cannot store
these edges.

Closes #2792

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(schema): declare the non-bridge structural pairs COBOL and Vue emit

The generated scope-resolution block closed the half of RELATION_SCHEMA a
label predicate can describe. The hand-declared half was still stale: with
#2791's Function->Variable fix applied, `analyze` continued to abort on this
repo's own test/fixtures/lang-resolution with

  Relationship label pair Module→Property is not declared

A full sweep (assertDeclaredPair patched to log-and-skip, run over the whole
fixture corpus) found 13 undeclared pairs over 106 edges. This branch already
covered 3 of them via the cross product; the remaining 10 come from emitters
outside the graph bridge:

  - cobol-processor.ts mints Module / Namespace / Record / Property /
    CodeElement and wires them with CONTAINS, CALLS and ACCESSES (9 pairs)
  - vue-sfc-extractor.ts emits BINDS_EVENT_HANDLER from a handler Function to
    the child component's File, the only edge whose target is a File (1 pair)

CodeElement, Namespace, Record and File are in neither scope-bridge label set,
so neither the generated block nor schema-pair-coverage.test.ts can reach them.

Adds test/integration/structural-pair-coverage.test.ts, which derives the
requirement from a corpus instead of a predicate: it runs the real pipeline
over the non-bridge fixtures and requires every FROM/TO pair they produce to
be declared. Mutation-checked — dropping `FROM Function TO File` fails it with
exactly Function|File.

Verified: cobol-app, vue-basic and php-transitive-traits now index instead of
aborting; the full lang-resolution corpus completes at 10,876 nodes / 18,517
edges; scrypster/muninndb at 0b7a4272 (the #2789 repro) completes at 20,069
nodes / 71,580 edges, matching #2791 exactly, so this supersedes that PR.

* refactor(test): simplify the structural pair coverage guard

Cleanup pass over the previous commit. No behaviour change to the schema.

- reuse `FIXTURES` and `runPipelineFromRepo` from resolvers/helpers.ts instead
  of re-deriving the fixture root and importing pipeline.js directly
- gate on `distWorkerExists()` like every other integration test that passes
  `workerUrlForTest`, so a missing dist skips rather than fails
- run the three fixtures with `it.concurrent.each`; they share nothing and the
  cost is almost all worker spawn plus grammar load, which overlaps well
  (tests phase 21-24s -> 5.6s measured)
- replace the sentinel-in-a-Set filter with a plain `.filter()` chain, matching
  the sibling unit test, and move the declared/table lookups off the per-edge
  path onto the deduped set
- move the pure string pin out of the integration tier into
  schema-pair-coverage.test.ts, where the identical construct already lives, so
  it needs no build and survives fixture deletion
- trim the schema and test prose that restated the code, and correct the
  BINDS_EVENT_HANDLER attribution: it is emitted by
  languages/vue/scope-resolver.ts, not vue-sfc-extractor.ts
- amend the v35 comment to mention the 10 structural pairs it now also stamps

Still mutation-checked: dropping `FROM Function TO File` now fails both the
integration sweep and the unit pin with exactly Function|File. 89 tests green.

* fix(schema): generate the attachment pair surface and close four analyze aborts

Review of the generated scope-bridge cross product found four `analyze`
hard-aborts still live at head, each reproduced end-to-end on the default
user path (`analyze --index-only --skip-git`):

  Method→Annotation   Spring `@Bean` + `@ConditionalOnMissingBean` (Java + Kotlin)
  Method→File         Vue Options-API `methods:` handler bound to a child event
  Namespace→Record    COBOL `DECLARATIVES` / `USE AFTER STANDARD ERROR ON <file>`
  Class→Tool          `@mcp.tool()` applied to a class

All four are pre-existing on main, and both existing guards were structurally
blind to them: the unit guard derives from LINKABLE_LABELS ∪ CALL_TARGET_TYPES
(none of Annotation/Tool/Record/File-as-target is a member) and the corpus
guard ran three fixtures that exercise none of these emitters. All 16 tests
passed while all four crashes were live.

The PR's model — "bridge endpoint × structural endpoint" — does not fit:
Namespace→Record is structural on both sides. The property that does hold is
that the ANCHOR is a lookup result, not a literal at the emit site, so the
emitter cannot constrain its label. That gives a second closed-form rule:

  DEFINITION_ANCHOR_LABELS × ATTACHMENT_TARGET_LABELS

DEFINITION_ANCHOR_LABELS is derived from NODE_TABLES by subtraction, so a new
node table joins automatically. 332 → 450 declared pairs.

Sized against a committed harness (gitnexus/bench/schema-pairs), real
@ladybugdb/core, identical data: 450 costs 0.93–1.05× of 332 on untyped-endpoint
anchored queries — inside noise — versus 1.22–1.43× at 641 and 2.03–2.34× at
1024. The harness reproduces the known #2792 cliff, which is what makes the 450
figure trustworthy.

Also in this change:

- Delete the 161 hand-declared pairs the rules already generate (233 → 72).
  The declared set is byte-identical at 450; those lines were load-bearing
  shadow, because the generator suppresses anything already declared
  structurally, so narrowing a rule later would silently keep pairs alive.
  A new guard fails CI if a hand-declared pair is ever re-added inside a rule.
- Import LINKABLE_LABELS / CALL_TARGET_TYPES instead of hand-copying them.
  The twins' stated justification ("the ingestion layer must not be imported
  here") is false: csv-generator.ts and lbug-adapter.ts, siblings in the same
  directory, already do, and no rule in AGENTS.md / ARCHITECTURE.md /
  CONTRIBUTING.md / GUARDRAILS.md states otherwise.
- Resolve `resolveStreamGraphEmit` after the guards that rebind `options.force`,
  not at function entry. It gates on `force`, and every freshness guard runs
  ~360 lines later, so the v34→v35 bump would have pushed every existing index
  down the non-streamed emit path — losing the #2680 memory streaming added for
  the #2649 kernel-scale OOM, for exactly the population most likely to be
  memory-constrained.
- `UndeclaredRelationPairError` now carries the relationship type, both node ids
  and the source file, with a matching CLI branch. The old message named only
  the abstract label pair, which a user could not act on. Found through the
  cause chain, since pipeline-phases/runner.ts rewraps every phase failure.
- Share one classifier (`relPairKeyFor`) across the router, both emit sinks and
  the corpus guard, which previously hand-mirrored the router's skip rule; one
  cause-chain walker in lib/utils.ts; one exported pair-matching regex.
- Corpus guard: four new fixtures reproducing the aborts, per-fixture sentinel
  pairs so a fixture that stops emitting fails loudly instead of passing
  vacuously on an empty graph.

The per-edge path stays allocation-free: the failure context is passed
positionally and the message is built only inside the throw.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0182jkjQqzACkJKYw4MLDnhX

* test(bench): re-baseline the COBOL capture fingerprint for the new fixture

`bench/scope-capture` globs `lang-resolution/cobol-*`, so the
`cobol-declaratives` fixture added in 81daf370e (to reproduce the
`Namespace→Record` analyze abort) joined that corpus and shifted the
fingerprint — 14 → 15 files.

Verified corpus-only, not a capture change: with that one fixture moved
aside the fingerprint is byte-identical to the prior baseline
(d45bb091…), and 81daf370e touches no COBOL capture code. The new value
reproduces CI's reported hash exactly. Scaling 0.677 < 1.5 budget.

`bench/scope-capture/measure.mjs --check` → PASS (15 languages).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0182jkjQqzACkJKYw4MLDnhX

---------

Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 16:26:25 +01:00
MyShining
de84ad6297
feat(spring): index @Bean factories and @Resource injection (#2740)
* feat(spring): index Bean factories and Resource injection

* fix(spring): address Bean and Resource review findings

* refactor(lbug): keep relation pair parsing in router

* test(lbug): preserve schema exports in WAL mocks

* test(cache): align schema bump pin

---------

Co-authored-by: Shining <xuenning@qiyi.com>
Co-authored-by: Gergo Magyar <gergomagyar0@gmail.com>
2026-07-31 10:33:42 +01:00
Gergő Magyar
df08ecc397
perf(lbug): cut graph-DB emit/persistence wall time (#2203) (#2215)
* perf(lbug): add PROF_LBUG_LOAD persistence-path timing breakdown (#2203 U1)

loadGraphToLbug is un-timed today; the analyze 'emit' number is the
scope-resolution emit bucket, not the CSV->COPY persistence path. Add a
zero-cost-when-off per-stage breakdown (csv-emit/copy-nodes/rel-split/
copy-rels/fallback/total + node/rel counts) gated by PROF_LBUG_LOAD=1,
mirroring the PROF_SCOPE_RESOLUTION pattern. Document the flag in README.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* perf(lbug): route relationships to per-pair CSVs in the emit pass (#2203 U2)

Relationships were written once to a monolithic relations.csv, then re-read
line-by-line (regex per edge) and re-split into per-FROM->TO-label-pair files
before COPY — writing and reading the entire ~1M-edge set twice. Route each
edge to its pair file directly during the single emit pass via a shared
RelPairRouter, eliminating the monolithic write + re-read + per-edge regex.

The router applies the SAME getNodeLabel + validTables filter as the legacy
splitRelCsvByLabelPair, which is retained as a differential oracle. A new
differential test asserts the direct-emit per-pair files are byte-for-byte
identical to the oracle's, with identical skip/total accounting. The prof
line (U1) drops its rel-split stage (routing now folds into csv-emit).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* perf(lbug): skip per-row microtask tick in BufferedCSVWriter (#2203 U3)

addRow awaited an already-resolved promise on every buffered row, scheduling
a microtask per node even when nothing flushed (millions at scale). It now
returns a promise ONLY when it flushes; the node-emit loop awaits once per
iteration after the switch. Flush/drain semantics are unchanged, so
backpressure on the rows that actually write is preserved and the emitted
CSV bytes are byte-identical (covered by the determinism + differential tests).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* bench(lbug): emit throughput + byte-identity gate for the persistence path (#2203 U4)

Build-free bench (bench/emit-persistence/measure.mjs) times streamAllCSVsToDisk
on a synthetic graph at two scales and gates: (1) an order-independent sha256
fingerprint over every emitted CSV line — the byte-identity guard for the U2/U3
emit optimisations — and (2) a scaling-ratio budget catching an O(n^2) emit
re-regression. Wired into ci-tests.yml alongside the cfg/scope-capture benches.
The LadybugDB COPY half needs a real DB, so its timing stays in PROF_LBUG_LOAD
+ the integration round-trip tests (documented in the bench README, with the
deferred COPY-parallelism follow-up).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(review): apply autofix feedback (#2203)

- P1: router backpressure drain-await rejected with a generic AbortError,
  masking the real EMFILE/disk-full error. Expose RelPairRouter.lastError and
  rethrow it in the emit catch — mirrors the oracle's throw streamError ?? err.
- P1: cover RelPairRouter error + backpressure + teardown paths with a new
  unit test (test/unit/rel-pair-routing.test.ts) using an injected mock stream.
- P2: wrap streamAllCSVsToDisk body in try/finally so the setMaxListeners bump
  is always restored (the U2 rel-routing throw path could leak it).
- P2: dedup WriteStreamFactory — re-export the canonical type from
  rel-pair-routing instead of a second identical declaration.
- P2: annotate splitRelCsvByLabelPair @internal as the retained differential
  oracle so a future dead-code sweep doesn't delete the byte-identity guard.
- P3: differential test now covers the proc_ prefix + clears
  GITNEXUS_SORT_GRAPH_OUTPUT to prevent env-leak desync.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(lbug): scope byte-identity to quote-free ids + lock the quote-in-id divergence (#2215 review)

The 'byte-identical' claim was unconditional, but the router derives labels from
the raw id while the retained splitRelCsvByLabelPair oracle re-derives them via a
regex over the escaped row — so for an id containing a double-quote they diverge
(the router is the more-correct path). Soften the wording in rel-pair-routing.ts,
the bench README, and the differential-test comment to document the exception,
and add a differential test asserting the intended divergence (router routes the
quote-in-id edge; oracle drops it) so a future change can't silently revert to
the buggy regex semantics.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* bench(lbug): per-file fingerprint so the gate catches pair-file mis-routing (#2215 review)

fingerprintEmit flattened every line of every per-pair file into one array,
sorted globally, and hashed — losing file boundaries, so a row routed to the
WRONG pair file produced an identical fingerprint. Hash a per-file digest
(filename + sha256(file bytes)) and combine the sorted entry list, so mis-routing
(and within-file row reordering) now changes the fingerprint. Baseline
regenerated; the new scheme yields a different hash on byte-identical emit,
confirming it is sensitive to file structure the old flatten ignored.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* bench(lbug): add absolute large-scale wall-time backstop to the emit gate (#2215 review)

The scaling-ratio gate only compares large/small, so a uniform Nx slowdown at
both scales passes with ratio ~1.0. Add an opt-in max_ms_large ceiling (1000ms
vs observed ~200ms — generous, host-noise-tolerant) that --check enforces
alongside the ratio, catching a gross absolute regression the ratio misses.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(lbug): cover the sorted-output path in the byte-identity differential (#2215 review)

The differential test only exercised the default insertion-order emit path. Add
a case under GITNEXUS_SORT_GRAPH_OUTPUT=1 that feeds the oracle the same
id-sorted order orderedRelationships() uses and asserts per-pair byte-identity,
so within-pair row reordering on the sorted path can't slip past the gate.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(lbug): cover the invalid-TO-label skip branch (#2215 review)

Only an invalid-FROM label was exercised; the validTables skip is an OR over
both endpoints, so the invalid-TO branch was untested (an inverted && would
have slipped through). Add a valid-FROM/invalid-TO edge to the differential
test and the router unit test, asserting it's skipped identically by router and
oracle.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(lbug): exercise the BufferedCSVWriter FLUSH_EVERY boundary in vitest (#2215 review)

The U3 addRow change (returns a flush promise only on flush; undefined when
buffered) and the loop's `if (pending) await pending` were only crossed by the
bench, never vitest (all fixtures are <500 nodes). Add a 600-node graph through
streamAllCSVsToDisk asserting all rows land exactly once across the 500-row
flush boundary — no drops, dups, or corruption.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(lbug): drop redundant step cast in buildRelRow (#2215 review)

GraphRelationship.step is already typed number?, so (rel as { step?: number }).step
was a no-op structural cast that obscured the shared-type coupling. Use rel.step
directly. Byte-identical — bench fingerprint unchanged, differential test green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(lbug): make the unknown-label node drop explicit (#2215 review)

With the U3 `let pending` switch idiom, a node whose label matches neither
codeWriterMap nor multiLangWriters left `pending` undefined and was silently
dropped — a footgun for a future node type. Add an explicit else with a comment
documenting that unknown labels are intentionally not persisted and that a new
type must be wired into a writer map. No behavior change (byte-identity + tests
unchanged).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(lbug): drop the unused WriteStreamFactory re-export (#2215 review)

The type was re-exported from lbug-adapter 'to preserve this module's surface,'
but no external code imports it by name from here (the only test reference is a
comment). Keep the import from rel-pair-routing.ts (its canonical home, still
used by splitRelCsvByLabelPair's signature) and drop the dead re-export.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 19:40:59 +01:00