GitNexus/gitnexus/test/integration/cobol-pipeline-benchmark.test.ts
Gergő Magyar 083aedbc41
refactor(ingestion): delete legacy call-resolution DAG + heritage processor (RING4-1, #942) (#2023)
* refactor(ingestion): delete legacy call-resolution DAG + heritage processor (#942)

RING4-1: all 16 production languages (incl. Vue #940) are registry-primary, so
the legacy resolution legs only ran under the now-removed CI parity gate. Calls
and inheritance now resolve exclusively through scope-resolution
(Registry.lookup, preEmitInheritanceEdges, emitHeritageEdges, buildMro →
MethodDispatchIndex).

Removed:
- Call-resolution DAG: call-processor.ts legacy body (processCalls,
  processCallsFromExtracted, resolveCallTarget + all resolver/dispatch/chain
  helpers), model/resolve.ts MRO-via-HeritageMap, model/heritage-map.ts,
  type-env DAG types; inferImplicitReceiver/selectDispatch LanguageProvider
  hooks + Ruby impls; DispatchDecision/ImplicitReceiverOverride/ReceiverEnriched.
- Legacy heritage path: heritage-processor.ts, heritage-types.ts,
  heritage-extractors/, @heritage.* tree-sitter queries, heritageExtractor/
  heritageDefaultEdge/interfaceNamePattern wiring, worker + parse-impl heritage
  passes (parse-worker/parsing-processor lockstep), cross-file-impl DAG pass.
- Scope-parity infrastructure entirely (no legacy↔registry parity left to run):
  scripts/run-parity.ts, scripts/ci-list-migrated-languages.ts,
  ci-scope-parity.yml, test:parity, and the scope-parity ci.yml gate. Resolver
  integration tests still run via the normal tests job.

Kept (shared infra, NOT call-DAG-only): type-env.ts buildTypeEnv (field
extraction / structure phase / embeddings), model/resolve.ts c3Linearize +
gatherAncestors (mro-processor mroPhase), route/fetch/exported-type-map helpers
in call-processor.ts, preEmitInheritanceEdges (legacy-edge dedup simplified).

Acceptance: grep for resolveCallTarget/inferImplicitReceiver/selectDispatch/
buildHeritageMap/HeritageMap/processHeritage/heritageExtractor/@heritage. is zero
across src + test. tsc clean (both packages); resolver integration suite green
(bit-compatible EXTENDS/IMPLEMENTS/CALLS); scope-capture fingerprints unchanged
(python re-baselined: removed redundant ignored captures). ARCHITECTURE.md
updated to scope-resolution-only.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(review): apply autofix feedback (#942)

ce-code-review autofix pass on the RING4-1 deletion:
- parse-cache.ts: bump SCHEMA_BUMP 2→3 — ParseWorkerResult lost its `heritage`
  field, so stale on-disk caches must invalidate (prevents a rollback replaying
  a heritage-less cache into legacy code) [api-contract P2].
- parse-impl.ts: drop 3 now-unused type imports (ExtractedCall,
  ExtractedAssignment, FileConstructorBindings) left by the deferred-block
  removal — would fail the eslint CI gate [correctness+maintainability P1].
- AGENTS.md / CLAUDE.md / scope-resolver.ts contract doc: fix stale pointers to
  the deleted "§ Call-Resolution DAG" section + removed hooks; preserve the
  language-neutrality rule [project-standards P1].
- registry-primary-flag.ts / cross-file.ts / parse-impl.ts: refresh stale
  comments referencing deleted symbols (legacy DAG, runCrossFileBindingPropagation).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(ingestion): remove the vestigial isRegistryPrimary flag (#942)

With the legacy call-resolution DAG deleted, the per-language
`REGISTRY_PRIMARY_<LANG>` / `isRegistryPrimary` / `MIGRATED_LANGUAGES` flag had
only one meaningful state — every production language resolves via
scope-resolution — and an explicit `=0` override could only *disable*
resolution with no fallback (a footgun the review flagged). Removing it.

- Delete `registry-primary-flag.ts` and the now-dead `shadow-harness.ts`
  (legacy↔registry shadow-parity tool) + its test.
- Collapse the three flag gates to their behavior-preserving outcome
  (`SCOPE_RESOLVERS == MIGRATED_LANGUAGES`, so this is a no-op):
  - scope-resolution phase now runs for every registered `SCOPE_RESOLVERS`
    entry (was `∩ MIGRATED_LANGUAGES`).
  - import-processor `addImportGraphEdge` + parse-impl `shouldAccumulate`:
    the legacy emit/accumulate paths were already inert for migrated
    languages (scope-resolution owns IMPORTS via the imports-to-edges bridge);
    drop the flag term.
- Collapse flag-branching tests to the scope-resolution path and delete the
  csharp legacy-`=0`-leg describe blocks; remove the ruby/rust-scope env-forcing
  hooks (no-ops now).
- Refresh docs/comments (ARCHITECTURE.md "one registration", scope-resolver
  cookbook, phase deps) — adding a language is now a single `SCOPE_RESOLVERS`
  registration.

Verified: tsc clean (both packages); resolver integration tests green
(747 assertions across cobol/csharp/ruby/rust/typescript/go, IMPORTS edges
intact); grep for the flag symbols is zero across src + test.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* style(format): prettier formatting on #942 changes

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(ci): drop legacy heritage-capture tests + re-baseline scope-capture fingerprints (#942)

Two CI failures from the #942 cleanup, surfaced by the tri-review + CI:

- tree-sitter-languages.test.ts: two tests asserted `@heritage.*` captures
  (Rust trait-impl, Dart extends/implements/with) that this PR removed. The
  acceptance grep used `@heritage\.` (with `@`); these reference the runtime
  capture name `heritage.trait` (no `@`), so they slipped the earlier sweep.
  Inheritance is now covered by the resolver integration suite. (fixed macos-latest)

- Re-baselined the scope-capture bench fingerprints for csharp/rust/ruby/java/
  javascript/kotlin (baselines.json) + python (python-scope/baseline-fingerprint.txt).
  The earlier test-cleanup reworded comments inside the lang-resolution fixture
  files (Shapes.cs, child.rs, derived.rb, IA.java/Plain.java, Service.js, F.kt,
  app.py) to scrub deleted-symbol references for the acceptance grep; those are
  the bench corpus, so capture node positions shifted. Capture LOGIC is
  unchanged — verified `--check` passes for all 14 langs + python. (fixed benchmarks)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs/chore: scrub remaining REGISTRY_PRIMARY + deleted-symbol references (#942)

Tri-review P3 follow-ups (verified):
- TESTING.md: rewrite the "Scope-resolution parity" section — the legacy
  dual-leg (REGISTRY_PRIMARY_<LANG>=0/1) and `npm run test:parity` no longer
  exist; resolver tests run once on the sole scope-resolution path in the
  normal tests job.
- scripts/bench-scope-resolution.ts: drop the inert `REGISTRY_PRIMARY_PYTHON=1`
  env set + usage hint (the flag is gone).
- ruby/scope-resolver.ts, php/captures.ts: re-point doc-comments off the
  deleted heritage-map.ts / heritage-processor.ts to the current behavior.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(ci): prettier format + regenerate scope-capture goldens (#942)

Two more CI failures, same root cause as the bench re-baseline (the
test-cleanup reworded comments in lang-resolution bench/golden-corpus fixtures):

- quality/format: prettier on tree-sitter-languages.test.ts (blank line left by
  the deleted heritage-capture tests) + TESTING.md (the rewritten section).
- tests/ubuntu/coverage: `csharp-captures-golden` (and python/ruby/rust) drifted
  because the edited fixtures feed the per-language capture-golden snapshots too
  (not just the bench). Regenerated via UPDATE_GOLDEN=1. Verified safe: only the
  edited-fixture entries changed; csharp `captureGroups` unchanged (38) — digest
  shifted from comment-position only; capture LOGIC untouched. 1168 scope-
  resolution tests pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(resolvers): drop createResolverParityIt wrapper, use vitest it directly

The parity-aware `it` wrapper became a no-op when #942 removed the legacy
call-resolution DAG (it just returned vitest's `it`). Remove it entirely so
the resolver tests call vitest's `it` directly instead of shadowing it with a
local `const it` (or `pit`/`rustParityIt`):

- helpers.ts: delete createResolverParityIt + its now-unused vitestIt import
  and VitestIt type.
- 16 files: drop `const it = createResolverParityIt('x')` and import `it`
  from vitest instead.
- ruby.test.ts (pit) + rust.test.ts (rustParityIt): rename calls to `it`.
- Scrub every comment that described the removed wrapper / dual-mode parity
  skip / legacy_skip gate (vue-scope, js/ts/dart/php/python headers, rust x2,
  cpp, swift x4, rust-coverage). Genuine test rationale is kept; only the
  vestigial two-leg framing is dropped. Accurate "legacy DAG (removed in
  #942)" historical notes are retained.

No fixtures touched (no bench/golden re-baseline). tsc clean; rust+ruby
resolver suites green (323 tests, incl. #1992 worker-path parity after a
local dist build).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 11:07:37 +01:00

251 lines
10 KiB
TypeScript

/**
* COBOL ingestion pipeline benchmark.
*
* Generates synthetic COBOL codebases at increasing scales and measures
* wall-clock time and peak heap through the full pipeline — scanning,
* preprocessing, COPY expansion, CALL resolution, and scope extraction.
*
* Run: GITNEXUS_BENCH=1 npx vitest run test/integration/cobol-pipeline-benchmark.test.ts
*
* COBOL is wired as a standalone provider, so the scope-resolution phase is
* skipped for it (standalone guard in phase.ts) and node/edge counts come
* entirely from cobolPhase.
*
* IMPORTANT — this benchmark measures scaling in FILE COUNT, so per-file work
* must stay constant as fileCount grows. Each program therefore COPYs a fixed
* number of shared copybooks (COPYBOOKS_PER_PROGRAM), independent of fileCount.
* Do NOT make every program COPY all copybooks: copybookCount grows as
* floor(fileCount/5), so copy-all makes emitted data-item nodes — and thus
* total work — O(fileCount²), which measures copybook fan-out rather than
* file-count scaling. The pipeline itself is O(fileCount) (verified: with
* constant fan-out, node count and wall-clock scale exactly linearly); the
* node-ratio assertion below guards against reintroducing the O(n²) pattern.
*/
import { describe, it, expect } from 'vitest';
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import { runPipelineFromRepo } from '../../src/core/ingestion/pipeline.js';
const BENCH_ENABLED = process.env.GITNEXUS_BENCH === '1';
interface BenchResult {
fileCount: number;
programCount: number;
paragraphCount: number;
copybookCount: number;
elapsedMs: number;
peakHeapMB: number;
nodeCount: number;
edgeCount: number;
}
function generateCobolFixture(
fileCount: number,
paragraphsPerProgram: number,
): { dir: string; programCount: number; paragraphCount: number; copybookCount: number } {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), `cobol-bench-${fileCount}-`));
const copybookDir = path.join(dir, 'copybooks');
fs.mkdirSync(copybookDir, { recursive: true });
const programCount = fileCount;
const paragraphCount = fileCount * paragraphsPerProgram;
// Generate shared copybooks (1 per 5 programs, at least 2)
const copybookCount = Math.max(2, Math.floor(fileCount / 5));
const copybookNames: string[] = [];
for (let c = 0; c < copybookCount; c++) {
const name = `BENCH${String(c + 1).padStart(4, '0')}`;
copybookNames.push(name);
const copyContent = [
` 01 ${name}-RECORD.`,
` 05 ${name}-KEY PIC X(10).`,
` 05 ${name}-VALUE PIC 9(08).`,
` 05 ${name}-FLAG PIC X(01).`,
'',
].join('\n');
fs.writeFileSync(path.join(copybookDir, `${name}.cpy`), copyContent);
}
for (let f = 0; f < fileCount; f++) {
const programName = `PGM${String(f + 1).padStart(4, '0')}`;
const paragraphs: string[] = [];
for (let p = 0; p < paragraphsPerProgram; p++) {
const paraName = `${String(p + 1).padStart(4, '0')}-PARA`;
// Every paragraph has a PERFORM to the next paragraph (or wraps around)
const nextParaIdx = (p + 1) % paragraphsPerProgram;
const nextParaName = `${String(nextParaIdx + 1).padStart(4, '0')}-PARA`;
const performLine = ` PERFORM ${nextParaName}.`;
// Cross-file CALL: every 3rd paragraph calls another program
const crossFileIdx = (f + p + 1) % fileCount;
const crossProgram = `PGM${String(crossFileIdx + 1).padStart(4, '0')}`;
const callLine =
p % 3 === 0
? ` CALL '${crossProgram}' USING ${copybookNames[p % copybookCount]}-KEY.`
: '';
// COPY in paragraphs adds preprocessing stress — non-idiomatic but
// exercises the preprocessor's expansion path per-paragraph.
const copyLine = ` COPY ${copybookNames[f % copybookCount]}.`;
paragraphs.push(
` ${paraName}.`,
copyLine,
performLine,
callLine,
` DISPLAY '${programName} ${paraName}'.`,
'',
);
}
// Each program COPYs a CONSTANT number of shared copybooks (independent of
// fileCount) so per-file work stays O(1) and the benchmark measures true
// file-count scaling. Copybooks are chosen by program index so they remain
// shared across programs (fan-in), still exercising cross-program copybook
// reuse and multi-COPY-per-program expansion. (Copying ALL copybooks here
// would make per-file work — and emitted data-item nodes — grow with
// fileCount, i.e. O(fileCount²); see the file header.)
const COPYBOOKS_PER_PROGRAM = 3;
const wsCopybooks = [
...new Set(
Array.from(
{ length: COPYBOOKS_PER_PROGRAM },
(_, k) => copybookNames[(f + k) % copybookCount],
),
),
];
const content = [
` IDENTIFICATION DIVISION.`,
` PROGRAM-ID. ${programName}.`,
` ENVIRONMENT DIVISION.`,
` DATA DIVISION.`,
` WORKING-STORAGE SECTION.`,
...wsCopybooks.map((n) => ` COPY ${n}.`),
` PROCEDURE DIVISION.`,
...paragraphs,
` STOP RUN.`,
` END PROGRAM ${programName}.`,
'',
].join('\n');
fs.writeFileSync(path.join(dir, `${programName}.cbl`), content);
}
return { dir, programCount, paragraphCount, copybookCount };
}
async function runBenchmark(
fileCount: number,
paragraphsPerProgram: number,
budgetMs: number,
): Promise<BenchResult> {
const { dir, programCount, paragraphCount, copybookCount } = generateCobolFixture(
fileCount,
paragraphsPerProgram,
);
let peakHeapMB = 0;
const heapSampler = setInterval(() => {
const heap = process.memoryUsage().heapUsed / 1024 / 1024;
if (heap > peakHeapMB) peakHeapMB = heap;
}, 50);
try {
const start = Date.now();
const result = await Promise.race([
runPipelineFromRepo(dir, () => {}, { skipGraphPhases: true }),
new Promise<never>((_, reject) =>
setTimeout(
() => reject(new Error(`Pipeline exceeded ${budgetMs}ms at ${fileCount} files`)),
budgetMs,
),
),
]);
const elapsedMs = Date.now() - start;
return {
fileCount,
programCount,
paragraphCount,
copybookCount,
elapsedMs,
peakHeapMB: Math.round(peakHeapMB),
nodeCount: result.graph.nodeCount,
edgeCount: result.graph.relationshipCount,
};
} finally {
clearInterval(heapSampler);
fs.rmSync(dir, { recursive: true, force: true });
}
}
function printResults(label: string, results: BenchResult[]) {
console.log(`\n${label}`);
console.log(
'┌──────────┬──────────┬────────────┬──────────┬───────────┬──────────┬───────┬───────┐',
);
console.log(
'│ Files │ Programs │ Paragraphs │ Copybooks│ Time (ms) │ Heap MB │ Nodes │ Edges │',
);
console.log(
'├──────────┼──────────┼────────────┼──────────┼───────────┼──────────┼───────┼───────┤',
);
for (const r of results) {
console.log(
`│ ${String(r.fileCount).padStart(8)} │ ${String(r.programCount).padStart(8)} │ ${String(r.paragraphCount).padStart(10)} │ ${String(r.copybookCount).padStart(8)} │ ${String(r.elapsedMs).padStart(9)} │ ${String(r.peakHeapMB).padStart(8)} │ ${String(r.nodeCount).padStart(5)} │ ${String(r.edgeCount).padStart(5)} │`,
);
}
console.log(
'└──────────┴──────────┴────────────┴──────────┴───────────┴──────────┴───────┴───────┘',
);
if (results.length >= 2) {
console.log('\nScaling ratios (time_ratio / file_ratio):');
for (let i = 1; i < results.length; i++) {
const fileRatio = results[i].fileCount / results[i - 1].fileCount;
const timeRatio = results[i].elapsedMs / results[i - 1].elapsedMs;
const scaling = timeRatio / fileRatio;
console.log(
` ${results[i - 1].fileCount} \u2192 ${results[i].fileCount}: ${scaling.toFixed(2)}x (${scaling < 1.5 ? 'linear' : scaling < 3 ? 'superlinear' : 'WARNING: quadratic'})`,
);
}
}
}
describe.skipIf(!BENCH_ENABLED)('COBOL pipeline benchmark', () => {
it('scales with file count', async () => {
const scales = [100, 250, 500, 1000];
const results: BenchResult[] = [];
for (const fileCount of scales) {
const paragraphsPerProgram = 3;
const result = await runBenchmark(fileCount, paragraphsPerProgram, 300_000);
results.push(result);
console.log(
` ${fileCount} files: ${result.elapsedMs}ms, ${result.peakHeapMB}MB heap, ${result.nodeCount} nodes, ${result.edgeCount} edges`,
);
}
printResults('COBOL Pipeline', results);
for (let i = 1; i < results.length; i++) {
const fileRatio = results[i].fileCount / results[i - 1].fileCount;
const timeRatio = results[i].elapsedMs / results[i - 1].elapsedMs;
// Wall-clock is noisy (GC/CI load); keep a coarse upper bound here.
expect(timeRatio / fileRatio).toBeLessThan(4);
// Deterministic regression guard: with constant per-program copybook
// fan-out the emitted node count is exactly linear in fileCount
// (ratio ≈ 1.0). If someone reintroduces O(fileCount²) work — e.g. by
// making every program COPY all copybooks — node growth jumps to ~2x
// per file-doubling and this fails. Node count is deterministic, so
// this is a non-flaky guard unlike the wall-clock check above.
const nodeRatio = results[i].nodeCount / results[i - 1].nodeCount;
expect(nodeRatio / fileRatio).toBeLessThan(1.3);
}
}, 600_000);
});