Merge remote-tracking branch 'origin/main' into feat/phase8-9-implementation

# Conflicts:
#	gitnexus/src/core/ingestion/languages/csharp.ts
#	gitnexus/src/core/ingestion/languages/dart.ts
#	gitnexus/src/core/ingestion/languages/kotlin.ts
#	gitnexus/src/core/ingestion/languages/python.ts
#	gitnexus/src/core/ingestion/languages/ruby.ts
#	gitnexus/src/core/ingestion/languages/rust.ts
#	gitnexus/src/core/ingestion/languages/typescript.ts
This commit is contained in:
Gergo Magyar 2026-03-26 14:06:25 +00:00
commit dddfd0d789
69 changed files with 11315 additions and 391 deletions

View file

@ -375,6 +375,7 @@ GitNexus builds a complete knowledge graph of your codebase through a multi-phas
| Swift | — | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| C | — | — | ✓ | — | ✓ | ✓ | — | ✓ | ✓ |
| C++ | — | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| Dart | ✓ | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
**Imports** — cross-file import resolution · **Named Bindings** — `import { X as Y }` / re-export tracking · **Exports** — public/exported symbol detection · **Heritage** — class inheritance, interfaces, mixins · **Type Annotations** — explicit type extraction for receiver resolution · **Constructor Inference** — infer receiver type from constructor calls (`self`/`this` resolution included for all languages) · **Config** — language toolchain config parsing (tsconfig, go.mod, etc.) · **Frameworks** — AST-based framework pattern detection · **Entry Points** — entry point scoring heuristics
@ -539,7 +540,7 @@ The wiki generator reads the indexed graph structure, groups files into modules
- [X] Constructor-Inferred Type Resolution, `self`/`this` Receiver Mapping
- [X] Wiki Generation, Multi-File Rename, Git-Diff Impact Analysis
- [X] Process-Grouped Search, 360-Degree Context, Claude Code Hooks
- [X] Multi-Repo MCP, Zero-Config Setup, 13 Language Support
- [X] Multi-Repo MCP, Zero-Config Setup, 14 Language Support
- [X] Community Detection, Process Detection, Confidence Scoring
- [X] Hybrid Search, Vector Index

View file

@ -0,0 +1,100 @@
# COBOL Code Indexing
GitNexus indexes COBOL codebases using a **regex-only extraction** strategy, bypassing tree-sitter entirely. This document explains why, how the pipeline works, and links to detailed sub-documents.
## Why Regex-Only?
The tree-sitter-cobol grammar (v0.0.1) has three critical limitations that make it unusable for production indexing:
| Issue | Impact | Severity |
|-------|--------|----------|
| External scanner hangs on ~5% of files | No timeout mechanism exists for the C scanner; the process blocks indefinitely | **Blocking** |
| Only ~15% of paragraph headers detected | Most procedure-division paragraphs are invisible to the grammar | High |
| Patch markers in cols 1-6 cause parse errors | Enterprise COBOL uses non-standard sequence area content (e.g., `mzADD`, `estero`, `#FIX`) | High |
Because the external scanner hang cannot be interrupted (there is no `setTimeoutMicros` equivalent for tree-sitter), using tree-sitter-cobol would hang the indexing pipeline on a non-trivial fraction of real-world files.
The regex-only approach provides:
- **Speed**: ~1ms per file average extraction time
- **Reliability**: zero hangs, zero crashes across 13,000+ files
- **Coverage**: captures all critical symbols -- program name, paragraphs, sections, CALL, PERFORM, COPY, data items (01-77, 88-level), file declarations, FD entries, EXEC SQL/CICS blocks, ENTRY points, and MOVE statements
## Architecture
```mermaid
flowchart TD
A[Repository Scan] --> B{File Detection}
B -->|Extension match| C[COBOL file]
B -->|GITNEXUS_COBOL_DIRS match| C
B -->|No match| Z[Skip]
C --> D{Copybook?}
D -->|Yes| E[Add to Copybook Map]
D -->|No| F[Source Program]
E --> G[COPY Expansion Engine]
F --> G
G -->|Inline copybook content| H[Expanded Source]
H --> I[Patch Marker Cleanup]
I --> J[Regex State Machine]
J --> K[Extracted Symbols]
K --> L[Graph Model Builder]
L --> M[Knowledge Graph]
subgraph "Per-Chunk Processing"
G
H
I
J
K
L
end
subgraph "Post-Processing"
M --> N[Community Detection]
M --> O[Process Detection]
M --> P[Contract Detection]
end
style J fill:#e8f5e9,stroke:#2e7d32
style G fill:#e3f2fd,stroke:#1565c0
```
## COBOL vs Tree-Sitter Languages
| Feature | COBOL (Regex) | Tree-Sitter Languages |
|---------|--------------|----------------------|
| Parser | Single-pass regex state machine | tree-sitter grammar + queries |
| Speed | ~1ms/file | ~5ms/file |
| AST available | No | Yes |
| COPY expansion | Yes (pre-processing step) | N/A |
| Deep indexing | Data items, SQL, CICS, FD, ENTRY | Type annotations, generics, etc. |
| Call extraction | PERFORM (intra-file) + CALL (cross-program) | AST-based call site detection |
| Import extraction | COPY statements | `import`/`require`/`use`/`#include` |
| Coverage | All critical symbols | Language-dependent query coverage |
| Failure mode | Never hangs | External scanner can hang (COBOL only) |
## Sub-Documents
| Document | Description |
|----------|-------------|
| [File Detection](./file-detection.md) | Extension mapping, `GITNEXUS_COBOL_DIRS`, copybook classification |
| [COPY Expansion](./copy-expansion.md) | Copybook inlining, REPLACING transformations, cycle detection |
| [Regex Extraction](./regex-extraction.md) | State machine, regex patterns, line processing |
| [Deep Indexing](./deep-indexing.md) | Data items, EXEC SQL/CICS, file declarations, FD, ENTRY, MOVE |
| [Graph Model](./graph-model.md) | COBOL-specific node types, edge types, full annotated example |
| [Performance](./performance.md) | Benchmarks, worker pool tuning, caps, troubleshooting |
## Key Source Files
| File | Purpose |
|------|---------|
| `gitnexus/src/core/ingestion/cobol-preprocessor.ts` | Patch marker cleanup + regex extraction engine |
| `gitnexus/src/core/ingestion/cobol-copy-expander.ts` | COPY statement expansion with REPLACING |
| `gitnexus/src/core/ingestion/utils.ts` | `getLanguageFromPath`, `getLanguageFromFilename` |
| `gitnexus/src/core/ingestion/pipeline.ts` | `isCobolCopybook`, `expandCobolCopies`, `detectCrossProgamContracts` |
| `gitnexus/src/core/ingestion/workers/parse-worker.ts` | `processCobolRegexOnly` -- graph model builder |
| `gitnexus/src/core/ingestion/workers/worker-pool.ts` | Configurable sub-batch size for COBOL |

View file

@ -0,0 +1,157 @@
# COBOL COPY Expansion
The COPY statement is COBOL's include mechanism -- analogous to `#include` in C or `import` in modern languages. GitNexus expands COPY statements **before** regex extraction so that symbols defined inside copybooks (data items, paragraphs, etc.) are visible in the program's extracted graph.
## Supported Syntax
### Basic COPY
```cobol
COPY CPSESP.
COPY "WORKGRID.CPY".
```
Inlines the content of the named copybook, replacing the COPY line(s).
### COPY with REPLACING
```cobol
COPY CPSESP REPLACING "ANAZI-KEY" BY "LK-KEY".
COPY CPSESP REPLACING LEADING "ESP-" BY "LK-ESP-"
LEADING "KPSESPL" BY "LK-KPSESPL".
COPY LINKAGE REPLACING TRAILING "-IN" BY "-OUT".
```
Three REPLACING types are supported:
| Type | Syntax | Behavior | Example |
| ------------ | ------------------------------------ | --------------------------------------- | -------------------------------- |
| **EXACT** | `REPLACING "OLD" BY "NEW"` | Replace exact identifier matches | `ANAZI-KEY` becomes `LK-KEY` |
| **LEADING** | `REPLACING LEADING "PFX-" BY "NEW-"` | Replace prefix on all COBOL identifiers | `ESP-NAME` becomes `LK-ESP-NAME` |
| **TRAILING** | `REPLACING TRAILING "-IN" BY "-OUT"` | Replace suffix on all COBOL identifiers | `DATA-IN` becomes `DATA-OUT` |
Multiple REPLACING clauses can appear in a single COPY statement. They are applied in order to each COBOL identifier in the copybook content.
### Multi-Line COPY
COPY statements can span multiple lines (standard COBOL continuation rules apply):
```cobol
COPY CPSESP REPLACING
- LEADING "ESP-" BY "LK-ESP-"
- LEADING "KPSESPL" BY "LK-KPSESPL".
```
Continuation lines (indicator `-` in column 7) are merged before COPY statement scanning.
## Expansion Flow
```mermaid
sequenceDiagram
participant Pipeline
participant Expander as COPY Expander
participant Resolver
participant Reader
Pipeline->>Pipeline: Identify all COBOL files
Pipeline->>Pipeline: Classify copybooks vs programs
Pipeline->>Reader: Read all copybook content upfront
Reader-->>Pipeline: Copybook content map (name -> content)
loop For each source file in chunk
Pipeline->>Expander: expandCopies(content, filePath, resolveFile, readFile)
Expander->>Expander: Merge continuation lines
Expander->>Expander: Detect COPY statements via regex
loop For each COPY statement (reverse order)
Expander->>Resolver: resolveFile(copyTarget)
Resolver-->>Expander: Copybook key or null
alt Resolved successfully
Expander->>Reader: readFile(resolvedKey)
Reader-->>Expander: Copybook content
Expander->>Expander: Apply REPLACING transformations
Expander->>Expander: Recurse for nested COPYs (depth + 1)
Expander->>Expander: Splice expanded content into output
else Not resolved
Expander->>Expander: Keep original COPY line
end
end
Expander-->>Pipeline: Expanded content + resolution metadata
Pipeline->>Pipeline: Replace file content with expanded content
end
```
The return type `CopyExpansionResult` contains `expandedContent` and `copyResolutions`. The `expansionDepth` field has been removed from the return type (it was unused by callers).
COPY statement line numbers in `CopyResolution` are 1-based (consistent with the preprocessor's line numbering). The splice operation that replaces COPY lines with expanded content adjusts for 0-based array indexing internally.
## Cycle Detection
Circular COPY references (e.g., copybook A includes copybook B which includes copybook A) are detected and handled:
1. Each expansion chain maintains a `visited` set of resolved copybook paths
2. If a copybook path is already in the visited set, the expansion is skipped
3. A `warnedCircular` set (internal to `expandCopies()`, not a parameter) deduplicates warning messages within a single file expansion
Known circular copybooks in PROJECT-NAME: `ANAZI`, `ANDIP`, `QDIPE` (self-referential includes).
## Max Depth
Nested COPY expansion is limited to **10 levels** (`DEFAULT_MAX_DEPTH`). If a COPY chain exceeds this depth, a warning is logged and the remaining COPY statements are left unexpanded.
## Max Total Expansions
A breadth amplification guard caps the total number of COPY expansions across all branches within a single file to **500** (`MAX_TOTAL_EXPANSIONS`). This prevents exponential blowup from diamond-shaped COPY graphs where N copybooks each include N other copybooks. Once the limit is reached, further COPY statements in that file are left unexpanded and a single warning is logged.
## REPLACING Application Detail
The REPLACING engine works by scanning all COBOL identifiers (matching `\b[A-Z][A-Z0-9-]*\b`) in the copybook content and applying each replacement rule:
```
Original copybook content:
05 ESP-NAME PIC X(30).
05 ESP-CODE PIC X(10).
05 KPSESPL-FLAG PIC X(01).
After REPLACING LEADING "ESP-" BY "LK-ESP-" LEADING "KPSESPL" BY "LK-KPSESPL":
05 LK-ESP-NAME PIC X(30).
05 LK-ESP-CODE PIC X(10).
05 LK-KPSESPL-FLAG PIC X(01).
```
For LEADING replacements, the engine checks if each identifier starts with the `from` prefix (case-insensitive) and replaces only the prefix portion, preserving the rest of the identifier.
For TRAILING replacements, the same logic applies to suffixes.
For EXACT replacements, only identifiers that match the `from` value exactly (case-insensitive) are replaced.
## Copybook Resolution
The resolver tries multiple strategies to match a COPY target name to a copybook file:
1. **Exact match**: `COPY CPSESP` resolves to copybook named `CPSESP`
2. **Strip extension**: `COPY WORKGRID.CPY` strips `.CPY` and resolves to `WORKGRID`
3. **Add extension**: `COPY CPSESP` tries `CPSESP.CPY` and `CPSESP.COPY`
If no match is found, the COPY statement is left in place (unexpanded) and a resolution record with `resolvedPath: null` is created.
## Pipeline Integration
The expansion runs **per chunk**, after file content is read but before dispatch to worker threads:
1. All copybook files are read upfront (they are typically small, collectively under 100MB)
2. Per chunk, the copybook map is merged with chunk content (in case a chunk contains copybooks)
3. Only programs (not copybooks themselves) undergo expansion
4. The expanded content replaces the original content in-place before worker dispatch
## Inline Comment Handling
The copy expander's `stripInlineComment()` helper is quote-aware: pipe characters (`|`) inside single- or double-quoted strings are preserved. This matches the same quote-aware logic used by the preprocessor.
## Source Files
- `gitnexus/src/core/ingestion/cobol-copy-expander.ts` -- `expandCopies()`, `parseReplacingClause()`, `applyReplacing()`
- `gitnexus/src/core/ingestion/pipeline.ts` -- `expandCobolCopies()`, copybook map construction, chunk integration

View file

@ -0,0 +1,312 @@
# COBOL Deep Indexing
Beyond basic symbol extraction (program name, paragraphs, CALL, PERFORM, COPY), GitNexus performs deep indexing of COBOL-specific constructs: data items, EXEC SQL/CICS blocks, file declarations, FD entries, ENTRY points, and MOVE statements.
## Data Items
### Level Numbers
| Level Range | Meaning | Graph Node Type |
|-------------|---------|-----------------|
| 01 | Record (group item) | `Record` |
| 02-49 | Elementary/group items | `Property` |
| 66 | RENAMES | `Property` |
| 77 | Independent item | `Property` |
| 88 | Condition name | `Const` |
FILLER items are skipped (no useful name for the graph).
### Clauses Parsed
The `parseDataItemClauses()` function extracts these clauses from the trailing text of a data item declaration:
| Clause | Pattern | Example |
|--------|---------|---------|
| `PIC` / `PICTURE` | `\bPIC(?:TURE)?\s+(?:IS\s+)?(\S+)` | `PIC X(30)`, `PICTURE IS 9(5)V99` |
| `USAGE` | `\bUSAGE\s+(?:IS\s+)?(COMP\|BINARY\|...)` | `USAGE IS COMP-3`, `BINARY` |
| `REDEFINES` | `\bREDEFINES\s+([A-Z][A-Z0-9-]+)` | `REDEFINES WK-DATE-NUM` |
| `OCCURS` | `\bOCCURS\s+(\d+)` | `OCCURS 12 TIMES` |
Standalone COMP variants (without the `USAGE` keyword) are also detected: `COMP`, `COMP-1` through `COMP-6`, `COMP-X`, `BINARY`, `PACKED-DECIMAL`.
### Data Hierarchy
Data items form a hierarchical structure based on level numbers. The extractor uses a **stack algorithm**:
```
Processing order:
01 WK-RECORD -> push {01, WK-RECORD} -> parent: Module
05 WK-NAME -> push {05, WK-NAME} -> parent: WK-RECORD (01 < 05)
10 WK-FIRST -> push {10, WK-FIRST} -> parent: WK-NAME (05 < 10)
10 WK-LAST -> pop WK-FIRST, push -> parent: WK-NAME (05 < 10)
05 WK-CODE -> pop WK-LAST, WK-NAME -> parent: WK-RECORD (01 < 05)
88 WK-ACTIVE -> (88 handled separately) -> parent: WK-CODE
```
The stack maintains items where each entry's level is strictly less than the next. When a new item arrives with a level <= the top of stack, items are popped until the stack top has a smaller level. A `CONTAINS` edge is created from the stack top to the new item.
For 88-level condition names, the parent is the immediately preceding non-88 data item (found by scanning backwards).
### Annotated Example
```cobol
01 WK-EMPLOYEE.
05 WK-EMP-ID PIC 9(6).
05 WK-EMP-NAME PIC X(30).
05 WK-EMP-STATUS PIC X(01).
88 WK-ACTIVE VALUE "A".
88 WK-INACTIVE VALUE "I".
05 WK-SALARY PIC 9(7)V99 COMP-3.
05 WK-DEPT PIC X(04) OCCURS 3 TIMES.
```
Produces:
- `Record` node: `WK-EMPLOYEE` (level 01, section: working-storage)
- `Property` nodes: `WK-EMP-ID`, `WK-EMP-NAME`, `WK-EMP-STATUS`, `WK-SALARY`, `WK-DEPT`
- `Const` nodes: `WK-ACTIVE` (values: `A`), `WK-INACTIVE` (values: `I`)
- `CONTAINS` edges: `WK-EMPLOYEE -> WK-EMP-ID`, `WK-EMPLOYEE -> WK-EMP-NAME`, etc.
- `CONTAINS` edges: `WK-EMP-STATUS -> WK-ACTIVE`, `WK-EMP-STATUS -> WK-INACTIVE`
### Data Item Cap
A maximum of **500 data items per file** (`MAX_DATA_ITEMS_PER_FILE`) are processed. Some COBOL programs (especially after COPY expansion) can have 10,000+ data items, which would cause graph bloat and push the V8 relationship Map past its 16.7M entry limit across thousands of files.
The cap applies after extraction: the first 500 items in source order are kept. Since 01-level records appear first, critical top-level structure is preserved.
## EXEC SQL
EXEC SQL blocks are accumulated across lines between `EXEC SQL` and `END-EXEC`, then parsed as a unit.
### Operation Classification
The first SQL keyword determines the operation:
| First Keyword | Operation |
|---------------|-----------|
| `SELECT` | SELECT |
| `INSERT` | INSERT |
| `UPDATE` | UPDATE |
| `DELETE` | DELETE |
| `DECLARE` | DECLARE |
| `OPEN` | OPEN |
| `CLOSE` | CLOSE |
| `FETCH` | FETCH |
| *(anything else)* | OTHER |
### Table Extraction
Tables are extracted from SQL clauses:
| Clause Pattern | Example |
|----------------|---------|
| `FROM <table>` | `SELECT * FROM EMPLOYEES` |
| `INSERT INTO <table>` | `INSERT INTO EMPLOYEES` |
| `UPDATE <table>` | `UPDATE EMPLOYEES SET ...` |
| `JOIN <table>` | `LEFT JOIN DEPARTMENTS ON ...` |
Note: The `INTO` pattern is restricted to `INSERT INTO` to avoid false positives from `FETCH ... INTO :host-var` and `SELECT ... INTO :host-var` statements, where `INTO` introduces host variables rather than table names.
### Cursor Detection
```cobol
EXEC SQL
DECLARE C-EMPLOYEES CURSOR FOR
SELECT EMP-ID, EMP-NAME FROM EMPLOYEES
WHERE DEPT = :WK-DEPT
END-EXEC
```
Extracts: cursor `C-EMPLOYEES`, table `EMPLOYEES`, host variable `WK-DEPT`.
### Host Variables
Host variables are COBOL variables referenced in SQL with a `:` prefix. The colon is stripped:
```sql
WHERE EMP-ID = :WK-EMP-ID AND DEPT = :WK-DEPT
```
Extracts: `WK-EMP-ID`, `WK-DEPT`.
### Graph Output
- `CodeElement` node per table, with description `sql-table op:{OP}`
- `CodeElement` node per cursor, with description `sql-cursor`
- `ACCESSES` edge from Module to each CodeElement
- Deduplication: if the same table appears in multiple SQL blocks, only one node is created
## EXEC CICS
EXEC CICS blocks are accumulated and parsed similarly to SQL blocks.
### Command Detection
Two-word commands are detected first (matched against the block start):
```
SEND MAP, RECEIVE MAP, SEND TEXT, SEND CONTROL, READ NEXT, READ PREV
```
If no two-word command matches, the first word is used (e.g., `LINK`, `XCTL`, `RETURN`, `READ`, `WRITE`).
### Extraction
| Element | Pattern | Example |
|---------|---------|---------|
| MAP name | `MAP('name')` or `MAP("name")` | `EXEC CICS SEND MAP('EMPMENU')` |
| PROGRAM name | `PROGRAM('name')` or `PROGRAM("name")` | `EXEC CICS LINK PROGRAM('BGTABUP')` |
| TRANSID | `TRANSID('name')` or `TRANSID("name")` | `EXEC CICS START TRANSID('EMP1')` |
### Graph Output
- MAP: `CodeElement` node with description `cics-map cmd:{CMD}` + `ACCESSES` edge from Module
- PROGRAM: `CALLS` edge (cross-program call via CICS LINK/XCTL)
- TRANSID: `CodeElement` node with description `cics-transid cmd:{CMD}` + `ACCESSES` edge from Module
### Annotated Example
```cobol
EXEC CICS
SEND MAP('EMPMENU')
MAPSET('EMPSET')
FROM(WK-MAP-DATA)
ERASE
END-EXEC
```
Produces:
- `CodeElement` node: `EMPMENU` (description: `cics-map cmd:SEND MAP`)
- `ACCESSES` edge: Module -> `EMPMENU`
## File Declarations
SELECT statements in the INPUT-OUTPUT SECTION are accumulated across multiple lines (until a period terminator) and parsed for:
| Clause | Pattern | Example |
|--------|---------|---------|
| SELECT | `SELECT <name>` | `SELECT MASTER-FILE` |
| ASSIGN | `ASSIGN TO <file>` | `ASSIGN TO "MASTER.DAT"` |
| ORGANIZATION | `ORGANIZATION IS <type>` | `ORGANIZATION IS INDEXED` |
| ACCESS | `ACCESS MODE IS <mode>` | `ACCESS MODE IS DYNAMIC` |
| RECORD KEY | `RECORD KEY IS <field>` | `RECORD KEY IS WK-EMP-ID` |
| FILE STATUS | `FILE STATUS IS <field>` | `FILE STATUS IS WK-FILE-STATUS` |
### Graph Output
- `CodeElement` node with description containing all parsed clauses (e.g., `select org:INDEXED access:DYNAMIC key:WK-EMP-ID status:WK-FILE-STATUS assign:MASTER.DAT`)
- `RECORD_KEY_OF` edge: from Property node to CodeElement (confidence 0.8)
- `FILE_STATUS_OF` edge: from Property node to CodeElement (confidence 0.8)
## FD Entries
FD (File Description) entries associate a file name with its record layout:
```cobol
FD MASTER-FILE.
01 MASTER-RECORD.
05 MR-EMP-ID PIC 9(6).
05 MR-EMP-NAME PIC X(30).
```
The extractor tracks `pendingFdName` state: when an `FD` line is seen, the next 01-level data item becomes its record.
### Graph Output
- `CodeElement` node with description `fd record:{recordName}`
- `CONTAINS` edge: FD CodeElement -> Record node
- `CONTAINS` edge: SELECT CodeElement -> FD CodeElement (linking file declaration to file description)
## ENTRY Points
The `ENTRY` statement defines additional entry points into a COBOL program (in addition to the main program entry):
```cobol
ENTRY "SUBPROG" USING WK-PARAM-1 WK-PARAM-2.
```
### Graph Output
- `Constructor` node with description `entry params:{param1},{param2}` (or just `entry` if no parameters)
- `CONTAINS` edge: Module -> Constructor
- Symbol table entry (so the entry point is discoverable by name)
## PROCEDURE DIVISION USING
```cobol
PROCEDURE DIVISION USING WK-INPUT-REC WK-OUTPUT-REC.
```
The USING clause identifies parameters received by the program from its caller.
### Graph Output
- `RECEIVES` edge: Module -> Property (for each parameter name, confidence 0.8)
## MOVE Statements
MOVE statements produce `ACCESSES` edges in the graph:
```cobol
MOVE WK-NAME TO OUT-NAME.
MOVE CORRESPONDING WK-INPUT TO WK-OUTPUT.
MOVE CORR WK-IN TO WK-OUT.
```
### Extraction Details
- Source and target identifiers are captured
- `CORRESPONDING` and its abbreviation `CORR` are both recognized (bulk field-by-field move)
- Figurative constants (SPACES, ZEROS, LOW-VALUES, HIGH-VALUES, QUOTES, ALL) are skipped
- The enclosing paragraph (`caller`) is tracked for context
### MOVE CORRESPONDING / CORR Edge Reasons
MOVE CORRESPONDING (and CORR) produces distinct edge reasons to differentiate from simple MOVE:
| Edge | Reason (simple MOVE) | Reason (CORRESPONDING/CORR) |
|------|---------------------|-----------------------------|
| Read (source) | `cobol-move-read` | `cobol-move-corresponding-read` |
| Write (target) | `cobol-move-write` | `cobol-move-corresponding-write` |
This distinction allows queries to find bulk field-by-field moves separately from simple variable assignments.
## GO TO DEPENDING ON
The `GO TO` statement with multiple targets and a `DEPENDING ON` clause is a computed branch:
```cobol
GO TO PARA-1 PARA-2 PARA-3
DEPENDING ON WK-SELECTOR.
```
All target paragraph names are extracted and emitted as separate `gotos` entries. Each target produces a `CALLS` edge in the graph (same semantics as PERFORM). The `DEPENDING ON` variable is not currently tracked as a data-flow dependency.
## SORT INPUT/OUTPUT PROCEDURE
SORT and MERGE statements can specify procedural entry points instead of file-based I/O:
```cobol
SORT SORT-FILE ON ASCENDING KEY SORT-KEY
INPUT PROCEDURE IS PREPARE-INPUT
OUTPUT PROCEDURE IS FORMAT-OUTPUT.
```
`INPUT PROCEDURE IS` and `OUTPUT PROCEDURE IS` targets are extracted as control-flow targets (same as PERFORM). They produce `performs` entries and corresponding `CALLS` edges in the graph.
## Fixed-Format Literal Continuation
In fixed-format COBOL, string literals can span multiple lines using the continuation indicator (`-` in column 7). When a continuation line starts with a quote character, the extractor joins it with the predecessor by removing the trailing quote from the previous line and the opening quote from the continuation:
```
Line N: MOVE "THIS IS A LONG STRI
Line N+1 (cont): - "NG VALUE" TO WK-FIELD.
Merged: MOVE "THIS IS A LONG STRING VALUE" TO WK-FIELD.
```
The trailing `"` on line N and the opening `"` on line N+1 are both removed, producing a seamless literal. If no matching quote is found on the predecessor line, the continuation is appended as-is.
## Source Files
- `gitnexus/src/core/ingestion/cobol-preprocessor.ts` -- All extraction logic, clause parsers, EXEC block parsers
- `gitnexus/src/core/ingestion/workers/parse-worker.ts` -- `processCobolRegexOnly()`, graph node/edge emission
- `gitnexus/src/core/ingestion/parsing-processor.ts` -- Sequential fallback with same `MAX_DATA_ITEMS_PER_FILE` cap

View file

@ -0,0 +1,126 @@
# COBOL File Detection
GitNexus detects COBOL files through two mechanisms: extension-based mapping and directory-based override for extensionless files. This document covers both, plus the copybook/program classification logic.
## Extension Mapping
### Program Extensions
| Extension | Type |
|-----------|------|
| `.cbl` | COBOL program |
| `.cob` | COBOL program |
| `.cobol` | COBOL program |
### Copybook Extensions
| Extension | Type | Notes |
|-----------|------|-------|
| `.cpy` | Copybook | Standard |
| `.copy` | Copybook | Standard |
| `.gnm` / `.GNM` | Copybook | Enterprise (GnuCOBOL naming) |
| `.fd` / `.FD` | Copybook | File Description fragment |
| `.wrk` / `.WRK` | Copybook | Working-Storage fragment |
| `.sel` / `.SEL` | Copybook | SELECT clause fragment |
| `.open` / `.OPEN` | Copybook | File OPEN fragment |
| `.close` / `.CLOSE` | Copybook | File CLOSE fragment |
| `.ini` / `.INI` | Copybook | Initialization fragment |
| `.def` / `.DEF` | Copybook | Definition fragment |
All extension matching is case-sensitive in `getLanguageFromFilename` (the extensions above are matched as written, including uppercase variants like `.GNM`).
## Extensionless File Detection: `GITNEXUS_COBOL_DIRS`
Many enterprise COBOL repositories use extensionless files -- the filename alone identifies the program (e.g., `s/BGTABFL` is the source for program `BGTABFL`). GitNexus handles this via the `GITNEXUS_COBOL_DIRS` environment variable.
### Configuration
Set `GITNEXUS_COBOL_DIRS` to a comma-separated list of directory names:
```bash
# Files in s/, c/, and wfproc/ directories (at any depth) are treated as COBOL
export GITNEXUS_COBOL_DIRS=s,c,wfproc
```
The matching is **case-insensitive** and checks all path segments:
- `/repo/s/BGTABFL` -- matches segment `s` -- COBOL
- `/repo/src/c/CPSESP` -- matches segment `c` -- COBOL
- `/repo/wfproc/WF001` -- matches segment `wfproc` -- COBOL
- `/repo/docs/README` -- no matching segment -- skipped
### Decision Tree
```mermaid
flowchart TD
A[getLanguageFromPath] --> B[getLanguageFromFilename]
B --> C{Known extension?}
C -->|Yes .cbl/.cob/.cobol/.cpy/...| D[Return COBOL]
C -->|Yes .ts/.py/.java/...| E[Return other language]
C -->|No match| F{Has extension?}
F -->|"Has dot in basename"| G[Return null]
F -->|"No dot = extensionless"| H{GITNEXUS_COBOL_DIRS set?}
H -->|No| G
H -->|Yes| I{Any path segment<br/>matches a configured dir?}
I -->|Yes| D
I -->|No| G
style D fill:#e8f5e9,stroke:#2e7d32
style G fill:#ffebee,stroke:#c62828
```
### Implementation Detail
The `GITNEXUS_COBOL_DIRS` value is parsed once (on first call) and cached in a `Set<string>`:
```typescript
// From gitnexus/src/core/ingestion/utils.ts
const getCobolDirs = (): Set<string> => {
if (_cobolDirs) return _cobolDirs;
const raw = process.env.GITNEXUS_COBOL_DIRS;
_cobolDirs = raw
? new Set(raw.split(',').map(d => d.trim().toLowerCase()))
: new Set();
return _cobolDirs;
};
```
The path segment check splits the full path on `/` and tests each segment against the cached set.
## Copybook vs Program Classification
After a file is identified as COBOL, it must be classified as either a **program** (to be parsed for symbols) or a **copybook** (to be loaded into the copybook map for COPY expansion).
### Classification Rules
A COBOL file is classified as a **copybook** if ANY of these conditions is true:
1. It has a recognized copybook extension (`.cpy`, `.copy`, `.gnm`, `.fd`, `.wrk`, `.sel`, `.open`, `.close`, `.ini`, `.def`)
2. It is an extensionless file whose path contains a directory segment matching one of: `c`, `copy`, `copybooks`, `copylib`, `cpy`
A file is classified as a **program** if:
1. It has a program extension (`.cbl`, `.cob`, `.cobol`), OR
2. It is extensionless and does NOT match any copybook directory pattern
### Copybook Name Resolution
Copybook names are derived from the filename:
- Strip the extension (if any)
- Convert to uppercase
Examples:
- `c/CPSESP` -- name: `CPSESP`
- `copy/workgrid.cpy` -- name: `WORKGRID`
- `c/ANAZI.GNM` -- name: `ANAZI`
This name is used to resolve `COPY CPSESP.` statements during expansion.
## Source Files
- `gitnexus/src/core/ingestion/utils.ts` -- `getLanguageFromPath()`, `getLanguageFromFilename()`, `getCobolDirs()`
- `gitnexus/src/core/ingestion/pipeline.ts` -- `isCobolCopybook()`, `getCopybookName()`, `COPYBOOK_EXTENSIONS`, `COBOL_PROGRAM_EXTENSIONS`

View file

@ -0,0 +1,193 @@
# COBOL Graph Model
This document describes the graph nodes and edges that GitNexus creates for COBOL codebases. The COBOL graph model is richer than most tree-sitter languages because it captures domain-specific constructs: file declarations, FD entries, data hierarchies, SQL tables, CICS maps, and cross-program contracts.
## Entity-Relationship Diagram
```mermaid
erDiagram
File ||--o{ Module : DEFINES
File ||--o{ Function : DEFINES
File ||--o{ Namespace : DEFINES
File ||--o{ Record : DEFINES
File ||--o{ Property : DEFINES
File ||--o{ Const : DEFINES
File ||--o{ CodeElement : DEFINES
File ||--o{ Constructor : DEFINES
File }o--o{ File : IMPORTS
Module ||--o{ Record : CONTAINS
Module ||--o{ Constructor : CONTAINS
Module }o--o{ CodeElement : ACCESSES
Module }o--o{ Module : CALLS
Module }o--o{ Module : CONTRACTS
Module }o--o{ Property : RECEIVES
Record ||--o{ Property : CONTAINS
Record ||--o{ Const : CONTAINS
Record }o--o{ Record : REDEFINES
Property ||--o{ Property : CONTAINS
Property ||--o{ Const : CONTAINS
Property }o--o{ Property : REDEFINES
Property }o--o{ CodeElement : RECORD_KEY_OF
Property }o--o{ CodeElement : FILE_STATUS_OF
CodeElement ||--o{ CodeElement : CONTAINS
CodeElement ||--o{ Record : CONTAINS
Function }o--o{ Function : CALLS
```
## Node Types
| Node Type | COBOL Concept | Created From | Example |
|-----------|--------------|--------------|---------|
| `Module` | PROGRAM-ID | `PROGRAM-ID. BGTABFL` | Name: `BGTABFL`, description may include author and date |
| `Function` | Paragraph | `PROCESS-RECORD.` at column 8 | Name: `PROCESS-RECORD` |
| `Namespace` | Procedure section | `MAIN-LOGIC SECTION.` at column 8 | Name: `MAIN-LOGIC` |
| `Record` | 01-level data item | `01 WK-EMPLOYEE.` | Description: `level:01 section:working-storage` |
| `Property` | 02-49/66/77 data item | `05 WK-NAME PIC X(30).` | Description: `level:05 pic:X(30) section:working-storage` |
| `Const` | 88-level condition | `88 WK-ACTIVE VALUE "A".` | Description: `level:88 values:A` |
| `CodeElement` | SELECT, FD, SQL table, CICS map, cursor, transid | Various | Description varies by subtype |
| `Constructor` | ENTRY point | `ENTRY "SUBPROG" USING WK-DATA` | Description: `entry params:WK-DATA` |
### CodeElement Subtypes
CodeElement is used for multiple COBOL constructs, distinguished by their description prefix:
| Subtype | ID Pattern | Description Format | Example |
|---------|-----------|-------------------|---------|
| File SELECT | `CodeElement:{path}:SELECT:{name}` | `select org:INDEXED access:DYNAMIC ...` | `SELECT MASTER-FILE` |
| FD entry | `CodeElement:{path}:FD:{name}` | `fd record:{recordName}` | `FD MASTER-FILE` |
| SQL table | `CodeElement:{path}:sql-table:{name}` | `sql-table op:SELECT` | Table `EMPLOYEES` |
| SQL cursor | `CodeElement:{path}:sql-cursor:{name}` | `sql-cursor` | Cursor `C-EMPLOYEES` |
| CICS map | `CodeElement:{path}:cics-map:{name}` | `cics-map cmd:SEND MAP` | Map `EMPMENU` |
| CICS transid | `CodeElement:{path}:cics-transid:{name}` | `cics-transid cmd:START` | Transid `EMP1` |
## Edge Types
| Edge Type | Source | Target | Created By | Confidence | Example |
|-----------|--------|--------|-----------|------------|---------|
| `DEFINES` | File | any node | File defines its symbols | 1.0 | File -> Module `BGTABFL` |
| `CALLS` | Function | Function | `PERFORM X [THRU Y]` | (via call-processor) | `PROCESS-RECORD` -> `CALC-TAX` |
| `CALLS` | Module | Module | `CALL "BGTABUP"` | (via call-processor) | `BGTABFL` -> `BGTABUP` |
| `CALLS` | Module | Module | `EXEC CICS LINK PROGRAM('X')` | (via call-processor) | `BGTABFL` -> `BGTABUP` |
| `IMPORTS` | File | File | `COPY copybook` | (via import-processor) | Source file -> Copybook file |
| `CONTAINS` | Module | Record | Data hierarchy root | 1.0 | `BGTABFL` -> `WK-EMPLOYEE` |
| `CONTAINS` | Record | Property | Data hierarchy | 1.0 | `WK-EMPLOYEE` -> `WK-NAME` |
| `CONTAINS` | Property | Property | Nested data items | 1.0 | `WK-ADDRESS` -> `WK-CITY` |
| `CONTAINS` | Record/Property | Const | 88-level parent | 1.0 | `WK-STATUS` -> `WK-ACTIVE` |
| `CONTAINS` | CodeElement (FD) | Record | FD record link | 1.0 | `FD:MASTER-FILE` -> `MASTER-RECORD` |
| `CONTAINS` | CodeElement (SELECT) | CodeElement (FD) | SELECT-FD link | 0.9 | `SELECT:MASTER-FILE` -> `FD:MASTER-FILE` |
| `CONTAINS` | Module | Constructor | ENTRY in module | 1.0 | `BGTABFL` -> `SUBPROG` |
| `REDEFINES` | Record | Record | `01 X REDEFINES Y` | 1.0 | `WK-DATE-NUM` -> `WK-DATE-ALPHA` |
| `REDEFINES` | Property | Property | `05 X REDEFINES Y` | 1.0 | `WK-CODE-NUM` -> `WK-CODE-ALPHA` |
| `RECORD_KEY_OF` | Property | CodeElement (SELECT) | `RECORD KEY IS field` | 0.8 | `WK-EMP-ID` -> `SELECT:MASTER-FILE` |
| `FILE_STATUS_OF` | Property | CodeElement (SELECT) | `FILE STATUS IS field` | 0.8 | `WK-FS` -> `SELECT:MASTER-FILE` |
| `ACCESSES` | Module | CodeElement | EXEC SQL/CICS | 0.9 | `BGTABFL` -> `sql-table:EMPLOYEES` |
| `RECEIVES` | Module | Property | `PROCEDURE USING` | 0.8 | `BGTABFL` -> `WK-INPUT-REC` |
| `CONTRACTS` | Module | Module | Shared copybook detection | 0.9 | `BGTABFL` -> `BGTABUP` (via `CPSESP`) |
## Full Annotated Example
Given this COBOL program:
```cobol
IDENTIFICATION DIVISION.
PROGRAM-ID. EMPMAINT.
AUTHOR. Development Team.
ENVIRONMENT DIVISION.
INPUT-OUTPUT SECTION.
FILE-CONTROL.
SELECT EMP-FILE
ASSIGN TO "EMPLOYEE.DAT"
ORGANIZATION IS INDEXED
ACCESS MODE IS DYNAMIC
RECORD KEY IS EMP-ID
FILE STATUS IS WS-FILE-STATUS.
DATA DIVISION.
FILE SECTION.
FD EMP-FILE.
01 EMP-RECORD.
05 EMP-ID PIC 9(6).
05 EMP-NAME PIC X(30).
WORKING-STORAGE SECTION.
01 WS-FLAGS.
05 WS-FILE-STATUS PIC X(02).
05 WS-EOF-FLAG PIC X(01).
88 WS-EOF VALUE "Y".
LINKAGE SECTION.
01 LK-SEARCH-KEY PIC 9(6).
PROCEDURE DIVISION USING LK-SEARCH-KEY.
MAIN-LOGIC SECTION.
MAIN-START.
PERFORM OPEN-FILE
PERFORM PROCESS-RECORDS
PERFORM CLOSE-FILE
STOP RUN.
OPEN-FILE.
OPEN I-O EMP-FILE.
PROCESS-RECORDS.
MOVE LK-SEARCH-KEY TO EMP-ID
EXEC SQL
SELECT EMP_SALARY INTO :WS-SALARY
FROM EMPLOYEES
WHERE EMP_ID = :EMP-ID
END-EXEC
CALL "EMPREPORT".
CLOSE-FILE.
CLOSE EMP-FILE.
```
The graph produced contains:
**Nodes:**
- `Module`: EMPMAINT (description: `author:Development Team`)
- `Namespace`: MAIN-LOGIC
- `Function`: MAIN-START, OPEN-FILE, PROCESS-RECORDS, CLOSE-FILE
- `Record`: EMP-RECORD, WS-FLAGS, LK-SEARCH-KEY
- `Property`: EMP-ID, EMP-NAME, WS-FILE-STATUS, WS-EOF-FLAG
- `Const`: WS-EOF (values: Y)
- `CodeElement`: SELECT:EMP-FILE, FD:EMP-FILE, sql-table:EMPLOYEES
- (COPY imports, if any, would produce File IMPORTS edges)
**Edges:**
- `DEFINES`: File -> all nodes
- `CONTAINS`: EMPMAINT -> EMP-RECORD, EMPMAINT -> WS-FLAGS, EMPMAINT -> LK-SEARCH-KEY
- `CONTAINS`: EMP-RECORD -> EMP-ID, EMP-RECORD -> EMP-NAME
- `CONTAINS`: WS-FLAGS -> WS-FILE-STATUS, WS-FLAGS -> WS-EOF-FLAG
- `CONTAINS`: WS-EOF-FLAG -> WS-EOF
- `CONTAINS`: FD:EMP-FILE -> EMP-RECORD
- `CONTAINS`: SELECT:EMP-FILE -> FD:EMP-FILE
- `CALLS`: MAIN-START -> OPEN-FILE, MAIN-START -> PROCESS-RECORDS, MAIN-START -> CLOSE-FILE
- `CALLS`: EMPMAINT -> EMPREPORT (external CALL)
- `ACCESSES`: EMPMAINT -> sql-table:EMPLOYEES
- `RECEIVES`: EMPMAINT -> LK-SEARCH-KEY (PROCEDURE USING)
- `RECORD_KEY_OF`: EMP-ID -> SELECT:EMP-FILE
- `FILE_STATUS_OF`: WS-FILE-STATUS -> SELECT:EMP-FILE
## How COBOL Differs from Tree-Sitter Languages
| Aspect | COBOL | Tree-Sitter Languages |
|--------|-------|----------------------|
| Node variety | 8 types (Module, Function, Namespace, Record, Property, Const, CodeElement, Constructor) | Typically 4-6 (Function, Class, Method, Interface, Module, Const) |
| Domain edges | RECORD_KEY_OF, FILE_STATUS_OF, ACCESSES, RECEIVES, CONTRACTS, REDEFINES | Primarily CALLS, IMPORTS, EXTENDS, IMPLEMENTS |
| Data hierarchy | Deep CONTAINS chains (01 -> 05 -> 10 -> 88) | Flat class members |
| Cross-program calls | CALL "name" + CICS LINK PROGRAM | Import-based resolution |
| Contract detection | Shared COPY copybook between caller/callee | Not applicable |
| Metadata | AUTHOR, DATE-WRITTEN on Module | JSDoc/docstring (not indexed) |
## Source Files
- `gitnexus/src/core/ingestion/workers/parse-worker.ts` -- `processCobolRegexOnly()`, node/edge emission logic
- `gitnexus/src/core/ingestion/pipeline.ts` -- `detectCrossProgamContracts()` for CONTRACTS edges
- `gitnexus/src/core/ingestion/cobol-preprocessor.ts` -- `CobolRegexResults` interface (all extracted data)

View file

@ -0,0 +1,261 @@
# COBOL Performance and Tuning
This document covers real-world benchmarks, worker pool configuration, memory management, known limitations, and troubleshooting for COBOL indexing.
## PROJECT-NAME Benchmark
The PROJECT-NAME project is a large Italian payroll system written in COBOL. It serves as the primary benchmark for COBOL indexing performance.
### Input
| Metric | Value |
| --------------------------- | ---------------------------------------------------------------------------- |
| Paths scanned | 14,217 |
| Parseable files | 13,129 |
| Total source size | 224 MB |
| Chunks | 12 (at 20 MB budget) |
| Copybooks loaded | 2,976 |
| Copybooks used in expansion | 2,955 |
| Key directories | `s/` (7773 programs), `c/` (3036 copybooks), `wfproc/` (1973 workflow files) |
### Output
| Metric | Value |
| ---------------------- | ------ |
| Graph nodes | 2.79M |
| Graph edges | 5.67M |
| Clusters (communities) | 16,679 |
| Execution flows | 300 |
### Timing
| Phase | Duration |
| ------------------------------- | ----------------- |
| Total | ~251s |
| KuzuDB write | 132s |
| Full-text search indexing | 6.7s |
| Regex extraction (avg per file) | ~1ms |
| COPY expansion + deep indexing | Remainder (~112s) |
### Indexing Command
```bash
cd /path/to/PROJECT-NAME
GITNEXUS_COBOL_DIRS=s,c,wfproc GITNEXUS_VERBOSE=1 node --max-old-space-size=8192 \
/path/to/gitnexus/dist/cli/index.js analyze --force
```
## Open-Source Benchmarks
### CardDemo (AWS)
| Metric | Value |
| ------ | ----- |
| Graph nodes | 12,323 |
| Graph edges | 8,893 |
| Total time | 7.4s |
### ACAS
| Metric | Value |
| ------ | ----- |
| Graph nodes | 14,016 |
| Graph edges | 15,452 |
| Total time | 9.3s |
### Micro-Benchmark (Single-File Extraction)
| Metric | Value |
| ------ | ----- |
| Per-iteration | 0.65ms |
| Throughput | ~382K lines/sec |
## Worker Pool Tuning
### Sub-Batch Size
The worker pool splits each worker's chunk into sub-batches to bound peak memory per `postMessage` serialization. COBOL repos use a smaller sub-batch size than the default:
| Parameter | Default | COBOL Mode |
| --------------------- | ----------- | ------------------- |
| Sub-batch size | 1,500 files | 200 files |
| Per sub-batch timeout | 120s | 120s (configurable) |
**Why 200?** COBOL regex extraction + preprocessing takes ~1ms per file on average, but with COPY expansion and deep indexing the effective time is ~150ms per file. At sub-batch size 1500, that would be ~225s per sub-batch, exceeding the 120s timeout.
COBOL mode is activated automatically when `GITNEXUS_COBOL_DIRS` is set:
```typescript
// From pipeline.ts
const cobolSubBatch = process.env.GITNEXUS_COBOL_DIRS ? 200 : undefined;
workerPool = createWorkerPool(workerUrl, undefined, cobolSubBatch);
```
### Worker Count
Workers default to `min(8, cpus - 1)`. For COBOL repos, this is usually sufficient since regex extraction is CPU-bound but fast. The bottleneck is typically KuzuDB write, not extraction.
### Timeout Configuration
| Environment Variable | Default | Purpose |
| ------------------------------------ | --------------- | --------------------------------------------------- |
| `GITNEXUS_WORKER_TIMEOUT_MS` | 120,000 (2 min) | Per sub-batch processing timeout |
| `GITNEXUS_WORKER_STARTUP_TIMEOUT_MS` | 60,000 (1 min) | Worker initialization timeout (tree-sitter loading) |
For COBOL-only repos, worker startup is faster because tree-sitter native modules are loaded lazily (skipped entirely if only COBOL files are present).
## Data Item Cap
### Configuration
```typescript
const MAX_DATA_ITEMS_PER_FILE = 500;
```
This constant appears in both `parse-worker.ts` (worker path) and `parsing-processor.ts` (sequential fallback).
### Rationale
Some COBOL programs, especially after COPY expansion, can have 10,000+ data items. At that scale:
- The in-memory relationship Map (for CONTAINS, REDEFINES, etc.) approaches the V8 16.7M entry limit across thousands of files
- KuzuDB write time increases linearly with edge count
- Most deep-nested items (level 20+) are rarely queried individually
### Impact
The cap truncates data items beyond the 500th in source order. Since 01-level Records appear first in COBOL source, the cap preserves:
- All 01-level record definitions
- The most important 02-49 level items (those closest to the record root)
- 88-level conditions associated with early items
To increase the cap for specific needs, modify the `MAX_DATA_ITEMS_PER_FILE` constant in both files.
## Memory Management
### COPY Expansion Breadth Guard
A per-file `MAX_TOTAL_EXPANSIONS = 500` limit prevents exponential blowup from diamond-shaped COPY graphs (e.g., N copybooks each containing N COPY statements). Once the limit is reached, further COPY statements in that file are left unexpanded. See [copy-expansion.md](copy-expansion.md) for details.
### COPY Expansion Memory
All copybook content is loaded upfront into a Map before chunk processing begins. For PROJECT-NAME:
- 2,976 copybooks, typically under 100MB total
- The Map is shared (read-only) across chunk iterations
- Per-chunk, the copybook map is merged with chunk file content (in case a chunk contains copybooks not in the pre-loaded set)
- After all chunks are processed, the copybook map is freed (`cobolCopybookContents = undefined`)
### Chunk Budget
Source files are grouped into chunks of max 20MB (`CHUNK_BYTE_BUDGET`). Each chunk's lifecycle:
1. Read file content into memory
2. Expand COPY statements (mutates content in-place)
3. Dispatch to workers for extraction
4. Workers return serialized results
5. Merge results into graph
6. Chunk content goes out of scope (GC reclaims)
This ensures only ~20MB of source + ~200-400MB of working memory (ASTs, extracted records, serialization) is active at any time.
### Shared Warning Deduplication
The `warnedCircular` set (used by the COPY expansion engine) is shared across all files in a chunk. This prevents the same circular copybook warning (e.g., `ANAZI includes itself`) from being logged thousands of times.
## Known Limitations
| Limitation | Impact | Workaround |
| ---------------------------------------- | --------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------- |
| tree-sitter-cobol hangs on ~5% of files | Cannot use tree-sitter for COBOL | Regex-only extraction (current approach) |
| Data item cap (500/file) | May miss deeply nested items in large programs | Increase `MAX_DATA_ITEMS_PER_FILE` in source |
| Circular copybooks (ANAZI, ANDIP, QDIPE) | Self-referential includes cannot be expanded | Detected and skipped with warning |
| wfproc/ files may not be pure COBOL | Workflow files may produce extraction noise | Exclude `wfproc` from `GITNEXUS_COBOL_DIRS` if problematic |
| No MOVE DATA_FLOW edges yet | Data flow between variables not in graph | Reserved for future release |
| Continuation line handling | Some complex multi-line continuations (especially in string literals spanning 3+ lines) may not merge correctly | Known edge case; affects <0.1% of lines |
| Single-line EXEC blocks | `EXEC SQL SELECT ... END-EXEC` on one line is handled, but pathological nesting is not | Extremely rare in practice |
| Extension case sensitivity | `.GNM` and `.gnm` are matched differently | Use the exact case from the codebase |
## Troubleshooting
### "COPY expansion failed"
```
[pipeline] COPY expansion failed for s/BGTABFL: Cannot read properties of null
```
**Cause:** A copybook referenced by a COPY statement cannot be found.
**Fix:**
1. Verify `GITNEXUS_COBOL_DIRS` includes the directory containing copybooks (typically `c`)
2. Check that copybook filenames match the COPY target (case-insensitive, after stripping extensions)
3. Ensure copybook files are not in `.gitignore`
### Worker sub-batch timeout
```
Worker 3 sub-batch timed out after 120s (chunk: 200 items)
```
**Cause:** A sub-batch took longer than the timeout. Typically happens when one file is extremely large (50,000+ lines after COPY expansion).
**Fix:** Increase the timeout:
```bash
GITNEXUS_WORKER_TIMEOUT_MS=300000 gitnexus analyze
```
### Memory errors (heap out of memory)
```
FATAL ERROR: CALL_AND_RETRY_LAST Allocation failed - JavaScript heap out of memory
```
**Fix:** Increase Node.js heap size:
```bash
node --max-old-space-size=16384 /path/to/gitnexus/dist/cli/index.js analyze
```
For very large repos (>500MB source), consider `--max-old-space-size=32768`.
### Concurrent analyze corruption
**Rule:** Only ONE `gitnexus analyze` process should run at a time per repository. Concurrent writes to KuzuDB corrupt the database.
If corruption occurs:
```bash
# Remove the KuzuDB directory and re-index
rm -rf .gitnexus/kuzu
gitnexus analyze --force
```
### Slow KuzuDB write phase
The KuzuDB write phase (132s for PROJECT-NAME) is the bottleneck for large COBOL repos. This is proportional to the number of nodes and edges being written. Reducing `MAX_DATA_ITEMS_PER_FILE` or excluding non-essential directories from `GITNEXUS_COBOL_DIRS` can help.
### Verbose output
Enable verbose logging to see per-phase timing and statistics:
```bash
GITNEXUS_VERBOSE=1 gitnexus analyze
```
This outputs:
- Scan statistics (paths, parseable files, chunk count)
- Worker pool configuration (worker count, sub-batch size)
- COPY expansion statistics (copybooks loaded, files expanded)
- Community and process detection results
- Contract detection results
## Source Files
- `gitnexus/src/core/ingestion/workers/worker-pool.ts` -- `DEFAULT_SUB_BATCH_SIZE`, `SUB_BATCH_TIMEOUT_MS`, `WORKER_STARTUP_TIMEOUT_MS`
- `gitnexus/src/core/ingestion/pipeline.ts` -- `CHUNK_BYTE_BUDGET`, COBOL sub-batch configuration, chunk lifecycle
- `gitnexus/src/core/ingestion/workers/parse-worker.ts` -- `MAX_DATA_ITEMS_PER_FILE`, `processCobolRegexOnly()`
- `gitnexus/src/core/ingestion/parsing-processor.ts` -- Sequential fallback `MAX_DATA_ITEMS_PER_FILE`

View file

@ -0,0 +1,206 @@
# COBOL Regex Extraction
The `extractCobolSymbolsWithRegex()` function in `cobol-preprocessor.ts` performs single-pass, state-machine-driven extraction of all COBOL symbols. This document describes the state machine, line processing flow, and every regex pattern used.
## State Machine: Division Tracking
The extractor tracks which COBOL division is currently being processed. Division transitions are detected by the `RE_DIVISION` pattern.
```mermaid
stateDiagram-v2
[*] --> null : Start of file
null --> identification : IDENTIFICATION DIVISION
identification --> environment : ENVIRONMENT DIVISION
environment --> data : DATA DIVISION
data --> procedure : PROCEDURE DIVISION
note right of identification
Extracts: PROGRAM-ID, AUTHOR, DATE-WRITTEN
end note
note right of environment
Extracts: SELECT ... ASSIGN ... (file declarations)
end note
note right of data
Extracts: FD entries, data items (01-77, 88), COPY
end note
note right of procedure
Extracts: paragraphs, sections, PERFORM, CALL,
ENTRY, MOVE, EXEC SQL/CICS
end note
```
## State Machine: Data Section Tracking
Within the DATA DIVISION, a secondary state machine tracks the current section to tag data items with their origin.
```mermaid
stateDiagram-v2
[*] --> unknown : DATA DIVISION entered
unknown --> working_storage : WORKING-STORAGE SECTION
unknown --> linkage : LINKAGE SECTION
unknown --> file : FILE SECTION
unknown --> local_storage : LOCAL-STORAGE SECTION
working_storage --> linkage : LINKAGE SECTION
working_storage --> file : FILE SECTION
linkage --> working_storage : WORKING-STORAGE SECTION
file --> working_storage : WORKING-STORAGE SECTION
file --> linkage : LINKAGE SECTION
local_storage --> working_storage : WORKING-STORAGE SECTION
```
Within the ENVIRONMENT DIVISION, the `currentEnvSection` tracks whether we are in `INPUT-OUTPUT` or `CONFIGURATION` section. SELECT statement accumulation only occurs in `INPUT-OUTPUT`.
## Line Processing Flow
Each raw source line goes through this pipeline:
```
Raw line
|
v
Length < 7? ---------> Skip (flush pending if any)
|
v
Indicator col 7
|
+-- '*' or '/' -----> Comment: skip entirely
|
+-- '-' ------------> Continuation: append to pending line
|
+-- other ----------> Normal: flush pending, strip inline comments (|),
buffer as new pending logical line
```
After all lines are processed, the final pending line is flushed, along with any accumulated SELECT statement, SORT/MERGE accumulator, and any open EXEC block (truncated file without `END-EXEC`).
### Inline Comment Stripping
Enterprise COBOL (particularly Italian dialect) uses the pipe character `|` as an inline comment marker. The `stripInlineComment()` helper is **quote-aware**: it tracks whether the scan position is inside a single- or double-quoted string and only treats `|` as a comment marker when outside quotes. Pipe characters inside string literals are preserved.
Free-format `*>` inline comment stripping uses the same quote-aware approach: the scanner walks character by character, toggling quote state, and only recognizes `*>` as a comment marker when not inside a quoted string.
### Patch Marker Handling
The `preprocessCobolSource()` function (run before extraction in the worker) replaces non-standard content in columns 1-6. Standard COBOL expects spaces or digit sequence numbers in this area. If any letter or `#` character is found, the entire sequence area is replaced with 6 spaces:
```
Before: mzADD MOVE WK-AMT TO WK-TOTAL
After: MOVE WK-AMT TO WK-TOTAL
```
This preserves exact line count for position mapping.
## Regex Pattern Reference
All patterns are compiled once as module-level constants and reused across calls.
### Division and Section Detection
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_DIVISION` | `\b(IDENTIFICATION\|ENVIRONMENT\|DATA\|PROCEDURE)\s+DIVISION\b` | Division boundary | `PROCEDURE DIVISION` |
| `RE_SECTION` | `\b(WORKING-STORAGE\|LINKAGE\|FILE\|LOCAL-STORAGE\|INPUT-OUTPUT\|CONFIGURATION)\s+SECTION\b` | Section boundary | `WORKING-STORAGE SECTION` |
### IDENTIFICATION DIVISION
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_PROGRAM_ID` | `\bPROGRAM-ID\.\s*([A-Z][A-Z0-9-]*)` | Program name | `PROGRAM-ID. BGTABFL` |
| `RE_AUTHOR` | `^\s+AUTHOR\.\s*(.+)` | Author metadata | `AUTHOR. D. Smith` |
| `RE_DATE_WRITTEN` | `^\s+DATE-WRITTEN\.\s*(.+)` | Date metadata | `DATE-WRITTEN. 2024-01-15` |
### ENVIRONMENT DIVISION
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_SELECT_START` | `\bSELECT\s+(?:OPTIONAL\s+)?([A-Z][A-Z0-9-]+)` | File SELECT start (with optional `SELECT OPTIONAL` support) | `SELECT MASTER-FILE`, `SELECT OPTIONAL TRANS-FILE` |
SELECT statements are accumulated across multiple lines until a period terminator is found, then parsed for ASSIGN, ORGANIZATION, ACCESS, RECORD KEY, and FILE STATUS clauses.
### DATA DIVISION
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_FD` | `^\s+FD\s+([A-Z][A-Z0-9-]+)` | File description | `FD MASTER-FILE` |
| `RE_DATA_ITEM` | `^\s+(\d{1,2})\s+([A-Z][A-Z0-9-]+)\s*(.*)` | Data item (01-77) | `05 WK-NAME PIC X(30)` |
| `RE_ANONYMOUS_REDEFINES` | `^\s+(\d{1,2})\s+REDEFINES\s+([A-Z][A-Z0-9-]+)` | Anonymous REDEFINES | `01 REDEFINES WK-REC` |
| `RE_88_LEVEL` | `^\s+88\s+([A-Z][A-Z0-9-]+)\s+VALUES?\s+(?:ARE\s+)?(.+)` | Condition name | `88 WK-ACTIVE VALUE "Y"` |
The trailing clauses of `RE_DATA_ITEM` are parsed by `parseDataItemClauses()` for PIC, USAGE, OCCURS, and REDEFINES.
### PROCEDURE DIVISION
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_PROC_SECTION` | `^ ([A-Z][A-Z0-9-]+)\s+SECTION\.\s*$` | Procedure section header | ` MAIN-LOGIC SECTION.` |
| `RE_PROC_PARAGRAPH` | `^ ([A-Z][A-Z0-9-]+)\.\s*$` | Paragraph header | ` PROCESS-RECORD.` |
| `RE_PERFORM` | `\bPERFORM\s+([A-Z][A-Z0-9-]+)(?:\s+THRU\s+([A-Z][A-Z0-9-]+))?` | PERFORM call | `PERFORM CALC-TAX THRU CALC-TAX-EXIT` |
| `RE_PROC_USING` | `\bPROCEDURE\s+DIVISION\s+USING\s+([\s\S]*?)(?:\.\|$)` | USING parameters | `PROCEDURE DIVISION USING WK-PARAM` |
| `RE_ENTRY` | `\bENTRY\s+"([^"]+)"(?:\s+USING\s+([\s\S]*?))?(?:\.\|$)` | ENTRY point | `ENTRY "SUBPROG" USING WK-DATA` |
| `RE_MOVE` | `\bMOVE\s+((?:CORRESPONDING\|CORR)\s+)?([A-Z][A-Z0-9-]+)\s+TO\s+(.+)` | MOVE statement (supports CORR abbreviation and multi-target) | `MOVE WK-NAME TO OUT-NAME`, `MOVE CORR WK-IN TO WK-OUT` |
The USING parameter list (`RE_PROC_USING`) is split on `\bRETURNING\b` before tokenization -- any RETURNING clause and everything after it is excluded from the parameter list (`.split(/\bRETURNING\b/i)[0]`).
Note: `RE_PROC_SECTION` and `RE_PROC_PARAGRAPH` require exactly 7 spaces of leading indentation (COBOL area A starting at column 8). This is the standard COBOL paragraph indentation.
### All-Division Patterns
These patterns are checked regardless of current division:
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_CALL` | `\bCALL\s+"([^"]+)"` | External program call | `CALL "BGTABUP"` |
| `RE_COPY_UNQUOTED` | `\bCOPY\s+([A-Z][A-Z0-9-]+)(?:\s\|\.)` | COPY (unquoted) | `COPY CPSESP.` |
| `RE_COPY_QUOTED` | `\bCOPY\s+"([^"]+)"(?:\s\|\.)` | COPY (quoted) | `COPY "WORKGRID.CPY".` |
### SORT/MERGE Support
| Constant | Purpose |
|----------|---------|
| `SORT_CLAUSE_NOISE` | Set of SORT/MERGE clause keywords filtered from USING/GIVING file lists: `ON`, `ASCENDING`, `DESCENDING`, `KEY`, `WITH`, `DUPLICATES`, `IN`, `ORDER`, `COLLATING`, `SEQUENCE`, `IS`, `THROUGH`, `THRU`, `INPUT`, `OUTPUT`, `PROCEDURE` |
SORT and MERGE statements are accumulated across multiple lines (like SELECT) until a period terminator is found, then parsed for USING/GIVING file lists and INPUT/OUTPUT PROCEDURE targets. The `flushSort()` helper encapsulates the flush-and-parse logic, mirroring the existing `flushSelect()` pattern. Both helpers are called at EOF to handle truncated files.
### GO TO Multi-Target
`RE_GOTO` captures all paragraph names in a `GO TO` statement, including the multi-target form `GO TO p1 p2 p3 DEPENDING ON x`. The captured group contains all target names (space-separated), which are split into individual targets. Each target produces a separate `gotos` entry.
### PROGRAM-ID Detection
PROGRAM-ID is detected regardless of the current division state. This handles sibling programs that appear after `END PROGRAM` and omit the `IDENTIFICATION DIVISION` header -- the extractor will still capture the PROGRAM-ID and push a new program boundary.
### EXEC Block Patterns
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_EXEC_SQL_START` | `\bEXEC\s+SQL\b` | Start of EXEC SQL block | `EXEC SQL` |
| `RE_EXEC_CICS_START` | `\bEXEC\s+CICS\b` | Start of EXEC CICS block | `EXEC CICS` |
| `RE_END_EXEC` | `\bEND-EXEC\b` | End of EXEC block | `END-EXEC` |
EXEC blocks accumulate all lines between `EXEC SQL/CICS` and `END-EXEC`, then delegate to `parseExecSqlBlock()` or `parseExecCicsBlock()` for detailed extraction.
## Excluded Paragraph Names
The following names are excluded from paragraph detection to avoid false positives from division/section headers:
```
DECLARATIVES, END, PROCEDURE, IDENTIFICATION,
ENVIRONMENT, DATA, WORKING-STORAGE, LINKAGE,
FILE, LOCAL-STORAGE, COMMUNICATION, REPORT,
SCREEN, INPUT-OUTPUT, CONFIGURATION
```
Additionally, paragraph candidates containing `DIVISION` or `SECTION` as substrings are excluded.
## MOVE Skip List (Figurative Constants)
MOVE statements where the source is a figurative constant are skipped:
```
SPACES, ZEROS, ZEROES, LOW-VALUES, LOW-VALUE,
HIGH-VALUES, HIGH-VALUE, QUOTES, QUOTE, ALL
```
## Source Files
- `gitnexus/src/core/ingestion/cobol-preprocessor.ts` -- `preprocessCobolSource()`, `extractCobolSymbolsWithRegex()`, all regex constants

View file

@ -0,0 +1,326 @@
---
title: "feat: Complete COBOL language feature coverage for maximum knowledge graph value"
type: feat
status: active
date: 2026-03-26
origin: Feature audit from v3-integration-architect agent (session 8642401e)
---
## Enhancement Summary
**Deepened on:** 2026-03-26
**Research agents used:** COBOL expert (Phase 1+2), graph value analyst, codebase explorer
**Sections enhanced:** Phase 1 (5 features), Phase 2 (4 features), graph value ranking
### Key Improvements from Research
1. **CALL USING** is the #1 highest-value edge type (9.2/10) — fixes ~40% of missing caller references
2. **EXEC DLI** requires dual-interface support (EXEC DLI + CBLTDLI CALL) for full IMS coverage
3. **DECLARATIVES** is lowest-risk Phase 2 item — existing section/paragraph detection already captures structure
4. **SET TO TRUE** accounts for 80-90% of all SET statements — prioritize this form
5. **INSPECT** needs multi-line accumulator (like SORT) — can span 5+ continuation lines
6. **Graph value ranking**: cobol-call-using (9.2) > cobol-error-handler (9.0) > dli-gu (8.2) > cobol-string (6.2)
### New Edge Cases Discovered
- CALL USING supports mixed modes: `USING BY REFERENCE WS-A BY CONTENT WS-B BY VALUE WS-C`
- CALL USING `ADDRESS OF` and `OMITTED` must be filtered from parameter lists
- EXEC DLI can have multiple SEGMENT levels in hierarchical retrieval (use matchAll)
- DECLARATIVES can have multiple USE sections (one per file + catch-all for INPUT/OUTPUT/I-O/EXTEND)
- INSPECT TALLYING can have multiple counters in a single statement
- STRING/UNSTRING can span multiple lines (need accumulator pattern)
---
# Complete COBOL Language Feature Coverage
## Overview
Implement the remaining 25 unhandled COBOL language features and fix 10 partial features to achieve ~95% coverage (up from 71.9%). The goal is to build the richest possible knowledge graph from COBOL codebases, enabling a future `modernize` MCP command (out of scope for this plan) that would use the graph to assist with COBOL-to-modern-language migration.
## Problem Statement
The COBOL processor currently handles 54 of 89 applicable language features (71.9%). The 25 unhandled features represent real data loss in the knowledge graph:
- **Cross-program data flow** is invisible (CALL ... USING parameters not extracted)
- **IMS/DB programs** produce empty graphs (EXEC DLI not recognized)
- **String transformation logic** is invisible (STRING/UNSTRING/INSPECT not tracked)
- **SQL copybook dependencies** are missing (EXEC SQL INCLUDE not mapped)
- **Error handling flows** are lost (DECLARATIVES/USE AFTER not captured)
## Proposed Solution
Implement features in 4 phases, ordered by graph value density (edges created per LOC of implementation). Each phase is independently shippable and testable.
## Technical Approach
### Phase 1: High-Value Data Flow Edges (~150 LOC, ~8 new edge types)
The highest-ROI features: they create new ACCESSES and IMPORTS edges that directly improve impact analysis.
**Critical research finding**: Multi-line statement accumulation is the dominant challenge. CALL USING, STRING/UNSTRING, and multi-line data item clauses all span multiple lines in production COBOL. The free-format path processes each line independently — these features need statement accumulators (like SORT/SELECT) or the free-format path needs multi-line awareness. Estimated LOC increased from 110 to 150 to account for accumulator infrastructure.
#### 1.1 EXEC SQL INCLUDE -> IMPORTS edges
- **File:** `cobol-preprocessor.ts` (parseExecSqlBlock)
- **What:** Detect `INCLUDE` as the operation, extract member name, emit as a `copies[]` entry
- **Graph:** IMPORTS edge from File to included copybook/SQLCA with reason `sql-include`
- **Tests:** Unit test for `EXEC SQL INCLUDE SQLCA END-EXEC` and `EXEC SQL INCLUDE CUSTCOPY END-EXEC`
**Research insights (EXEC SQL INCLUDE):**
- DB2 member names can contain underscores: `EXEC SQL INCLUDE CUST_TBL_DCL END-EXEC` — regex must use `[A-Z][A-Z0-9_-]+`
- Quoted literal form: `EXEC SQL INCLUDE 'DBRMLIB.MEMBER' END-EXEC` (z/OS PDS qualified name)
- SQLCA/SQLDA are DB2 builtins — won't resolve to repo files. Emit unresolved IMPORTS edge (still valuable)
- No REPLACING support on EXEC SQL INCLUDE (unlike COPY)
- Add `INCLUDE` to `OP_MAP` in `parseExecSqlBlock`; extract member via `RE_SQL_INCLUDE = /^INCLUDE\s+(?:'([^']+)'|"([^"]+)"|([A-Z][A-Z0-9_-]+))/i`
#### 1.2 CALL ... USING parameter extraction -> ACCESSES edges (Graph value: 9.2/10)
- **File:** `cobol-preprocessor.ts` (processLogicalLine CALL section)
- **What:** After capturing CALL target, scan for USING clause. Extract parameter names (reuse USING_KEYWORDS filter). Store as `calls[].parameters: string[]`
- **Interface:** Add `parameters?: string[]` to calls array type in CobolRegexResults
- **File:** `cobol-processor.ts` (CALL edge block)
- **Graph:** For each USING parameter, create ACCESSES edge from caller to data item Property node with reason `cobol-call-using`
- **Tests:** `CALL 'AUDITLOG' USING CUST-ID WS-AMOUNT` -> 2 ACCESSES edges
**Research insights (CALL USING forms):**
- Mixed modes: `CALL 'PGM' USING BY REFERENCE WS-A BY CONTENT WS-B BY VALUE WS-C`
- Pointer passing: `CALL 'PGM' USING ADDRESS OF WS-A`
- Placeholder: `CALL 'PGM' USING OMITTED WS-B`
- Filter keywords: add `ADDRESS`, `OMITTED`, `LENGTH` to USING_KEYWORDS (already has BY/VALUE/REFERENCE/CONTENT)
- **Impact tool enhancement:** CALL-USING edges enable BFS traversal through parameter data flow — single most impactful edge type for COBOL impact analysis
#### 1.3 STRING/UNSTRING data flow -> ACCESSES edges
- **File:** `cobol-preprocessor.ts` (new section in extractProcedure)
- **What:** Accumulate multi-line STRING/UNSTRING until period or END-STRING/END-UNSTRING. Extract sources and INTO targets.
- **Interface:** Add `strings: Array<{ sources: string[]; target: string; type: 'string' | 'unstring'; line: number; caller: string | null }>` to CobolRegexResults
- **Graph:** read-ACCESSES on sources, write-ACCESSES on INTO target with reason `cobol-string-read` / `cobol-string-write`
- **Tests:** 2 unit tests + integration test assertions
**Research insights (STRING/UNSTRING):**
- **Needs statement accumulator** — STRING/UNSTRING always span multiple lines in production
- Terminate accumulation at: period, END-STRING/END-UNSTRING, or start of next COBOL verb
- STRING sources: identifiers before each `DELIMITED BY`. Filter: STRING, DELIMITED, BY, SIZE, ALL, INTO, WITH, POINTER, ON, OVERFLOW, NOT, END-STRING
- UNSTRING: source is first identifier after UNSTRING; INTO targets are identifiers after INTO. Filter: DELIMITER, IN, COUNT, TALLYING, OR
- WITH POINTER field is both read AND written (starting position updated)
- TALLYING IN / COUNT IN fields are write targets
- Literal sources (`'text'`) must be filtered — quote-aware tokenization needed
- **Edge case**: STRING terminated by next verb, not period — existing fixture has `STRING ... DISPLAY` without period between them
#### 1.4 OCCURS DEPENDING ON -> ACCESSES edge
- **File:** `cobol-preprocessor.ts` (parseDataItemClauses)
- **What:** Extend OCCURS regex to capture DEPENDING ON field, KEY fields, and INDEXED BY names
- **Interface:** Add `dependingOn?: string`, `occursMax?: number`, `occursKeys?: Array<{direction: string; fields: string[]}>`, `indexedBy?: string[]` to data items
- **Graph:** ACCESSES edge from table item to controlling field with reason `cobol-depends-on`
- **Tests:** `05 WS-TABLE OCCURS 100 DEPENDING ON WS-COUNT` -> edge
**Research insights (OCCURS):**
- IBM allows `OCCURS 0 TO n DEPENDING ON` (zero minimum) and `OCCURS UNBOUNDED DEPENDING ON` (V6.4)
- Subscripted controlling fields: `DEPENDING ON WS-COUNT(WS-IDX)` — strip subscripts before storing
- **Pre-existing gap**: Multi-line data item clauses without continuation indicator are NOT captured. `05 WS-TABLE\n OCCURS 100\n DEPENDING ON WS-COUNT.` — the current RE_DATA_ITEM only gets the first line, `rest` is empty. Fixing properly requires a data item accumulator (like SELECT). **Defer full fix to Phase 3; implement same-line capture now.**
- KEY IS fields: `ASCENDING KEY IS WS-KEY-1 WS-KEY-2` — capture for SEARCH ALL resolution
- INDEXED BY: `INDEXED BY IDX-1 IDX-2` — capture for SET/SEARCH context
#### 1.5 VALUE clause for standard data items
- **File:** `cobol-preprocessor.ts` (parseDataItemClauses)
- **What:** Extract VALUE using a pragmatic function that handles quoted strings, numerics, figurative constants, hex/national literals
- **Interface:** Already exists as `values?: string[]` on data items (currently only populated for 88-level)
- **Graph:** Stored in Property node description (no new edges)
- **Tests:** `01 WS-STATUS PIC X VALUE 'A'` -> values: ['A']
**Research insights (VALUE forms):**
- Hex literals: `VALUE X'F1F2F3F4'`, National: `VALUE N'text'`, DBCS: `VALUE G'text'`
- Figurative constants: SPACES, ZEROS, ZEROES, LOW-VALUES, HIGH-VALUES, QUOTES, NULL, NULLS
- ALL literal: `VALUE ALL '*'`
- Numeric with sign/decimal: `VALUE -123.45`, `VALUE +1`
- `VALUE IS` optional — both `VALUE 'A'` and `VALUE IS 'A'` valid
- **Decimal vs period ambiguity**: `VALUE 100.` — is `.` decimal or terminator? `parseDataItemClauses` already strips trailing period, so this is handled
- IBM V6.4: floating-point `VALUE 1.0E5` — extend numeric regex if needed
- Implementation: use a pragmatic `extractValue(rest)` function, not a single complex regex
### Phase 2: EXEC DLI + DECLARATIVES (~90 LOC, ~4 new edge types)
IMS/DB support and error handling flows.
#### 2.1 EXEC DLI (IMS/DB) -> ACCESSES edges (Graph value: 8.2/10)
- **File:** `cobol-preprocessor.ts` (processLogicalLine — add RE_EXEC_DLI_START check alongside SQL/CICS)
- **What:** Accumulate EXEC DLI blocks like EXEC SQL. Parse DLI verbs (GU, GN, GNP, GHU, GHN, GHNP, ISRT, DLET, REPL, CHKP, SCHD, TERM). Extract segment name, PCB number, INTO/FROM areas, WHERE fields, PSB name.
- **Interface:** Add `execDliBlocks: Array<{ line: number; verb: string; pcbNumber?: number; segmentName?: string; intoField?: string; fromField?: string; whereField?: string; psbName?: string }>` to CobolRegexResults
- **Graph:** CodeElement node + ACCESSES edge to `<ims>:<segmentName>` Record node with reason `dli-{verb}`; ACCESSES edges to INTO/FROM data areas; PSB ACCESSES for SCHD
- **Tests:** `EXEC DLI GU USING PCB(1) SEGMENT(CUSTOMER) INTO(WS-CUST) END-EXEC`
**Research insights (dual IMS interface):**
- **EXEC DLI**: Embedded command interface for CICS-DL/I programs only
- **CBLTDLI CALL**: Batch interface via `CALL 'CBLTDLI' USING function-code PCB io-area SSA1..SSA15`
- CBLTDLI is already captured as a CALL to 'CBLTDLI' — enrich with USING parameter semantics later
- Multiple SEGMENT levels in hierarchical retrieval — use `matchAll` on segment regex
- DLI verbs: GU (most common), GN, GNP, GHU, GHN, GHNP, ISRT, REPL, DLET, CHKP, SCHD, TERM, ROLL, ROLB
- **Edge case**: DLET/REPL have no SEGMENT clause (operate on current position)
- **Recommended order**: Implement AFTER DECLARATIVES and SET (lower risk, higher frequency)
#### 2.2 DECLARATIVES / USE AFTER STANDARD EXCEPTION (Graph value: 9.0/10)
- **File:** `cobol-preprocessor.ts` (processLogicalLine — detect DECLARATIVES keyword, track USE AFTER blocks)
- **What:** When `DECLARATIVES.` is encountered, switch to declaratives mode. Extract USE statements binding sections to files/modes.
- **Interface:** Add `declaratives: Array<{ sectionName: string; useType: 'error' | 'debug' | 'label' | 'reporting'; target: string; line: number }>` to CobolRegexResults
- **Graph:** ACCESSES edge from declarative Namespace to file Record with reason `cobol-declarative-error-handler`
- **Tests:** Unit test with DECLARATIVES section, integration test for error flow
**Research insights (DECLARATIVES syntax):**
- `USE AFTER STANDARD {EXCEPTION|ERROR} ON {file-name|INPUT|OUTPUT|I-O|EXTEND}`
- EXCEPTION and ERROR are synonymous; STANDARD is optional in IBM dialects
- Multiple USE sections allowed (one per file + catch-all for I/O modes)
- `END DECLARATIVES.` must NOT reset PROCEDURE DIVISION state
- `DECLARATIVES` is already in EXCLUDED_PARA_NAMES — no false paragraph risk
- Existing section/paragraph detection already captures structural elements — just need USE binding
- **Lowest risk Phase 2 item** — implement first
#### 2.3 SET statement -> ACCESSES edges
- **File:** `cobol-preprocessor.ts` (extractProcedure — new RE_SET regex)
- **Interface:** Add `sets: Array<{ targets: string[]; form: 'to-true'|'to-value'|'up-by'|'down-by'|'address-of'|'to-null'|'to-entry'; value?: string; entryTarget?: string; entryIsLiteral?: boolean; line: number; caller: string | null }>` to CobolRegexResults
- **Graph:** ACCESSES write edge with reason `cobol-set-condition` (TO TRUE), `cobol-set-index` (TO/UP/DOWN), `cobol-set-address` (ADDRESS OF). SET ENTRY with literal -> CALLS edge.
- **Tests:** `SET WS-EOF TO TRUE`, `SET IDX-1 TO 5`, `SET IDX-1 UP BY 1`
**Research insights (SET forms by frequency):**
- `SET condition TO TRUE` — 80-90% of all SET usage. Multiple targets: `SET COND-A COND-B TO TRUE`
- `SET index TO/UP BY/DOWN BY` — ~8%. Multiple indices: `SET IDX-1 IDX-2 UP BY 1`
- `SET pointer TO ADDRESS OF data-item` / `SET ADDRESS OF data-item TO pointer` — ~2%
- `SET proc-ptr TO ENTRY "PROGNAME"` — rare but creates CALLS edge (like dynamic CALL)
- Filter OF/IN qualifiers: `SET COND-A OF WS-RECORD TO TRUE` (strip OF WS-RECORD)
- **Prioritize**: SET TO TRUE alone covers 80-90% — implement this form first
#### 2.4 INSPECT -> ACCESSES edges
- **File:** `cobol-preprocessor.ts` (extractProcedure — new `inspectAccum` accumulator like SORT)
- **What:** Accumulate multi-line INSPECT until period. Extract inspected field + tally counters.
- **Interface:** Add `inspects: Array<{ inspectedField: string; counters: string[]; form: 'tallying'|'replacing'|'converting'|'tallying-replacing'; line: number; caller: string | null }>` to CobolRegexResults
- **Graph:** ACCESSES read on inspected field always; write if REPLACING/CONVERTING. Write edges for tally counters. Reason: `cobol-inspect-read`/`cobol-inspect-write`/`cobol-inspect-tally`
- **Tests:** `INSPECT WS-FIELD TALLYING WS-COUNT FOR ALL 'A'` -> read on WS-FIELD, write on WS-COUNT
**Research insights (INSPECT forms by frequency):**
- REPLACING (~60%): `INSPECT WS-STR REPLACING ALL 'A' BY 'B'`
- TALLYING (~25%): `INSPECT WS-STR TALLYING WS-CNT FOR ALL 'A'` — multiple counters possible
- CONVERTING (~10%): `INSPECT WS-STR CONVERTING 'abc' TO 'ABC'`
- Combined (~5%): TALLYING + REPLACING in single statement
- **Needs multi-line accumulator** — INSPECT frequently spans 3-5 lines in production
- Extract tally counters with `([A-Z][A-Z0-9-]+)\s+FOR\b` matchAll pattern
- Filter figurative constants (SPACES, ZEROS) using existing MOVE_SKIP set
### Phase 3: Completeness Fixes (~60 LOC)
Fix the 10 partial features and small gaps.
#### 3.1 CALL ... RETURNING extraction
- Extend RE_CALL processing to capture RETURNING target after the USING clause
- Store as `calls[].returning?: string`
- Graph: ACCESSES write edge with reason `cobol-call-returning`
#### 3.2 SELECT OPTIONAL flag preservation
- Store `isOptional: boolean` in FileDeclaration interface
- Include in Record node description
#### 3.3 ALTERNATE RECORD KEY extraction
- Add regex in parseSelectStatement: `/\bALTERNATE\s+RECORD\s+KEY\s+(?:IS\s+)?([A-Z][A-Z0-9-]+)/i`
- Store as `alternateKeys?: string[]`
#### 3.4 COMMON attribute on nested programs
- Extend RE_PROGRAM_ID: `/\bPROGRAM-ID\.\s*([A-Z][A-Z0-9-]+)(?:\s+IS\s+COMMON)?/i`
- Store `isCommon: boolean` on Module node
- Affects cross-program CALL resolution scope
#### 3.5 IS EXTERNAL / IS GLOBAL as first-class properties
- Change from usage string hack to proper boolean fields on data items
- Add `isExternal?: boolean`, `isGlobal?: boolean` to data item interface
#### 3.6 AUTHOR / DATE-WRITTEN mapped to Module node
- Already extracted as programMetadata — map to Module node properties
- `graph.addNode({ ..., properties: { ..., author, dateWritten } })`
#### 3.7 REPLACE statement
- Track REPLACE / REPLACE OFF state in preprocessor
- Apply text substitutions during preprocessing (before regex extraction)
- Complex: requires careful scoping rules
### Phase 4: Niche Features (~30 LOC)
Low-priority but nice for completeness.
#### 4.1 INITIALIZE statement -> write ACCESSES
- `/\bINITIALIZE\s+([A-Z][A-Z0-9-]+)/i`
- ACCESSES write edge with reason `cobol-initialize`
#### 4.2 Remaining IDENTIFICATION DIVISION paragraphs
- DATE-COMPILED, INSTALLATION, SECURITY, REMARKS
- Map to Module node description properties
#### 4.3 EXEC SQL INCLUDE -> IMPORTS edge (expansion)
- For EXEC SQL INCLUDE inside EXEC blocks that reference copybooks containing SQL
- Create IMPORTS edge similar to COPY
## Acceptance Criteria
### Functional Requirements
- [ ] Phase 1: All 5 features implemented with unit + integration tests
- [ ] Phase 2: All 4 features implemented with unit + integration tests
- [ ] Phase 3: All 7 partial features fixed
- [ ] Phase 4: At least 2 of 3 niche features implemented
- [ ] All existing 145 tests continue to pass
- [ ] TypeScript compiles cleanly
### Non-Functional Requirements
- [ ] No performance regression: CardDemo benchmark stays under 8s
- [ ] No file exceeds 1500 LOC (preprocessor currently 1326)
- [ ] ACAS benchmark shows increased node/edge counts (more data extracted)
- [ ] CardDemo benchmark shows increased edge counts (CALL USING, STRING, etc.)
### Quality Gates
- [ ] Each phase has its own commit
- [ ] Integration test assertions updated with exact counts per phase
- [ ] Benchmark run after each phase to track graph growth
## Dependencies & Risks
### Dependencies
- None. All changes are additive to existing COBOL processor code.
- No LanguageProvider changes needed.
- No graph schema changes needed (all new constructs map to existing node labels + edge types).
### Risks
- **preprocessor.ts size**: Currently 1326 LOC. Phase 1+2 adds ~200 LOC -> 1526 LOC. May need to extract helpers into a separate `cobol-data-flow.ts` module if it exceeds 1500.
- **REPLACE statement** (Phase 3.7) is the most complex feature — requires tracking text substitution state across logical lines. Consider deferring to a separate PR if it takes >100 LOC.
- **EXEC DLI** (Phase 2.1) is only testable against IMS codebases. Need fixture data or synthetic test cases.
## Graph Value Ranking by MCP Tool Impact
Research agent analyzed all 5 MCP tools (query, context, impact, detect_changes, rename) against planned edge types:
| Edge Type | QUERY | CONTEXT | IMPACT | DETECT | RENAME | **Overall** |
|-----------|-------|---------|--------|--------|--------|-------------|
| `cobol-call-using` | 4/5 | 5/5 | 5/5 | 4/5 | 4/5 | **9.2/10** |
| `cobol-error-handler` | 5/5 | 4/5 | 5/5 | 5/5 | 2/5 | **9.0/10** |
| `dli-*` (IMS verbs) | 4/5 | 4/5 | 5/5 | 4/5 | 2/5 | **8.2/10** |
| `cobol-string-*` | 4/5 | 3/5 | 3/5 | 3/5 | 2/5 | **6.2/10** |
**Key finding**: `cobol-call-using` alone would fix ~40% of missing caller references in COBOL graphs.
## Future Considerations
This plan provides the graph data foundation for a future `modernize` MCP command (out of scope) that would:
- Use CALL USING edges to map data contracts between programs
- Use STRING/UNSTRING edges to identify data transformation logic
- Use EXEC SQL/DLI edges to map database access patterns
- Use DECLARATIVES to understand error handling architecture
- Use the complete knowledge graph to generate migration plans
**MCP tool enhancements needed** (after this plan ships):
- Add `cobol-call-using`, `cobol-error-handler`, `dli-*` to IMPACT tool's default `relationTypes` for COBOL repos
- Add confidence floors for new edge types in `IMPACT_RELATION_CONFIDENCE`
- Register new edge types in `VALID_RELATION_TYPES` set (`local-backend.ts:52`)
## Sources & References
### Internal References
- Feature audit: session 8642401e (COBOL expert agent, 123 features audited)
- Prior plans: `docs/plans/2026-03-25-feat-cobol-100-percent-feature-coverage-plan.md`
- Architecture: `docs/code-indexing/cobol/` (7 documentation files)
### External References
- COBOL features reference: mainframestechhelp.com/tutorials/cobol/features.htm
- COBOL-85 standard: ISO/IEC 1989:1985
- IBM Enterprise COBOL reference

View file

@ -0,0 +1,55 @@
---
title: "Field Extractors for All Supported Languages"
type: feat
status: active
date: 2026-03-26
---
# Field Extractors for All Supported Languages
## Overview
PR #494 adds a `FieldExtractor` infrastructure with only a TypeScript implementation. This plan fills the registry for all 14 supported languages using a table-driven generic extractor, plus unit tests.
## Approach: Generic Table-Driven Extractor
Instead of 14 separate 300+ line files, create a `GenericFieldExtractor` configured via a per-language `FieldExtractionConfig`. Each config specifies:
- AST node types for type declarations (class_declaration, struct_item, etc.)
- AST node types for field declarations within bodies
- How to extract field name, type, visibility, static, readonly from the AST
- Default visibility and body node type
## Implementation
### Phase 1: Generic Field Extractor
**Create:** `field-extractors/generic.ts`
A single `createFieldExtractor(config)` factory that returns a `FieldExtractor` for any language.
### Phase 2: Language Configs
**Create:** `field-extractors/configs.ts`
Export configs for all 14 languages. Each config is ~20-40 lines of node type mappings.
### Phase 3: Register All in Index
**Modify:** `field-extractors/index.ts`
Register all 14 extractors. Keep the TypeScript-specific class for backwards compatibility.
### Phase 4: Unit Tests
**Create:** `test/unit/field-extraction-all-languages.test.ts`
For each language, parse a small code snippet and verify the extractor produces correct FieldInfo[].
## Acceptance Criteria
- [x] GenericFieldExtractor created with table-driven config
- [ ] Configs for: TS, JS, Python, Java, Kotlin, Go, Rust, C#, C++, C, PHP, Ruby, Swift, Dart
- [ ] All extractors registered in index.ts
- [ ] Unit tests for all languages
- [ ] `npx tsc --noEmit` passes
- [ ] All tests pass

View file

@ -42,4 +42,6 @@ export enum SupportedLanguages {
Kotlin = 'kotlin',
Swift = 'swift',
Dart = 'dart',
/** Standalone regex processor — no tree-sitter, no LanguageProvider. */
Cobol = 'cobol',
}

View file

@ -34,7 +34,15 @@ export const createKnowledgeGraph = (): KnowledgeGraph => {
};
/**
* Remove all nodes (and their relationships) belonging to a file
* Remove a single relationship by id.
* Returns true if the relationship existed and was removed, false otherwise.
*/
const removeRelationship = (relationshipId: string): boolean => {
return relationshipMap.delete(relationshipId);
};
/**
* Remove all nodes (and their relationships) belonging to a file.
*/
const removeNodesByFile = (filePath: string): number => {
let removed = 0;
@ -75,6 +83,7 @@ export const createKnowledgeGraph = (): KnowledgeGraph => {
addRelationship,
removeNode,
removeNodesByFile,
removeRelationship,
};
};

View file

@ -146,4 +146,5 @@ export interface KnowledgeGraph {
addRelationship: (relationship: GraphRelationship) => void,
removeNode: (nodeId: string) => boolean,
removeNodesByFile: (filePath: string) => number,
removeRelationship: (relationshipId: string) => boolean,
}

View file

@ -11,7 +11,6 @@ import { getLanguageFromFilename } from './utils/language-detection.js';
import { isVerboseIngestionEnabled } from './utils/verbose.js';
import { yieldToEventLoop } from './utils/event-loop.js';
import { FUNCTION_NODE_TYPES, extractFunctionName, findEnclosingClassId } from './utils/ast-helpers.js';
import { isBuiltInOrNoise } from './utils/noise-filter.js';
import {
countCallArguments,
inferCallForm,
@ -210,6 +209,26 @@ const findEnclosingFunction = (
return generateId(finalLabel, `${filePath}:${funcName}`);
}
}
// Language-specific enclosing function resolution (e.g., Dart where
// function_body is a sibling of function_signature, not a child).
if (provider.enclosingFunctionFinder) {
const customResult = provider.enclosingFunctionFinder(current);
if (customResult) {
// Try SymbolTable first (same pattern as the FUNCTION_NODE_TYPES branch above).
const resolved = ctx.resolve(customResult.funcName, filePath);
if (resolved?.tier === 'same-file' && resolved.candidates.length > 0) {
return resolved.candidates[0].nodeId;
}
let finalLabel = customResult.label;
if (provider.labelOverride) {
const override = provider.labelOverride(current.previousSibling!, finalLabel);
if (override !== null) finalLabel = override;
}
return generateId(finalLabel, `${filePath}:${customResult.funcName}`);
}
}
current = current.parent;
}
@ -497,7 +516,7 @@ export const processCalls = async (
}
}
if (isBuiltInOrNoise(calledName)) return;
if (provider.isBuiltInName(calledName)) return;
const callNode = captureMap['call'];
const callForm = inferCallForm(callNode, nameNode);

File diff suppressed because it is too large Load diff

View file

@ -0,0 +1,501 @@
/**
* COBOL COPY statement expansion engine.
*
* Expands COPY statements by inlining copybook content, applying REPLACING
* transformations (LEADING, TRAILING, EXACT), and handling nested copies
* with cycle detection.
*
* This is a preprocessing step that runs BEFORE extractCobolSymbolsWithRegex.
* The caller should run preprocessCobolSource first to clean patch markers.
*
* Supported syntax:
* COPY CPSESP.
* COPY "WORKGRID.CPY".
* COPY CPSESP REPLACING LEADING "ESP-" BY "LK-ESP-"
* LEADING "KPSESPL" BY "LK-KPSESPL".
* COPY ANAZI REPLACING "ANAZI-KEY" BY "LK-KEY".
*/
// ---------------------------------------------------------------------------
// Public interfaces
// ---------------------------------------------------------------------------
export interface CopyReplacing {
type: 'LEADING' | 'TRAILING' | 'EXACT';
from: string;
to: string;
isPseudotext?: boolean;
}
export interface CopyResolution {
copyTarget: string;
resolvedPath: string | null;
line: number;
replacing: CopyReplacing[];
library?: string;
}
export interface CopyExpansionResult {
expandedContent: string;
copyResolutions: CopyResolution[];
}
// ---------------------------------------------------------------------------
// Constants
// ---------------------------------------------------------------------------
export const DEFAULT_MAX_DEPTH = 10;
/** COBOL identifier pattern: starts with letter, contains letters, digits, hyphens. */
const RE_COBOL_IDENTIFIER = /\b([A-Z][A-Z0-9-]*)\b/gi;
// ---------------------------------------------------------------------------
// Private helpers
// ---------------------------------------------------------------------------
/**
* Strip inline comments (Italian-style `|` comments).
* Only strips if `|` appears in the code area (col 7+).
*/
function stripInlineComment(line: string): string {
let inQuote: string | null = null;
for (let i = 0; i < line.length; i++) {
const ch = line[i];
if (inQuote) {
if (ch === inQuote) inQuote = null;
} else if (ch === '"' || ch === "'") {
inQuote = ch;
} else if (ch === '|') {
return line.substring(0, i);
}
}
return line;
}
/**
* Check if a line is a COBOL comment (indicator in col 7 is `*` or `/`).
*/
function isCommentLine(line: string): boolean {
return line.length >= 7 && (line[6] === '*' || line[6] === '/');
}
/**
* Check if a line is a continuation line (indicator in col 7 is `-`).
*/
function isContinuationLine(line: string): boolean {
return line.length >= 7 && line[6] === '-';
}
/**
* Merge continuation lines into their predecessors.
* Returns an array of logical lines with their original starting line numbers.
*/
function mergeLogicalLines(
rawLines: string[],
): Array<{ text: string; lineNum: number }> {
const logical: Array<{ text: string; lineNum: number }> = [];
for (let i = 0; i < rawLines.length; i++) {
const raw = rawLines[i];
// Skip comment lines
if (isCommentLine(raw)) {
logical.push({ text: '', lineNum: i + 1 });
continue;
}
// Continuation: merge into previous logical line
if (isContinuationLine(raw)) {
if (logical.length > 0) {
const prev = logical[logical.length - 1];
const continuation = raw.length > 7 ? raw.substring(7).trimStart() : '';
prev.text += continuation;
}
// Push empty placeholder to preserve line count
logical.push({ text: '', lineNum: i + 1 });
continue;
}
// Normal line: strip inline comments
const cleaned = stripInlineComment(raw);
logical.push({ text: cleaned, lineNum: i + 1 });
}
return logical;
}
// ---------------------------------------------------------------------------
// COPY statement parsing
// ---------------------------------------------------------------------------
interface ParsedCopyStatement {
startLine: number;
endLine: number;
target: string;
replacing: CopyReplacing[];
library?: string;
}
/**
* Parse REPLACING clause text into structured replacements.
*
* Input examples:
* LEADING "ESP-" BY "LK-ESP-" LEADING "KPSESPL" BY "LK-KPSESPL"
* "ANAZI-KEY" BY "LK-KEY"
* TRAILING "-IN" BY "-OUT"
* ==CUST-== BY ==WS-CUST-==
* ==OLD-TEXT== BY ====
*/
export function parseReplacingClause(text: string): CopyReplacing[] {
const replacings: CopyReplacing[] = [];
if (!text || text.trim().length === 0) return replacings;
// Tokenize: ==pseudotext==, "quoted strings", or bare words.
// Pseudotext can contain spaces and single = chars but not ==.
interface TokenInfo { value: string; isPseudotext: boolean; }
const tokens: TokenInfo[] = [];
const tokenRe = /==((?:[^=]|=[^=])*)==|"([^"]*)"|(\S+)/g;
let tm: RegExpExecArray | null;
while ((tm = tokenRe.exec(text)) !== null) {
if (tm[1] !== undefined) {
// Pseudotext: trim leading/trailing whitespace
tokens.push({ value: tm[1].trim(), isPseudotext: true });
} else if (tm[2] !== undefined) {
tokens.push({ value: tm[2], isPseudotext: false });
} else {
tokens.push({ value: tm[3], isPseudotext: false });
}
}
// Parse token stream: [LEADING|TRAILING]? <from> BY <to>
let i = 0;
while (i < tokens.length) {
let type: CopyReplacing['type'] = 'EXACT';
// Check for type modifier (only on non-pseudotext tokens)
if (!tokens[i].isPseudotext) {
const upper = tokens[i].value.toUpperCase();
if (upper === 'LEADING') {
type = 'LEADING';
i++;
} else if (upper === 'TRAILING') {
type = 'TRAILING';
i++;
}
}
if (i >= tokens.length) break;
const fromToken = tokens[i];
i++;
// Pseudotext always forces EXACT type
if (fromToken.isPseudotext) type = 'EXACT';
// Expect BY keyword
if (i >= tokens.length) break;
if (tokens[i].value.toUpperCase() !== 'BY') {
// Malformed — skip this token and try to resync
continue;
}
i++; // skip BY
if (i >= tokens.length) break;
const toToken = tokens[i];
i++;
replacings.push({ type, from: fromToken.value, to: toToken.value, isPseudotext: fromToken.isPseudotext || undefined });
}
return replacings;
}
/**
* Scan logical lines for COPY statements.
* COPY statements can span multiple lines and terminate with a period.
*/
function parseCopyStatements(
logicalLines: Array<{ text: string; lineNum: number }>,
): ParsedCopyStatement[] {
const results: ParsedCopyStatement[] = [];
let accumulator: string | null = null;
let startLine = 0;
let endLine = 0;
for (let i = 0; i < logicalLines.length; i++) {
const { text, lineNum } = logicalLines[i];
if (text.length === 0) continue;
// Check for COPY keyword start (not inside a string context)
const copyStart = text.match(/\bCOPY\b/i);
if (accumulator === null) {
if (!copyStart) continue;
// Start accumulating from the COPY keyword onwards
const copyIdx = copyStart.index!;
accumulator = text.substring(copyIdx);
startLine = lineNum;
endLine = lineNum;
} else {
// Continue accumulating
accumulator += ' ' + text.trim();
endLine = lineNum;
}
// Check if statement terminates (period at end of accumulated text)
if (accumulator !== null && /\.\s*$/.test(accumulator)) {
const parsed = parseSingleCopyStatement(accumulator, startLine, endLine);
if (parsed) {
results.push(parsed);
}
accumulator = null;
}
}
// If there's an unterminated COPY (missing period), try to parse what we have
if (accumulator !== null) {
const parsed = parseSingleCopyStatement(accumulator, startLine, endLine);
if (parsed) {
results.push(parsed);
}
}
return results;
}
/**
* Parse a single complete COPY statement string.
*
* Formats:
* COPY target.
* COPY "target".
* COPY target REPLACING ... .
*/
function parseSingleCopyStatement(
stmt: string,
startLine: number,
endLine: number,
): ParsedCopyStatement | null {
// Strip terminating period
const text = stmt.replace(/\.\s*$/, '').trim();
// Extract target: COPY <target> or COPY "<target>" or COPY '<target>'
// Optionally followed by IN/OF <library-name> (COBOL-85 standard: IN and OF are synonyms)
const targetMatch = text.match(
/^COPY\s+(?:"([^"]+)"|'([^']+)'|([A-Z][A-Z0-9-]*))(?:\s+(?:IN|OF)\s+([A-Z][A-Z0-9-]*))?/i,
);
if (!targetMatch) return null;
const target = targetMatch[1] ?? targetMatch[2] ?? targetMatch[3];
const library = targetMatch[4] || undefined;
// Extract REPLACING clause if present
let replacing: CopyReplacing[] = [];
const replacingIdx = text.search(/\bREPLACING\b/i);
if (replacingIdx >= 0) {
const replacingText = text.substring(replacingIdx + 'REPLACING'.length);
replacing = parseReplacingClause(replacingText);
}
return { startLine, endLine, target, replacing, library };
}
// ---------------------------------------------------------------------------
// REPLACING application
// ---------------------------------------------------------------------------
/**
* Apply REPLACING transformations to copybook content.
*
* LEADING: replace prefix in COBOL identifiers.
* TRAILING: replace suffix in COBOL identifiers.
* EXACT: replace exact token matches.
*/
function applyReplacing(content: string, replacings: CopyReplacing[]): string {
if (replacings.length === 0) return content;
// First pass: handle EXACT replacements that contain spaces or non-identifier
// characters (pseudotext). These cannot be handled by identifier-level matching.
let result = content;
for (const r of replacings) {
if (r.type === 'EXACT' && (r.isPseudotext || r.from.includes(' ') || !/^[A-Z][A-Z0-9-]*$/i.test(r.from))) {
const escaped = r.from.replace(/[.*+?^${}()|[\]\\]/g, '\\$&');
const re = new RegExp(escaped, 'gi');
result = result.replace(re, r.to);
}
}
// Second pass: identifier-level replacements (LEADING, TRAILING, single-word EXACT)
const identifierReplacings = replacings.filter(
r => !(r.type === 'EXACT' && (r.isPseudotext || r.from.includes(' ') || !/^[A-Z][A-Z0-9-]*$/i.test(r.from))),
);
if (identifierReplacings.length === 0) return result;
return result.replace(RE_COBOL_IDENTIFIER, (match) => {
for (const r of identifierReplacings) {
const upper = match.toUpperCase();
const from = r.from.toUpperCase();
const to = r.to.toUpperCase();
switch (r.type) {
case 'LEADING':
if (upper.startsWith(from)) {
return to + match.substring(from.length);
}
break;
case 'TRAILING':
if (upper.endsWith(from)) {
return match.substring(0, match.length - from.length) + to;
}
break;
case 'EXACT':
if (upper === from) {
return to;
}
break;
}
}
return match;
});
}
// ---------------------------------------------------------------------------
// Main expansion engine
// ---------------------------------------------------------------------------
/**
* Expand COBOL COPY statements by inlining copybook content.
*
* @param content - Source COBOL content (after preprocessCobolSource)
* @param filePath - Path of the source file (for diagnostics)
* @param resolveFile - Maps a COPY target name to a filesystem path, or null if not found
* @param readFile - Reads file content by path, or null if unreadable
* @param maxDepth - Maximum nesting depth for recursive expansion (default: 10)
* @returns Expanded content and resolution metadata
*/
export function expandCopies(
content: string,
filePath: string,
resolveFile: (name: string) => string | null,
readFile: (path: string) => string | null,
maxDepth: number = DEFAULT_MAX_DEPTH,
): CopyExpansionResult {
const allResolutions: CopyResolution[] = [];
const warnedCircular = new Set<string>();
let totalExpansions = 0;
const MAX_TOTAL_EXPANSIONS = 500;
const expanded = expandRecursive(content, filePath, 0, new Set<string>());
return {
expandedContent: expanded,
copyResolutions: allResolutions,
};
/**
* Recursively expand COPY statements in content.
*
* @param src - Source content to expand
* @param srcPath - Path of the file being expanded (for cycle detection logging)
* @param depth - Current recursion depth
* @param visited - Set of already-visited copybook paths (cycle detection)
*/
function expandRecursive(
src: string,
srcPath: string,
depth: number,
visited: Set<string>,
): string {
const rawLines = src.split(/\r?\n/);
const logicalLines = mergeLogicalLines(rawLines);
const copyStatements = parseCopyStatements(logicalLines);
// No COPY statements — return as-is
if (copyStatements.length === 0) return src;
// Process COPY statements in reverse order so line numbers stay valid
// as we splice content
const outputLines = [...rawLines];
for (let ci = copyStatements.length - 1; ci >= 0; ci--) {
const cs = copyStatements[ci];
// Resolve the copybook path
const resolvedPath = resolveFile(cs.target);
// Record resolution metadata
allResolutions.push({
copyTarget: cs.target,
resolvedPath,
line: cs.startLine,
replacing: cs.replacing,
library: cs.library,
});
// Cannot resolve — keep original lines
if (resolvedPath === null) {
continue;
}
// Cycle detection
if (visited.has(resolvedPath)) {
if (!warnedCircular.has(resolvedPath)) {
warnedCircular.add(resolvedPath);
console.warn(
`[cobol-copy-expander] Circular COPY detected: ${cs.target} (${resolvedPath}) ` +
`includes itself. Skipping expansion.`,
);
}
continue;
}
// Max depth exceeded — keep unexpanded
if (depth >= maxDepth) {
console.warn(
`[cobol-copy-expander] Max expansion depth (${maxDepth}) reached for ` +
`COPY ${cs.target} in ${srcPath}. Skipping expansion.`,
);
continue;
}
// Guard against exponential breadth amplification (N copybooks each with N COPYs)
if (++totalExpansions > MAX_TOTAL_EXPANSIONS) {
if (!warnedCircular.has('__max_total__')) {
warnedCircular.add('__max_total__');
console.warn(
`[cobol-copy-expander] Max total expansions (${MAX_TOTAL_EXPANSIONS}) reached ` +
`in ${srcPath}. Skipping further expansions.`,
);
}
continue;
}
// Read the copybook content
const copybookContent = readFile(resolvedPath);
if (copybookContent === null) {
continue;
}
// Apply REPLACING transformations
const replaced = applyReplacing(copybookContent, cs.replacing);
// Recurse into the copybook for nested COPYs
const nestedVisited = new Set(visited);
nestedVisited.add(resolvedPath);
const expandedCopybook = expandRecursive(
replaced,
resolvedPath,
depth + 1,
nestedVisited,
);
// Splice: replace the COPY statement lines with expanded content
// startLine/endLine are 1-based; convert to 0-based array index
const expansionLines = expandedCopybook.split('\n');
const removeCount = cs.endLine - cs.startLine + 1;
outputLines.splice(cs.startLine - 1, removeCount, ...expansionLines);
}
return outputLines.join('\n');
}
}

File diff suppressed because it is too large Load diff

View file

@ -0,0 +1,263 @@
/**
* JCL Parser — Regex single-pass extraction.
*
* Extracts JCL constructs from mainframe job streams:
* - JOB statements (job name, CLASS, MSGCLASS)
* - EXEC statements (step -> program or proc)
* - DD statements (dataset references, DISP)
* - PROC definitions (in-stream and catalogued)
* - INCLUDE MEMBER= directives
* - SET symbolic parameters
* - IF/ELSE/ENDIF conditional execution
* - JCLLIB ORDER= search paths
*
* Pattern follows cobol-preprocessor.ts — regex-only, no tree-sitter.
*/
export interface JclParseResults {
jobs: Array<{ name: string; line: number; class?: string; msgclass?: string }>;
steps: Array<{ name: string; jobName: string; program?: string; proc?: string; line: number }>;
ddStatements: Array<{ ddName: string; stepName: string; dataset?: string; disp?: string; line: number }>;
procs: Array<{ name: string; line: number; isInStream: boolean }>;
includes: Array<{ member: string; line: number }>;
sets: Array<{ variable: string; value: string; line: number }>;
jcllib: Array<{ order: string[]; line: number }>;
conditionals: Array<{ type: 'IF' | 'ELSE' | 'ENDIF'; condition?: string; line: number }>;
}
// ── JCL statement patterns ─────────────────────────────────────────────
// JCL continuation: line ends with a non-blank in col 72, next line starts with //
// We handle continuations by joining lines before matching.
/** Match //jobname JOB ... */
const JOB_RE = /^\/\/(\w{1,8})\s+JOB\s+(.*)/i;
/** Match //stepname EXEC PGM=program or //stepname EXEC procname */
const EXEC_RE = /^\/\/(\w{1,8})\s+EXEC\s+(.*)/i;
/** Match //ddname DD ... */
const DD_RE = /^\/\/(\w{1,8})\s+DD\s+(.*)/i;
/** Match // JCLLIB ORDER=(lib1,lib2,...) */
const JCLLIB_RE = /^\/\/\s+JCLLIB\s+ORDER=\(([^)]+)\)/i;
/** Match // IF condition THEN */
const IF_RE = /^\/\/\s+IF\s+(.+)\s+THEN/i;
/** Match // ELSE */
const ELSE_RE = /^\/\/\s+ELSE\b/i;
/** Match // ENDIF */
const ENDIF_RE = /^\/\/\s+ENDIF\b/i;
/** Match // INCLUDE MEMBER=name */
const INCLUDE_RE = /^\/\/\s+INCLUDE\s+MEMBER=(\w+)/i;
/** Match // SET var=value */
const SET_RE = /^\/\/\s+SET\s+(\w+)=(.+)/i;
/** Match // PROC or //name PROC */
const PROC_RE = /^\/\/(\w*)\s+PROC\b/i;
/** Match // PEND */
const PEND_RE = /^\/\/\s+PEND\b/i;
// ── Parameter extractors ───────────────────────────────────────────────
function extractParam(params: string, key: string): string | undefined {
// Match KEY=VALUE or KEY='VALUE' in JCL parameter string
const re = new RegExp(`${key}=(?:'([^']*)'|(\\S+?))(?:[,\\s]|$)`, 'i');
const m = params.match(re);
return m ? (m[1] ?? m[2]) : undefined;
}
function extractPgm(params: string): string | undefined {
return extractParam(params, 'PGM');
}
function extractProc(params: string): string | undefined {
// If no PGM= keyword, the first positional parameter is the proc name
if (/PGM=/i.test(params)) return undefined;
const cleaned = params.replace(/,.*/, '').trim();
// Proc name is the first token (no = sign)
if (cleaned && !cleaned.includes('=')) {
return cleaned.replace(/[,\s].*/s, '').toUpperCase();
}
return undefined;
}
function extractDsn(params: string): string | undefined {
return extractParam(params, 'DSN') ?? extractParam(params, 'DSNAME');
}
function extractDisp(params: string): string | undefined {
const m = params.match(/DISP=\(?\s*([^),\s]+)/i);
return m ? m[1] : undefined;
}
/**
* Parse a JCL file and extract all constructs.
*
* @param content - Raw JCL file content
* @param filePath - Path for diagnostics (not used in extraction)
* @returns Parsed JCL results
*/
export function parseJcl(content: string, filePath: string): JclParseResults {
const results: JclParseResults = {
jobs: [],
steps: [],
ddStatements: [],
procs: [],
includes: [],
sets: [],
jcllib: [],
conditionals: [],
};
const rawLines = content.split(/\r?\n/);
// Join continuation lines: a line ending with non-blank in col 71 (0-indexed)
// followed by a line starting with // is a continuation.
const lines: Array<{ text: string; lineNum: number }> = [];
let i = 0;
while (i < rawLines.length) {
let line = rawLines[i];
const lineNum = i + 1;
// JCL continuation: if line is exactly 72+ chars and col 72 is non-blank
// and the next line starts with //, join them.
while (
i + 1 < rawLines.length &&
line.length >= 72 &&
line[71] !== ' ' &&
rawLines[i + 1].startsWith('//')
) {
i++;
// Continuation text starts after // and leading spaces
const contText = rawLines[i].substring(2).replace(/^\s+/, ' ');
// Remove the continuation marker (col 72+) from current line
line = line.substring(0, 71).trimEnd() + contText;
}
lines.push({ text: line, lineNum });
i++;
}
let currentJobName = '';
let currentStepName = '';
let inStreamProcName = '';
for (const { text, lineNum } of lines) {
// Skip JCL comments (starting with //* )
if (text.startsWith('//*')) continue;
// Skip non-JCL lines (don't start with //)
if (!text.startsWith('//')) continue;
// PROC definition (in-stream)
const procMatch = text.match(PROC_RE);
if (procMatch) {
const procName = procMatch[1] || inStreamProcName;
if (procName) {
results.procs.push({ name: procName.toUpperCase(), line: lineNum, isInStream: true });
}
inStreamProcName = procName?.toUpperCase() || '';
continue;
}
// PEND (end of in-stream proc)
if (PEND_RE.test(text)) {
inStreamProcName = '';
continue;
}
// JCLLIB ORDER=
const jcllibMatch = text.match(JCLLIB_RE);
if (jcllibMatch) {
const libs = jcllibMatch[1].split(',').map(s => s.trim().replace(/'/g, ''));
results.jcllib.push({ order: libs, line: lineNum });
continue;
}
// IF/ELSE/ENDIF
const ifMatch = text.match(IF_RE);
if (ifMatch) {
results.conditionals.push({ type: 'IF', condition: ifMatch[1].trim(), line: lineNum });
continue;
}
if (ELSE_RE.test(text)) {
results.conditionals.push({ type: 'ELSE', line: lineNum });
continue;
}
if (ENDIF_RE.test(text)) {
results.conditionals.push({ type: 'ENDIF', line: lineNum });
continue;
}
// INCLUDE MEMBER=
const includeMatch = text.match(INCLUDE_RE);
if (includeMatch) {
results.includes.push({ member: includeMatch[1].toUpperCase(), line: lineNum });
continue;
}
// SET var=value
const setMatch = text.match(SET_RE);
if (setMatch) {
results.sets.push({
variable: setMatch[1].toUpperCase(),
value: setMatch[2].trim().replace(/,\s*$/, ''),
line: lineNum,
});
continue;
}
// JOB statement
const jobMatch = text.match(JOB_RE);
if (jobMatch) {
currentJobName = jobMatch[1].toUpperCase();
const params = jobMatch[2];
results.jobs.push({
name: currentJobName,
line: lineNum,
class: extractParam(params, 'CLASS'),
msgclass: extractParam(params, 'MSGCLASS'),
});
continue;
}
// EXEC statement
const execMatch = text.match(EXEC_RE);
if (execMatch) {
currentStepName = execMatch[1].toUpperCase();
const params = execMatch[2];
const pgm = extractPgm(params);
const proc = pgm ? undefined : extractProc(params);
results.steps.push({
name: currentStepName,
jobName: currentJobName,
program: pgm?.toUpperCase(),
proc: proc?.toUpperCase(),
line: lineNum,
});
continue;
}
// DD statement
const ddMatch = text.match(DD_RE);
if (ddMatch) {
const ddName = ddMatch[1].toUpperCase();
const params = ddMatch[2];
results.ddStatements.push({
ddName,
stepName: currentStepName,
dataset: extractDsn(params)?.toUpperCase(),
disp: extractDisp(params)?.toUpperCase(),
line: lineNum,
});
continue;
}
}
return results;
}

View file

@ -0,0 +1,274 @@
/**
* JCL Processor — Converts JCL parse results into graph nodes and edges.
*
* Maps JCL entities to existing graph types (no new tables):
* - Job -> CodeElement (description: "jcl-job class:A msgclass:X")
* - Step -> CodeElement (description: "jcl-step pgm:PROGRAMNAME")
* - Dataset -> CodeElement (description: "jcl-dataset disp:SHR")
* - PROC -> Module
*
* Edges:
* - Job CONTAINS Step
* - Step CALLS Module (when PGM= matches an indexed program)
* - Step references Dataset (CALLS edge with reason "jcl-dd")
* - Job/Step IMPORTS PROC
*
* Pattern follows detectCrossProgamContracts() in pipeline.ts.
*/
import { parseJcl, type JclParseResults } from './jcl-parser.js';
import type { KnowledgeGraph } from '../../graph/types.js';
import { generateId } from '../../../lib/utils.js';
export interface JclProcessResult {
jobCount: number;
stepCount: number;
datasetCount: number;
programLinks: number;
}
/**
* Process JCL files and integrate into the knowledge graph.
*
* @param graph - The in-memory knowledge graph
* @param jclPaths - File paths of JCL files
* @param jclContents - Map of path -> file content
* @returns Summary of what was added
*/
export function processJclFiles(
graph: KnowledgeGraph,
jclPaths: string[],
jclContents: Map<string, string>,
): JclProcessResult {
let jobCount = 0;
let stepCount = 0;
let datasetCount = 0;
let programLinks = 0;
// Collect all Module names for step -> program linking
const moduleNames = new Map<string, string>(); // uppercase name -> node id
graph.forEachNode(node => {
if (node.label === 'Module') {
const nodeName = node.properties.name;
if (typeof nodeName === 'string') {
moduleNames.set(nodeName.toUpperCase(), node.id);
}
}
});
for (const filePath of jclPaths) {
const content = jclContents.get(filePath);
if (!content) continue;
const parsed = parseJcl(content, filePath);
const result = integrateJclResults(graph, parsed, filePath, moduleNames);
jobCount += result.jobCount;
stepCount += result.stepCount;
datasetCount += result.datasetCount;
programLinks += result.programLinks;
}
return { jobCount, stepCount, datasetCount, programLinks };
}
function integrateJclResults(
graph: KnowledgeGraph,
parsed: JclParseResults,
filePath: string,
moduleNames: Map<string, string>,
): JclProcessResult {
let jobCount = 0;
let stepCount = 0;
let datasetCount = 0;
let programLinks = 0;
// Track step node IDs for DD -> step linking
const stepNodeIds = new Map<string, string>(); // stepName -> nodeId
// 1. Create Job nodes
for (const job of parsed.jobs) {
const jobId = generateId('CodeElement', `${filePath}:job:${job.name}`);
const classPart = job.class ? ` class:${job.class}` : '';
const msgPart = job.msgclass ? ` msgclass:${job.msgclass}` : '';
graph.addNode({
id: jobId,
label: 'CodeElement',
properties: {
name: job.name,
filePath,
startLine: job.line,
endLine: job.line,
description: `jcl-job${classPart}${msgPart}`,
},
});
// Link File -> Job (CONTAINS)
const fileId = generateId('File', filePath);
graph.addRelationship({
id: `${fileId}_contains_${jobId}`,
type: 'CONTAINS',
sourceId: fileId,
targetId: jobId,
confidence: 1.0,
reason: 'jcl-job',
});
jobCount++;
}
// 1.5 Pre-register in-stream PROCs so steps can reference them
// (fixes ordering bug: steps processed before PROCs were registered)
for (const proc of parsed.procs) {
const procId = generateId('Module', `${filePath}:proc:${proc.name}`);
moduleNames.set(proc.name.toUpperCase(), procId);
}
// 2. Create Step nodes and link to programs
for (const step of parsed.steps) {
const stepId = generateId('CodeElement', `${filePath}:step:${step.jobName}:${step.name}`);
const pgmPart = step.program ? ` pgm:${step.program}` : '';
const procPart = step.proc ? ` proc:${step.proc}` : '';
graph.addNode({
id: stepId,
label: 'CodeElement',
properties: {
name: step.name,
filePath,
startLine: step.line,
endLine: step.line,
description: `jcl-step${pgmPart}${procPart}`,
},
});
stepNodeIds.set(step.name, stepId);
// Link Job -> Step (CONTAINS)
if (step.jobName) {
const jobId = generateId('CodeElement', `${filePath}:job:${step.jobName}`);
graph.addRelationship({
id: `${jobId}_contains_${stepId}`,
type: 'CONTAINS',
sourceId: jobId,
targetId: stepId,
confidence: 1.0,
reason: 'jcl-step',
});
}
// Link Step -> Module (CALLS) when PGM= matches an indexed program
if (step.program) {
const moduleId = moduleNames.get(step.program.toUpperCase());
if (moduleId) {
graph.addRelationship({
id: `${stepId}_calls_${moduleId}`,
type: 'CALLS',
sourceId: stepId,
targetId: moduleId,
confidence: 0.95,
reason: 'jcl-exec-pgm',
});
programLinks++;
}
}
// Link Step -> PROC (CALLS) — PROC as Module
if (step.proc) {
const procModuleId = moduleNames.get(step.proc.toUpperCase());
if (procModuleId) {
graph.addRelationship({
id: `${stepId}_calls_proc_${procModuleId}`,
type: 'CALLS',
sourceId: stepId,
targetId: procModuleId,
confidence: 0.9,
reason: 'jcl-exec-proc',
});
}
}
stepCount++;
}
// 3. Create Dataset nodes from DD statements
const seenDatasets = new Set<string>();
for (const dd of parsed.ddStatements) {
if (!dd.dataset) continue;
// Create dataset node (deduplicated per file)
const datasetKey = `${filePath}:dataset:${dd.dataset}`;
const datasetId = generateId('CodeElement', datasetKey);
if (!seenDatasets.has(dd.dataset)) {
const dispPart = dd.disp ? ` disp:${dd.disp}` : '';
graph.addNode({
id: datasetId,
label: 'CodeElement',
properties: {
name: dd.dataset,
filePath,
startLine: dd.line,
endLine: dd.line,
description: `jcl-dataset${dispPart}`,
},
});
seenDatasets.add(dd.dataset);
datasetCount++;
}
// Link Step -> Dataset (CALLS with reason jcl-dd)
const stepId = stepNodeIds.get(dd.stepName);
if (stepId) {
graph.addRelationship({
id: `${stepId}_dd_${dd.ddName}_${datasetId}`,
type: 'CALLS',
sourceId: stepId,
targetId: datasetId,
confidence: 0.85,
reason: `jcl-dd:${dd.ddName}`,
});
}
}
// 4. Create PROC nodes (in-stream procs as Module)
for (const proc of parsed.procs) {
if (!proc.isInStream) continue;
const procId = generateId('Module', `${filePath}:proc:${proc.name}`);
graph.addNode({
id: procId,
label: 'Module',
properties: {
name: proc.name,
filePath,
startLine: proc.line,
endLine: proc.line,
description: 'jcl-proc-instream',
},
});
// Register for step linking
moduleNames.set(proc.name.toUpperCase(), procId);
}
// 5. INCLUDE directives -> IMPORTS edges
for (const inc of parsed.includes) {
const moduleId = moduleNames.get(inc.member.toUpperCase());
if (moduleId) {
const fileId = generateId('File', filePath);
graph.addRelationship({
id: `${fileId}_includes_${moduleId}`,
type: 'IMPORTS',
sourceId: fileId,
targetId: moduleId,
confidence: 0.9,
reason: 'jcl-include',
});
}
}
return { jobCount, stepCount, datasetCount, programLinks };
}

View file

@ -226,6 +226,7 @@ export const ENTRY_POINT_PATTERNS = {
/^onEvent$/, // BLoC event handler
/^mapEventToState$/, // Legacy BLoC pattern
],
[SupportedLanguages.Cobol]: [], // Standalone regex processor — no tree-sitter entry points
} satisfies Record<SupportedLanguages, RegExp[]>;
/** Pre-computed merged patterns (universal + language-specific) to avoid per-call array allocation. */
@ -325,7 +326,7 @@ export function calculateEntryPointScore(
// Check positive patterns
const allPatterns = MERGED_ENTRY_POINT_PATTERNS[language];
if (allPatterns.some(p => p.test(name))) {
if (allPatterns?.some(p => p.test(name))) {
nameMultiplier = 1.5; // Bonus for matching entry point pattern
reasons.push('entry-pattern');
}

View file

@ -601,6 +601,7 @@ export const AST_FRAMEWORK_PATTERNS_BY_LANGUAGE = {
{ framework: 'flutter', entryPointMultiplier: 2.5, reason: 'flutter-widget', patterns: FRAMEWORK_AST_PATTERNS.flutter },
{ framework: 'riverpod', entryPointMultiplier: 2.8, reason: 'riverpod-pattern', patterns: FRAMEWORK_AST_PATTERNS.riverpod },
],
[SupportedLanguages.Cobol]: [], // Standalone regex processor — no AST framework patterns
} satisfies Record<SupportedLanguages, AstFrameworkPatternConfig[]>;
/** Pre-lowercased patterns for O(1) pattern matching at runtime */

View file

@ -39,6 +39,12 @@ export function resolveDartImport(
return null;
}
// Relative imports — use standard resolution
return resolveStandard(stripped, filePath, ctx, SupportedLanguages.Dart);
// Relative imports — use standard resolution.
// Dart relative imports don't require a leading "./" (e.g. `import 'models.dart'`).
// The standard resolver only recognises paths starting with "." as relative, so
// prepend "./" when the path doesn't already start with "." to ensure correct
// same-directory resolution (without this, "models.dart" would be mangled by the
// generic dot-to-slash conversion intended for Java-style package imports).
const relPath = stripped.startsWith('.') ? stripped : './' + stripped;
return resolveStandard(relPath, filePath, ctx, SupportedLanguages.Dart);
}

View file

@ -41,7 +41,12 @@ interface LanguageProviderConfig {
readonly extensions: readonly string[];
// ── Parser ────────────────────────────────────────────────────────
/** Tree-sitter query strings for definitions, imports, calls, heritage */
/** Parse strategy: 'tree-sitter' (default) uses AST parsing via tree-sitter.
* 'standalone' means the language has its own regex-based processor and
* should be skipped by the tree-sitter pipeline (e.g., COBOL, Markdown). */
readonly parseStrategy?: 'tree-sitter' | 'standalone';
/** Tree-sitter query strings for definitions, imports, calls, heritage.
* Required for tree-sitter languages; empty string for standalone processors. */
readonly treeSitterQueries: string;
// ── Core (required) ───────────────────────────────────────────────
@ -80,6 +85,17 @@ interface LanguageProviderConfig {
projectConfig: unknown,
) => void;
// ── Enclosing function resolution ───────────────────────────────
/** Resolve the enclosing function name + label from an AST ancestor node
* that is NOT a standard FUNCTION_NODE_TYPE. For languages where the
* function body is a sibling of the signature (e.g. Dart: function_body ↔
* function_signature are siblings under program/class_body), the default
* parent walk cannot find the enclosing function. This hook lets the
* language provider inspect each ancestor and return the resolved result.
* Return null to continue the default walk.
* Default: undefined (standard parent walk only). */
readonly enclosingFunctionFinder?: (ancestorNode: SyntaxNode) => { funcName: string; label: NodeLabel } | null;
// ── Labels ────────────────────────────────────────────────────────
/** Override the default node label for definition.function captures.
* Return null to skip (C/C++ duplicate), a different label to reclassify
@ -115,6 +131,11 @@ interface LanguageProviderConfig {
* When true, the worker extracts routes via the language's route extraction logic.
* Default: undefined (no route files). */
readonly isRouteFile?: (filePath: string) => boolean;
// ── Noise filtering ────────────────────────────────────────────────
/** Built-in/stdlib names that should be filtered from the call graph for this language.
* Default: undefined (no language-specific filtering). */
readonly builtInNames?: ReadonlySet<string>;
}
/** Runtime type — same as LanguageProviderConfig but with defaults guaranteed present. */
@ -124,6 +145,8 @@ export interface LanguageProvider extends Omit<LanguageProviderConfig,
readonly importSemantics: ImportSemantics;
readonly heritageDefaultEdge: 'EXTENDS' | 'IMPLEMENTS';
readonly mroStrategy: MroStrategy;
/** Check if a name is a built-in/stdlib function that should be filtered from the call graph. */
readonly isBuiltInName: (name: string) => boolean;
}
const DEFAULTS: Pick<LanguageProvider, 'importSemantics' | 'heritageDefaultEdge' | 'mroStrategy'> = {
@ -134,5 +157,10 @@ const DEFAULTS: Pick<LanguageProvider, 'importSemantics' | 'heritageDefaultEdge'
/** Define a language provider — required fields must be supplied, optional fields get sensible defaults. */
export function defineLanguage(config: LanguageProviderConfig): LanguageProvider {
return { ...DEFAULTS, ...config };
const builtIns = config.builtInNames;
return {
...DEFAULTS,
...config,
isBuiltInName: builtIns ? (name: string) => builtIns.has(name) : () => false,
};
}

View file

@ -20,6 +20,28 @@ import type { LanguageProvider } from '../language-provider.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { cConfig as cFieldConfig, cppConfig as cppFieldConfig } from '../field-extractors/configs/c-cpp.js';
const C_BUILT_INS: ReadonlySet<string> = new Set([
'printf', 'fprintf', 'sprintf', 'snprintf', 'vprintf', 'vfprintf', 'vsprintf', 'vsnprintf',
'scanf', 'fscanf', 'sscanf',
'malloc', 'calloc', 'realloc', 'free', 'memcpy', 'memmove', 'memset', 'memcmp',
'strlen', 'strcpy', 'strncpy', 'strcat', 'strncat', 'strcmp', 'strncmp', 'strstr', 'strchr', 'strrchr',
'atoi', 'atol', 'atof', 'strtol', 'strtoul', 'strtoll', 'strtoull', 'strtod',
'sizeof', 'offsetof', 'typeof',
'assert', 'abort', 'exit', '_exit',
'fopen', 'fclose', 'fread', 'fwrite', 'fseek', 'ftell', 'rewind', 'fflush', 'fgets', 'fputs',
'likely', 'unlikely', 'BUG', 'BUG_ON', 'WARN', 'WARN_ON', 'WARN_ONCE',
'IS_ERR', 'PTR_ERR', 'ERR_PTR', 'IS_ERR_OR_NULL',
'ARRAY_SIZE', 'container_of', 'list_for_each_entry', 'list_for_each_entry_safe',
'min', 'max', 'clamp', 'abs', 'swap',
'pr_info', 'pr_warn', 'pr_err', 'pr_debug', 'pr_notice', 'pr_crit', 'pr_emerg',
'printk', 'dev_info', 'dev_warn', 'dev_err', 'dev_dbg',
'GFP_KERNEL', 'GFP_ATOMIC',
'spin_lock', 'spin_unlock', 'spin_lock_irqsave', 'spin_unlock_irqrestore',
'mutex_lock', 'mutex_unlock', 'mutex_init',
'kfree', 'kmalloc', 'kzalloc', 'kcalloc', 'krealloc', 'kvmalloc', 'kvfree',
'get', 'put',
]);
/** Label override shared by C and C++: skip function_definition captures inside class/struct
* bodies (they're duplicates of definition.method captures). */
const cppLabelOverride: NonNullable<LanguageProvider['labelOverride']> = (functionNode, defaultLabel) => {
@ -37,6 +59,7 @@ export const cProvider = defineLanguage({
importSemantics: 'wildcard',
fieldExtractor: createFieldExtractor(cFieldConfig),
labelOverride: cppLabelOverride,
builtInNames: C_BUILT_INS,
});
export const cppProvider = defineLanguage({
@ -50,4 +73,5 @@ export const cppProvider = defineLanguage({
mroStrategy: 'leftmost-base',
fieldExtractor: createFieldExtractor(cppFieldConfig),
labelOverride: cppLabelOverride,
builtInNames: C_BUILT_INS,
});

View file

@ -0,0 +1,27 @@
/**
* COBOL Language Provider
*
* Standalone regex-based processor — no tree-sitter grammar.
* COBOL files (.cbl, .cob, .cobol, .cpy, .copybook) are detected and
* processed by cobol-processor.ts in pipeline Phase 2.6, not by the
* tree-sitter pipeline.
*
* This provider exists to satisfy the SupportedLanguages exhaustiveness
* checks and to declare parseStrategy: 'standalone'.
*/
import { SupportedLanguages } from '../../../config/supported-languages.js';
import { defineLanguage } from '../language-provider.js';
export const cobolProvider = defineLanguage({
id: SupportedLanguages.Cobol,
parseStrategy: 'standalone',
extensions: [], // COBOL files detected by cobol-processor's isCobolFile/isJclFile
treeSitterQueries: '',
typeConfig: {
declarationNodeTypes: new Set(),
extractDeclaration: () => null,
extractParameter: () => null,
},
exportChecker: () => false,
importResolver: () => null,
});

View file

@ -16,6 +16,27 @@ import { CSHARP_QUERIES } from '../tree-sitter-queries.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { csharpConfig as csharpFieldConfig } from '../field-extractors/configs/csharp.js';
const BUILT_INS: ReadonlySet<string> = new Set([
'Console', 'WriteLine', 'ReadLine', 'Write',
'Task', 'Run', 'Wait', 'WhenAll', 'WhenAny', 'FromResult', 'Delay', 'ContinueWith',
'ConfigureAwait', 'GetAwaiter', 'GetResult',
'ToString', 'GetType', 'Equals', 'GetHashCode', 'ReferenceEquals',
'Add', 'Remove', 'Contains', 'Clear', 'Count', 'Any', 'All',
'Where', 'Select', 'SelectMany', 'OrderBy', 'OrderByDescending', 'GroupBy',
'First', 'FirstOrDefault', 'Single', 'SingleOrDefault', 'Last', 'LastOrDefault',
'ToList', 'ToArray', 'ToDictionary', 'AsEnumerable', 'AsQueryable',
'Aggregate', 'Sum', 'Average', 'Min', 'Max', 'Distinct', 'Skip', 'Take',
'String', 'Format', 'IsNullOrEmpty', 'IsNullOrWhiteSpace', 'Concat', 'Join',
'Trim', 'TrimStart', 'TrimEnd', 'Split', 'Replace', 'StartsWith', 'EndsWith',
'Convert', 'ToInt32', 'ToDouble', 'ToBoolean', 'ToByte',
'Math', 'Abs', 'Ceiling', 'Floor', 'Round', 'Pow', 'Sqrt',
'Dispose', 'Close',
'TryParse', 'Parse',
'AddRange', 'RemoveAt', 'RemoveAll', 'FindAll', 'Exists', 'TrueForAll',
'ContainsKey', 'TryGetValue', 'AddOrUpdate',
'Throw', 'ThrowIfNull',
]);
export const csharpProvider = defineLanguage({
id: SupportedLanguages.CSharp,
extensions: ['.cs'],
@ -27,4 +48,5 @@ export const csharpProvider = defineLanguage({
interfaceNamePattern: /^I[A-Z]/,
mroStrategy: 'implements-split',
fieldExtractor: createFieldExtractor(csharpFieldConfig),
builtInNames: BUILT_INS,
});

View file

@ -5,8 +5,14 @@
* - importSemantics: 'wildcard' (Dart imports bring everything public into scope)
* - exportChecker: public if no leading underscore
* - Dart SDK imports (dart:*) and external packages are skipped
* - enclosingFunctionFinder: Dart's tree-sitter grammar places function_body
* as a sibling of function_signature/method_signature (not as a child).
* The hook resolves the enclosing function by inspecting the previous sibling.
*/
import type { SyntaxNode } from '../utils/ast-helpers.js';
import type { NodeLabel } from '../../graph/types.js';
import { FUNCTION_NODE_TYPES, extractFunctionName } from '../utils/ast-helpers.js';
import { SupportedLanguages } from '../../../config/supported-languages.js';
import { defineLanguage } from '../language-provider.js';
import { typeConfig as dartConfig } from '../type-extractors/dart.js';
@ -16,6 +22,32 @@ import { DART_QUERIES } from '../tree-sitter-queries.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { dartConfig as dartFieldConfig } from '../field-extractors/configs/dart.js';
/**
* Resolve the enclosing function from a `function_body` node by looking at its
* previous sibling. In Dart's tree-sitter grammar, function_signature and
* function_body are siblings under program or class_body, unlike most languages
* where the function declaration wraps both.
*
* Delegates name extraction to the shared `extractFunctionName` which already
* handles Dart's function_signature and method_signature node types.
*/
const dartEnclosingFunctionFinder = (node: SyntaxNode): { funcName: string; label: NodeLabel } | null => {
if (node.type !== 'function_body') return null;
const prev = node.previousSibling;
if (!prev || !FUNCTION_NODE_TYPES.has(prev.type)) return null;
const { funcName, label } = extractFunctionName(prev);
return funcName ? { funcName, label } : null;
};
const BUILT_INS: ReadonlySet<string> = new Set([
'setState', 'mounted', 'debugPrint',
'runApp', 'showDialog', 'showModalBottomSheet',
'Navigator', 'push', 'pushNamed', 'pushReplacement', 'pop', 'maybePop',
'ScaffoldMessenger', 'showSnackBar',
'deactivate', 'reassemble', 'debugDumpApp', 'debugDumpRenderTree',
'then', 'catchError', 'whenComplete', 'listen',
]);
export const dartProvider = defineLanguage({
id: SupportedLanguages.Dart,
extensions: ['.dart'],
@ -25,4 +57,6 @@ export const dartProvider = defineLanguage({
importResolver: resolveDartImport,
importSemantics: 'wildcard',
fieldExtractor: createFieldExtractor(dartFieldConfig),
enclosingFunctionFinder: dartEnclosingFunctionFinder,
builtInNames: BUILT_INS,
});

View file

@ -23,6 +23,7 @@ import { phpProvider } from './php.js';
import { rubyProvider } from './ruby.js';
import { swiftProvider } from './swift.js';
import { dartProvider } from './dart.js';
import { cobolProvider } from './cobol.js';
export const providers = {
[SupportedLanguages.JavaScript]: javascriptProvider,
@ -39,6 +40,7 @@ export const providers = {
[SupportedLanguages.Ruby]: rubyProvider,
[SupportedLanguages.Swift]: swiftProvider,
[SupportedLanguages.Dart]: dartProvider,
[SupportedLanguages.Cobol]: cobolProvider,
} satisfies Record<SupportedLanguages, LanguageProvider>;
/** Get provider by language enum (always succeeds for SupportedLanguages). */

View file

@ -19,6 +19,21 @@ import { isKotlinClassMethod } from '../utils/ast-helpers.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { kotlinConfig } from '../field-extractors/configs/jvm.js';
const BUILT_INS: ReadonlySet<string> = new Set([
'println', 'print', 'readLine', 'require', 'requireNotNull', 'check', 'assert', 'lazy', 'error',
'listOf', 'mapOf', 'setOf', 'mutableListOf', 'mutableMapOf', 'mutableSetOf',
'arrayOf', 'sequenceOf', 'also', 'apply', 'run', 'with', 'takeIf', 'takeUnless',
'TODO', 'buildString', 'buildList', 'buildMap', 'buildSet',
'repeat', 'synchronized',
'launch', 'async', 'runBlocking', 'withContext', 'coroutineScope',
'supervisorScope', 'delay',
'flow', 'flowOf', 'collect', 'emit', 'onEach', 'catch',
'buffer', 'conflate', 'distinctUntilChanged',
'flatMapLatest', 'flatMapMerge', 'combine',
'stateIn', 'shareIn', 'launchIn',
'to', 'until', 'downTo', 'step',
]);
export const kotlinProvider = defineLanguage({
id: SupportedLanguages.Kotlin,
extensions: ['.kt', '.kts'],
@ -30,6 +45,7 @@ export const kotlinProvider = defineLanguage({
importPathPreprocessor: appendKotlinWildcard,
mroStrategy: 'implements-split',
fieldExtractor: createFieldExtractor(kotlinConfig),
builtInNames: BUILT_INS,
labelOverride: (functionNode, defaultLabel) => {
if (defaultLabel !== 'Function') return defaultLabel;
if (isKotlinClassMethod(functionNode)) return 'Method';

View file

@ -18,6 +18,25 @@ import type { NodeLabel } from '../../graph/types.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { phpConfig as phpFieldConfig } from '../field-extractors/configs/php.js';
const BUILT_INS: ReadonlySet<string> = new Set([
'echo', 'isset', 'empty', 'unset', 'list', 'array', 'compact', 'extract',
'count', 'strlen', 'strpos', 'strrpos', 'substr', 'strtolower', 'strtoupper', 'trim',
'ltrim', 'rtrim', 'str_replace', 'str_contains', 'str_starts_with', 'str_ends_with',
'sprintf', 'vsprintf', 'printf', 'number_format',
'array_map', 'array_filter', 'array_reduce', 'array_push', 'array_pop', 'array_shift',
'array_unshift', 'array_slice', 'array_splice', 'array_merge', 'array_keys', 'array_values',
'array_key_exists', 'in_array', 'array_search', 'array_unique', 'usort', 'rsort',
'json_encode', 'json_decode', 'serialize', 'unserialize',
'intval', 'floatval', 'strval', 'boolval', 'is_null', 'is_string', 'is_int', 'is_array',
'is_object', 'is_numeric', 'is_bool', 'is_float',
'var_dump', 'print_r', 'var_export',
'date', 'time', 'strtotime', 'mktime', 'microtime',
'file_exists', 'file_get_contents', 'file_put_contents', 'is_file', 'is_dir',
'preg_match', 'preg_match_all', 'preg_replace', 'preg_split',
'header', 'session_start', 'session_destroy', 'ob_start', 'ob_end_clean', 'ob_get_clean',
'dd', 'dump',
]);
/** Eloquent model properties whose array values are worth indexing. */
const ELOQUENT_ARRAY_PROPS = new Set(['fillable', 'casts', 'hidden', 'guarded', 'with', 'appends']);
@ -133,4 +152,5 @@ export const phpProvider = defineLanguage({
fieldExtractor: createFieldExtractor(phpFieldConfig),
descriptionExtractor: phpDescriptionExtractor,
isRouteFile: isPhpRouteFile,
builtInNames: BUILT_INS,
});

View file

@ -20,6 +20,13 @@ import { PYTHON_QUERIES } from '../tree-sitter-queries.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { pythonConfig as pythonFieldConfig } from '../field-extractors/configs/python.js';
const BUILT_INS: ReadonlySet<string> = new Set([
'print', 'len', 'range', 'str', 'int', 'float', 'list', 'dict', 'set', 'tuple',
'append', 'extend', 'update',
'type', 'isinstance', 'issubclass', 'getattr', 'setattr', 'hasattr',
'enumerate', 'zip', 'sorted', 'reversed', 'min', 'max', 'sum', 'abs',
]);
export const pythonProvider = defineLanguage({
id: SupportedLanguages.Python,
extensions: ['.py'],
@ -31,4 +38,5 @@ export const pythonProvider = defineLanguage({
importSemantics: 'namespace',
mroStrategy: 'c3',
fieldExtractor: createFieldExtractor(pythonFieldConfig),
builtInNames: BUILT_INS,
});

View file

@ -17,6 +17,22 @@ import { RUBY_QUERIES } from '../tree-sitter-queries.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { rubyConfig as rubyFieldConfig } from '../field-extractors/configs/ruby.js';
const BUILT_INS: ReadonlySet<string> = new Set([
'puts', 'p', 'pp', 'raise', 'fail',
'require', 'require_relative', 'load', 'autoload',
'include', 'extend', 'prepend',
'attr_accessor', 'attr_reader', 'attr_writer',
'public', 'private', 'protected', 'module_function',
'lambda', 'proc', 'block_given?',
'nil?', 'is_a?', 'kind_of?', 'instance_of?', 'respond_to?',
'freeze', 'frozen?', 'dup', 'tap', 'yield_self',
'each', 'select', 'reject', 'detect', 'collect',
'inject', 'flat_map', 'each_with_object', 'each_with_index',
'any?', 'all?', 'none?', 'count', 'first', 'last',
'sort_by', 'min_by', 'max_by',
'group_by', 'partition', 'compact', 'flatten', 'uniq',
]);
export const rubyProvider = defineLanguage({
id: SupportedLanguages.Ruby,
extensions: ['.rb', '.rake', '.gemspec'],
@ -27,4 +43,5 @@ export const rubyProvider = defineLanguage({
callRouter: routeRubyCall,
importSemantics: 'wildcard',
fieldExtractor: createFieldExtractor(rubyFieldConfig),
builtInNames: BUILT_INS,
});

View file

@ -20,6 +20,19 @@ import { RUST_QUERIES } from '../tree-sitter-queries.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { rustConfig as rustFieldConfig } from '../field-extractors/configs/rust.js';
const BUILT_INS: ReadonlySet<string> = new Set([
'unwrap', 'expect', 'unwrap_or', 'unwrap_or_else', 'unwrap_or_default',
'ok', 'err', 'is_ok', 'is_err', 'map', 'map_err', 'and_then', 'or_else',
'clone', 'to_string', 'to_owned', 'into', 'from', 'as_ref', 'as_mut',
'iter', 'into_iter', 'collect', 'filter', 'fold', 'for_each',
'len', 'is_empty', 'push', 'pop', 'insert', 'remove', 'contains',
'format', 'write', 'writeln', 'panic', 'unreachable', 'todo', 'unimplemented',
'vec', 'println', 'eprintln', 'dbg',
'lock', 'read', 'try_lock',
'spawn', 'join', 'sleep',
'Some', 'None', 'Ok', 'Err',
]);
export const rustProvider = defineLanguage({
id: SupportedLanguages.Rust,
extensions: ['.rs'],
@ -30,4 +43,5 @@ export const rustProvider = defineLanguage({
namedBindingExtractor: extractRustNamedBindings,
mroStrategy: 'qualified-syntax',
fieldExtractor: createFieldExtractor(rustFieldConfig),
builtInNames: BUILT_INS,
});

View file

@ -103,6 +103,34 @@ function wireSwiftImplicitImports(
}
}
const BUILT_INS: ReadonlySet<string> = new Set([
'print', 'debugPrint', 'dump', 'fatalError', 'precondition', 'preconditionFailure',
'assert', 'assertionFailure', 'NSLog',
'abs', 'min', 'max', 'zip', 'stride', 'sequence', 'repeatElement',
'swap', 'withUnsafePointer', 'withUnsafeMutablePointer', 'withUnsafeBytes',
'autoreleasepool', 'unsafeBitCast', 'unsafeDowncast', 'numericCast',
'type', 'MemoryLayout',
'map', 'flatMap', 'compactMap', 'filter', 'reduce', 'forEach', 'contains',
'first', 'last', 'prefix', 'suffix', 'dropFirst', 'dropLast',
'sorted', 'reversed', 'enumerated', 'joined', 'split',
'append', 'insert', 'remove', 'removeAll', 'removeFirst', 'removeLast',
'isEmpty', 'count', 'index', 'startIndex', 'endIndex',
'addSubview', 'removeFromSuperview', 'layoutSubviews', 'setNeedsLayout',
'layoutIfNeeded', 'setNeedsDisplay', 'invalidateIntrinsicContentSize',
'addTarget', 'removeTarget', 'addGestureRecognizer',
'addConstraint', 'addConstraints', 'removeConstraint', 'removeConstraints',
'NSLocalizedString', 'Bundle',
'reloadData', 'reloadSections', 'reloadRows', 'performBatchUpdates',
'register', 'dequeueReusableCell', 'dequeueReusableSupplementaryView',
'beginUpdates', 'endUpdates', 'insertRows', 'deleteRows', 'insertSections', 'deleteSections',
'present', 'dismiss', 'pushViewController', 'popViewController', 'popToRootViewController',
'performSegue', 'prepare',
'DispatchQueue', 'async', 'sync', 'asyncAfter',
'Task', 'withCheckedContinuation', 'withCheckedThrowingContinuation',
'sink', 'store', 'assign', 'receive', 'subscribe',
'addObserver', 'removeObserver', 'post', 'NotificationCenter',
]);
export const swiftProvider = defineLanguage({
id: SupportedLanguages.Swift,
extensions: ['.swift'],
@ -114,4 +142,5 @@ export const swiftProvider = defineLanguage({
heritageDefaultEdge: 'IMPLEMENTS',
fieldExtractor: createFieldExtractor(swiftFieldConfig),
implicitImportWirer: wireSwiftImplicitImports,
builtInNames: BUILT_INS,
});

View file

@ -18,6 +18,27 @@ import { typescriptFieldExtractor } from '../field-extractors/typescript.js';
import { createFieldExtractor } from '../field-extractors/generic.js';
import { javascriptConfig } from '../field-extractors/configs/typescript-javascript.js';
const BUILT_INS: ReadonlySet<string> = new Set([
'console', 'log', 'warn', 'error', 'info', 'debug',
'setTimeout', 'setInterval', 'clearTimeout', 'clearInterval',
'parseInt', 'parseFloat', 'isNaN', 'isFinite',
'encodeURI', 'decodeURI', 'encodeURIComponent', 'decodeURIComponent',
'JSON', 'parse', 'stringify',
'Object', 'Array', 'String', 'Number', 'Boolean', 'Symbol', 'BigInt',
'Map', 'Set', 'WeakMap', 'WeakSet',
'Promise', 'resolve', 'reject', 'then', 'catch', 'finally',
'Math', 'Date', 'RegExp', 'Error',
'require', 'import', 'export', 'fetch', 'Response', 'Request',
'useState', 'useEffect', 'useCallback', 'useMemo', 'useRef', 'useContext',
'useReducer', 'useLayoutEffect', 'useImperativeHandle', 'useDebugValue',
'createElement', 'createContext', 'createRef', 'forwardRef', 'memo', 'lazy',
'map', 'filter', 'reduce', 'forEach', 'find', 'findIndex', 'some', 'every',
'includes', 'indexOf', 'slice', 'splice', 'concat', 'join', 'split',
'push', 'pop', 'shift', 'unshift', 'sort', 'reverse',
'keys', 'values', 'entries', 'assign', 'freeze', 'seal',
'hasOwnProperty', 'toString', 'valueOf',
]);
export const typescriptProvider = defineLanguage({
id: SupportedLanguages.TypeScript,
extensions: ['.ts', '.tsx'],
@ -27,6 +48,7 @@ export const typescriptProvider = defineLanguage({
importResolver: resolveTypescriptImport,
namedBindingExtractor: extractTsNamedBindings,
fieldExtractor: typescriptFieldExtractor,
builtInNames: BUILT_INS,
});
export const javascriptProvider = defineLanguage({
@ -38,4 +60,5 @@ export const javascriptProvider = defineLanguage({
importResolver: resolveJavascriptImport,
namedBindingExtractor: extractTsNamedBindings,
fieldExtractor: createFieldExtractor(javascriptConfig),
builtInNames: BUILT_INS,
});

View file

@ -1,6 +1,7 @@
import { createKnowledgeGraph } from '../graph/graph.js';
import { processStructure } from './structure-processor.js';
import { processMarkdown } from './markdown-processor.js';
import { processCobol, isCobolFile, isJclFile } from './cobol-processor.js';
import { processParsing } from './parsing-processor.js';
import {
processImports,
@ -464,6 +465,14 @@ async function runScanAndStructure(
stats: { filesProcessed: totalFiles, totalFiles, nodesCreated: graph.nodeCount },
});
// ── Custom (non-tree-sitter) processors ─────────────────────────────
// Each custom processor follows the pattern in markdown-processor.ts:
// 1. Export a process function: (graph, files, allPathSet) => result
// 2. Export a file detection function: (path) => boolean
// 3. Filter files by extension, write nodes/edges directly to graph
// To add a new language: create a new processor file, import it here,
// and add a filter-read-call-log block following the pattern below.
// ── Phase 2.5: Markdown processing (headings + cross-links) ────────
const mdScanned = scannedFiles.filter(f => f.path.endsWith('.md') || f.path.endsWith('.mdx'));
if (mdScanned.length > 0) {
@ -478,6 +487,26 @@ async function runScanAndStructure(
}
}
// ── Phase 2.6: COBOL processing (regex extraction, no tree-sitter) ──
const cobolScanned = scannedFiles.filter(f => isCobolFile(f.path) || isJclFile(f.path));
if (cobolScanned.length > 0) {
const cobolContents = await readFileContents(repoPath, cobolScanned.map(f => f.path));
const cobolFiles = cobolScanned
.filter(f => cobolContents.has(f.path))
.map(f => ({ path: f.path, content: cobolContents.get(f.path)! }));
const allPathSet = new Set(allPaths);
const cobolResult = processCobol(graph, cobolFiles, allPathSet);
if (isDev) {
console.log(` COBOL: ${cobolResult.programs} programs, ${cobolResult.paragraphs} paragraphs, ${cobolResult.sections} sections from ${cobolFiles.length} files`);
if (cobolResult.execSqlBlocks > 0 || cobolResult.execCicsBlocks > 0 || cobolResult.entryPoints > 0) {
console.log(` COBOL enriched: ${cobolResult.execSqlBlocks} SQL blocks, ${cobolResult.execCicsBlocks} CICS blocks, ${cobolResult.entryPoints} entry points, ${cobolResult.moves} moves, ${cobolResult.fileDeclarations} file declarations`);
}
if (cobolResult.jclJobs > 0) {
console.log(` JCL: ${cobolResult.jclJobs} jobs, ${cobolResult.jclSteps} steps`);
}
}
}
return { scannedFiles, allPaths, totalFiles };
}

View file

@ -808,7 +808,7 @@ export const RUBY_QUERIES = `
; NOTE: This may over-capture variable reads as calls (e.g. 'result' at
; statement level). Ruby's grammar makes bare identifiers ambiguous — they
; could be local variables or zero-arity method calls. Post-processing via
; isBuiltInOrNoise and symbol resolution filtering suppresses most false
; provider.isBuiltInName and symbol resolution filtering suppresses most false
; positives, but a variable name that coincidentally matches a method name
; elsewhere may produce a false CALLS edge.
(body_statement
@ -1054,6 +1054,20 @@ export const DART_QUERIES = `
(factory_constructor_signature
(identifier) @name . (formal_parameter_list))) @definition.constructor
; ── Field declarations (String name = '', Address address = Address()) ──────
(declaration
(type_identifier)
(initialized_identifier_list
(initialized_identifier
(identifier) @name))) @definition.property
; ── Nullable field declarations (String? name) ──────────────────────────────
(declaration
(nullable_type)
(initialized_identifier_list
(initialized_identifier
(identifier) @name))) @definition.property
; ── Getters ──────────────────────────────────────────────────────────────────
(method_signature
(getter_signature
@ -1097,6 +1111,22 @@ export const DART_QUERIES = `
(library_export
(configurable_uri) @import.source)) @import
; ── Write access: obj.field = value ──────────────────────────────────────────
(assignment_expression
left: (assignable_expression
(identifier) @assignment.receiver
(unconditional_assignable_selector
(identifier) @assignment.property))
right: (_)) @assignment
; ── Write access: this.field = value ─────────────────────────────────────────
(assignment_expression
left: (assignable_expression
(this) @assignment.receiver
(unconditional_assignable_selector
(identifier) @assignment.property))
right: (_)) @assignment
; ── Heritage: extends ────────────────────────────────────────────────────────
(class_definition
name: (identifier) @heritage.class
@ -1134,4 +1164,5 @@ export const LANGUAGE_QUERIES: Record<SupportedLanguages, string> = {
[SupportedLanguages.Ruby]: RUBY_QUERIES,
[SupportedLanguages.Swift]: SWIFT_QUERIES,
[SupportedLanguages.Dart]: DART_QUERIES,
[SupportedLanguages.Cobol]: '', // Standalone regex processor — no tree-sitter queries
};

View file

@ -1,6 +1,5 @@
import { type SyntaxNode, FUNCTION_NODE_TYPES, extractFunctionName, CLASS_CONTAINER_TYPES } from './utils/ast-helpers.js';
import { CALL_EXPRESSION_TYPES } from './utils/call-analysis.js';
import { isBuiltInOrNoise } from './utils/noise-filter.js';
import { SupportedLanguages } from '../../config/supported-languages.js';
import { TYPED_PARAMETER_TYPES } from './type-extractors/shared.js';
import { getProvider } from './languages/index.js';
@ -732,7 +731,7 @@ export const buildTypeEnv = (
lookupReturnType(callee: string): string | undefined {
// SymbolTable is authoritative when it has an unambiguous match
if (symbolTable) {
if (isBuiltInOrNoise(callee)) return undefined;
if (provider.isBuiltInName(callee)) return undefined;
const callables = symbolTable.lookupFuzzyCallable(callee);
if (callables.length === 1) {
const rawReturn = callables[0].returnType;
@ -746,7 +745,7 @@ export const buildTypeEnv = (
},
lookupRawReturnType(callee: string): string | undefined {
if (symbolTable) {
if (isBuiltInOrNoise(callee)) return undefined;
if (provider.isBuiltInName(callee)) return undefined;
const callables = symbolTable.lookupFuzzyCallable(callee);
if (callables.length === 1) return callables[0].returnType;
// Ambiguous (2+) → return undefined (conservative, no cross-file fallback)

View file

@ -111,6 +111,25 @@ function hasDartTypeAnnotation(node: SyntaxNode): boolean {
// ── Tier 0: Explicit Type Annotations ───────────────────────────────────
const extractDartDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
// initialized_identifier: comma-separated variable (String a, b, c) — type is on parent
if (node.type === 'initialized_identifier') {
const parent = node.parent;
if (!parent) return;
let typeNode = findChild(parent, 'type_identifier');
if (!typeNode) {
const nullable = findChild(parent, 'nullable_type');
if (nullable) typeNode = findChild(nullable, 'type_identifier');
}
if (!typeNode) return;
const typeName = extractSimpleTypeName(typeNode);
if (!typeName || typeName === 'dynamic') return;
const nameNode = findChild(node, 'identifier');
if (!nameNode) return;
const varName = extractVarName(nameNode);
if (varName) env.set(varName, typeName);
return;
}
let typeNode = findChild(node, 'type_identifier');
if (!typeNode) {
const nullable = findChild(node, 'nullable_type');

View file

@ -82,6 +82,9 @@ export const FUNCTION_NODE_TYPES = new Set([
// Ruby
'method', // def foo
'singleton_method', // def self.foo
// Dart
'function_signature',
'method_signature',
]);
/**
@ -430,6 +433,34 @@ export const extractFunctionName = (node: SyntaxNode): { funcName: string | null
}
funcName = nameNode?.text;
label = 'Method';
} else if (node.type === 'function_signature') {
// Dart: top-level function signatures
let nameNode = node.childForFieldName?.('name');
if (!nameNode) {
for (let i = 0; i < node.childCount; i++) {
const c = node.child(i);
if (c?.type === 'identifier') { nameNode = c; break; }
}
}
funcName = nameNode?.text ?? null;
} else if (node.type === 'method_signature') {
// Dart: method_signature wraps function_signature
let funcSig: SyntaxNode | null = null;
for (let i = 0; i < node.childCount; i++) {
const c = node.child(i);
if (c?.type === 'function_signature') { funcSig = c; break; }
}
if (funcSig) {
let nameNode = funcSig.childForFieldName?.('name');
if (!nameNode) {
for (let i = 0; i < funcSig.childCount; i++) {
const c = funcSig.child(i);
if (c?.type === 'identifier') { nameNode = c; break; }
}
}
funcName = nameNode?.text ?? null;
}
label = 'Method';
}
return { funcName, label };
@ -471,6 +502,7 @@ export const extractMethodSignature = (node: SyntaxNode | null | undefined): Met
const paramListTypes = new Set([
'formal_parameters', 'parameters', 'parameter_list',
'function_parameters', 'method_parameters', 'function_value_parameters',
'formal_parameter_list', // Dart
]);
// Node types that indicate variadic/rest parameters

View file

@ -64,6 +64,7 @@ const MEMBER_ACCESS_NODE_TYPES = new Set([
'selector_expression', // Go: obj.Method()
'navigation_suffix', // Kotlin/Swift: obj.method() — nameNode sits inside navigation_suffix
'member_binding_expression', // C#: user?.Method() — null-conditional access
'unconditional_assignable_selector', // Dart: obj.method() — nameNode inside selector > unconditional_assignable_selector
]);
/**
@ -208,6 +209,16 @@ export const extractReceiverName = (
}
}
// Dart: unconditional_assignable_selector is inside a `selector`, which is a sibling
// of the receiver in the expression_statement. For `user.save()`, the previous named
// sibling of the `selector` is `identifier [user]`.
if (!receiver && parent.type === 'unconditional_assignable_selector') {
const selectorNode = parent.parent; // selector [.save]
if (selectorNode) {
receiver = selectorNode.previousNamedSibling;
}
}
// C# null-conditional: user?.Save() → conditional_access_expression wraps member_binding_expression
if (!receiver && parent.type === 'member_binding_expression') {
const condAccess = parent.parent;
@ -291,6 +302,14 @@ export const extractReceiverNode = (
}
}
// Dart: unconditional_assignable_selector — receiver is previous sibling of the selector
if (!receiver && parent.type === 'unconditional_assignable_selector') {
const selectorNode = parent.parent;
if (selectorNode) {
receiver = selectorNode.previousNamedSibling;
}
}
if (!receiver && parent.type === 'member_binding_expression') {
const condAccess = parent.parent;
if (condAccess?.type === 'conditional_access_expression') {
@ -527,6 +546,27 @@ export function extractMixedChain(
} else {
return { chain, baseReceiverName: innerObject.text || undefined };
}
} else if (current.type === 'selector') {
// ── Dart: flat selector siblings (user.address.save() uses selector nodes) ──
// Extract field name from unconditional_assignable_selector child
const uas = current.namedChildren?.find(
(c: SyntaxNode) => c.type === 'unconditional_assignable_selector',
);
const propertyName = uas?.namedChildren?.find(
(c: SyntaxNode) => c.type === 'identifier',
)?.text;
if (!propertyName) break;
chain.unshift({ kind: 'field', name: propertyName });
// Walk to previous sibling for the next step in the chain
const prev = current.previousNamedSibling;
if (!prev) break;
if (prev.type === 'selector') {
current = prev;
} else {
// Base receiver (identifier or other terminal)
return { chain, baseReceiverName: prev.text || undefined };
}
} else {
// Simple identifier — this is the base receiver
return chain.length > 0

View file

@ -1,175 +0,0 @@
/**
* Built-in name filtering — identifies standard library functions and common noise
* that should not be tracked as call targets in the knowledge graph.
*
* Covers: JS/TS, Python, Kotlin, C/C++, C#, PHP, Swift, Rust, Ruby, Dart/Flutter standard libraries.
*/
export const BUILT_IN_NAMES = new Set([
// JavaScript/TypeScript
'console', 'log', 'warn', 'error', 'info', 'debug',
'setTimeout', 'setInterval', 'clearTimeout', 'clearInterval',
'parseInt', 'parseFloat', 'isNaN', 'isFinite',
'encodeURI', 'decodeURI', 'encodeURIComponent', 'decodeURIComponent',
'JSON', 'parse', 'stringify',
'Object', 'Array', 'String', 'Number', 'Boolean', 'Symbol', 'BigInt',
'Map', 'Set', 'WeakMap', 'WeakSet',
'Promise', 'resolve', 'reject', 'then', 'catch', 'finally',
'Math', 'Date', 'RegExp', 'Error',
'require', 'import', 'export', 'fetch', 'Response', 'Request',
'useState', 'useEffect', 'useCallback', 'useMemo', 'useRef', 'useContext',
'useReducer', 'useLayoutEffect', 'useImperativeHandle', 'useDebugValue',
'createElement', 'createContext', 'createRef', 'forwardRef', 'memo', 'lazy',
'map', 'filter', 'reduce', 'forEach', 'find', 'findIndex', 'some', 'every',
'includes', 'indexOf', 'slice', 'splice', 'concat', 'join', 'split',
'push', 'pop', 'shift', 'unshift', 'sort', 'reverse',
'keys', 'values', 'entries', 'assign', 'freeze', 'seal',
'hasOwnProperty', 'toString', 'valueOf',
// Python
'print', 'len', 'range', 'str', 'int', 'float', 'list', 'dict', 'set', 'tuple',
'append', 'extend', 'update',
// NOTE: 'open', 'read', 'write', 'close' removed — these are real C POSIX syscalls
'type', 'isinstance', 'issubclass', 'getattr', 'setattr', 'hasattr',
'enumerate', 'zip', 'sorted', 'reversed', 'min', 'max', 'sum', 'abs',
// Kotlin stdlib
'println', 'print', 'readLine', 'require', 'requireNotNull', 'check', 'assert', 'lazy', 'error',
'listOf', 'mapOf', 'setOf', 'mutableListOf', 'mutableMapOf', 'mutableSetOf',
'arrayOf', 'sequenceOf', 'also', 'apply', 'run', 'with', 'takeIf', 'takeUnless',
'TODO', 'buildString', 'buildList', 'buildMap', 'buildSet',
'repeat', 'synchronized',
// Kotlin coroutine builders & scope functions
'launch', 'async', 'runBlocking', 'withContext', 'coroutineScope',
'supervisorScope', 'delay',
// Kotlin Flow operators
'flow', 'flowOf', 'collect', 'emit', 'onEach', 'catch',
'buffer', 'conflate', 'distinctUntilChanged',
'flatMapLatest', 'flatMapMerge', 'combine',
'stateIn', 'shareIn', 'launchIn',
// Kotlin infix stdlib functions
'to', 'until', 'downTo', 'step',
// C/C++ standard library
'printf', 'fprintf', 'sprintf', 'snprintf', 'vprintf', 'vfprintf', 'vsprintf', 'vsnprintf',
'scanf', 'fscanf', 'sscanf',
'malloc', 'calloc', 'realloc', 'free', 'memcpy', 'memmove', 'memset', 'memcmp',
'strlen', 'strcpy', 'strncpy', 'strcat', 'strncat', 'strcmp', 'strncmp', 'strstr', 'strchr', 'strrchr',
'atoi', 'atol', 'atof', 'strtol', 'strtoul', 'strtoll', 'strtoull', 'strtod',
'sizeof', 'offsetof', 'typeof',
'assert', 'abort', 'exit', '_exit',
'fopen', 'fclose', 'fread', 'fwrite', 'fseek', 'ftell', 'rewind', 'fflush', 'fgets', 'fputs',
// Linux kernel common macros/helpers (not real call targets)
'likely', 'unlikely', 'BUG', 'BUG_ON', 'WARN', 'WARN_ON', 'WARN_ONCE',
'IS_ERR', 'PTR_ERR', 'ERR_PTR', 'IS_ERR_OR_NULL',
'ARRAY_SIZE', 'container_of', 'list_for_each_entry', 'list_for_each_entry_safe',
'min', 'max', 'clamp', 'abs', 'swap',
'pr_info', 'pr_warn', 'pr_err', 'pr_debug', 'pr_notice', 'pr_crit', 'pr_emerg',
'printk', 'dev_info', 'dev_warn', 'dev_err', 'dev_dbg',
'GFP_KERNEL', 'GFP_ATOMIC',
'spin_lock', 'spin_unlock', 'spin_lock_irqsave', 'spin_unlock_irqrestore',
'mutex_lock', 'mutex_unlock', 'mutex_init',
'kfree', 'kmalloc', 'kzalloc', 'kcalloc', 'krealloc', 'kvmalloc', 'kvfree',
'get', 'put',
// C# / .NET built-ins
'Console', 'WriteLine', 'ReadLine', 'Write',
'Task', 'Run', 'Wait', 'WhenAll', 'WhenAny', 'FromResult', 'Delay', 'ContinueWith',
'ConfigureAwait', 'GetAwaiter', 'GetResult',
'ToString', 'GetType', 'Equals', 'GetHashCode', 'ReferenceEquals',
'Add', 'Remove', 'Contains', 'Clear', 'Count', 'Any', 'All',
'Where', 'Select', 'SelectMany', 'OrderBy', 'OrderByDescending', 'GroupBy',
'First', 'FirstOrDefault', 'Single', 'SingleOrDefault', 'Last', 'LastOrDefault',
'ToList', 'ToArray', 'ToDictionary', 'AsEnumerable', 'AsQueryable',
'Aggregate', 'Sum', 'Average', 'Min', 'Max', 'Distinct', 'Skip', 'Take',
'String', 'Format', 'IsNullOrEmpty', 'IsNullOrWhiteSpace', 'Concat', 'Join',
'Trim', 'TrimStart', 'TrimEnd', 'Split', 'Replace', 'StartsWith', 'EndsWith',
'Convert', 'ToInt32', 'ToDouble', 'ToBoolean', 'ToByte',
'Math', 'Abs', 'Ceiling', 'Floor', 'Round', 'Pow', 'Sqrt',
'Dispose', 'Close',
'TryParse', 'Parse',
'AddRange', 'RemoveAt', 'RemoveAll', 'FindAll', 'Exists', 'TrueForAll',
'ContainsKey', 'TryGetValue', 'AddOrUpdate',
'Throw', 'ThrowIfNull',
// PHP built-ins
'echo', 'isset', 'empty', 'unset', 'list', 'array', 'compact', 'extract',
'count', 'strlen', 'strpos', 'strrpos', 'substr', 'strtolower', 'strtoupper', 'trim',
'ltrim', 'rtrim', 'str_replace', 'str_contains', 'str_starts_with', 'str_ends_with',
'sprintf', 'vsprintf', 'printf', 'number_format',
'array_map', 'array_filter', 'array_reduce', 'array_push', 'array_pop', 'array_shift',
'array_unshift', 'array_slice', 'array_splice', 'array_merge', 'array_keys', 'array_values',
'array_key_exists', 'in_array', 'array_search', 'array_unique', 'usort', 'rsort',
'json_encode', 'json_decode', 'serialize', 'unserialize',
'intval', 'floatval', 'strval', 'boolval', 'is_null', 'is_string', 'is_int', 'is_array',
'is_object', 'is_numeric', 'is_bool', 'is_float',
'var_dump', 'print_r', 'var_export',
'date', 'time', 'strtotime', 'mktime', 'microtime',
'file_exists', 'file_get_contents', 'file_put_contents', 'is_file', 'is_dir',
'preg_match', 'preg_match_all', 'preg_replace', 'preg_split',
'header', 'session_start', 'session_destroy', 'ob_start', 'ob_end_clean', 'ob_get_clean',
'dd', 'dump',
// Swift/iOS built-ins and standard library
'print', 'debugPrint', 'dump', 'fatalError', 'precondition', 'preconditionFailure',
'assert', 'assertionFailure', 'NSLog',
'abs', 'min', 'max', 'zip', 'stride', 'sequence', 'repeatElement',
'swap', 'withUnsafePointer', 'withUnsafeMutablePointer', 'withUnsafeBytes',
'autoreleasepool', 'unsafeBitCast', 'unsafeDowncast', 'numericCast',
'type', 'MemoryLayout',
// Swift collection/string methods (common noise)
'map', 'flatMap', 'compactMap', 'filter', 'reduce', 'forEach', 'contains',
'first', 'last', 'prefix', 'suffix', 'dropFirst', 'dropLast',
'sorted', 'reversed', 'enumerated', 'joined', 'split',
'append', 'insert', 'remove', 'removeAll', 'removeFirst', 'removeLast',
'isEmpty', 'count', 'index', 'startIndex', 'endIndex',
// UIKit/Foundation common methods (noise in call graph)
'addSubview', 'removeFromSuperview', 'layoutSubviews', 'setNeedsLayout',
'layoutIfNeeded', 'setNeedsDisplay', 'invalidateIntrinsicContentSize',
'addTarget', 'removeTarget', 'addGestureRecognizer',
'addConstraint', 'addConstraints', 'removeConstraint', 'removeConstraints',
'NSLocalizedString', 'Bundle',
'reloadData', 'reloadSections', 'reloadRows', 'performBatchUpdates',
'register', 'dequeueReusableCell', 'dequeueReusableSupplementaryView',
'beginUpdates', 'endUpdates', 'insertRows', 'deleteRows', 'insertSections', 'deleteSections',
'present', 'dismiss', 'pushViewController', 'popViewController', 'popToRootViewController',
'performSegue', 'prepare',
// GCD / async
'DispatchQueue', 'async', 'sync', 'asyncAfter',
'Task', 'withCheckedContinuation', 'withCheckedThrowingContinuation',
// Combine
'sink', 'store', 'assign', 'receive', 'subscribe',
// Notification / KVO
'addObserver', 'removeObserver', 'post', 'NotificationCenter',
// Rust standard library (common noise in call graphs)
'unwrap', 'expect', 'unwrap_or', 'unwrap_or_else', 'unwrap_or_default',
'ok', 'err', 'is_ok', 'is_err', 'map', 'map_err', 'and_then', 'or_else',
'clone', 'to_string', 'to_owned', 'into', 'from', 'as_ref', 'as_mut',
'iter', 'into_iter', 'collect', 'map', 'filter', 'fold', 'for_each',
'len', 'is_empty', 'push', 'pop', 'insert', 'remove', 'contains',
'format', 'write', 'writeln', 'panic', 'unreachable', 'todo', 'unimplemented',
'vec', 'println', 'eprintln', 'dbg',
'lock', 'read', 'write', 'try_lock',
'spawn', 'join', 'sleep',
'Some', 'None', 'Ok', 'Err',
// Ruby built-ins and Kernel methods
'puts', 'p', 'pp', 'raise', 'fail',
'require', 'require_relative', 'load', 'autoload',
'include', 'extend', 'prepend',
'attr_accessor', 'attr_reader', 'attr_writer',
'public', 'private', 'protected', 'module_function',
'lambda', 'proc', 'block_given?',
'nil?', 'is_a?', 'kind_of?', 'instance_of?', 'respond_to?',
'freeze', 'frozen?', 'dup', 'tap', 'yield_self',
// Dart / Flutter
'setState', 'mounted', 'debugPrint',
'runApp', 'showDialog', 'showModalBottomSheet',
'Navigator', 'push', 'pushNamed', 'pushReplacement', 'pop', 'maybePop',
'ScaffoldMessenger', 'showSnackBar',
'deactivate', 'reassemble', 'debugDumpApp', 'debugDumpRenderTree',
// Dart async
'then', 'catchError', 'whenComplete', 'listen',
// Ruby enumerables
'each', 'select', 'reject', 'detect', 'collect',
'inject', 'flat_map', 'each_with_object', 'each_with_index',
'any?', 'all?', 'none?', 'count', 'first', 'last',
'sort_by', 'min_by', 'max_by',
'group_by', 'partition', 'compact', 'flatten', 'uniq',
]);
/** Check if a name is a built-in function or common noise that should be filtered out */
export const isBuiltInOrNoise = (name: string): boolean => BUILT_IN_NAMES.has(name);

View file

@ -29,7 +29,6 @@ try { Dart = _require('tree-sitter-dart'); } catch {}
let Kotlin: any = null;
try { Kotlin = _require('tree-sitter-kotlin'); } catch {}
import { getLanguageFromFilename } from '../utils/language-detection.js';
import { isBuiltInOrNoise } from '../utils/noise-filter.js';
import {
FUNCTION_NODE_TYPES,
extractFunctionName,
@ -388,6 +387,23 @@ const findEnclosingFunctionId = (node: any, filePath: string, provider: Language
return result;
}
}
// Language-specific enclosing function resolution (e.g., Dart where
// function_body is a sibling of function_signature, not a child).
if (provider.enclosingFunctionFinder) {
const customResult = provider.enclosingFunctionFinder(current);
if (customResult) {
let finalLabel: NodeLabel = customResult.label;
if (provider.labelOverride) {
const override = provider.labelOverride(current.previousSibling, finalLabel);
if (override !== null) finalLabel = override;
}
const result = generateId(finalLabel, `${filePath}:${customResult.funcName}`);
functionIdCache.set(node, result);
return result;
}
}
current = current.parent;
}
functionIdCache.set(node, null);
@ -1268,7 +1284,7 @@ const processFileGroup = (
// kind === 'call' — fall through to normal call processing below
}
if (!isBuiltInOrNoise(calledName)) {
if (!provider.isBuiltInName(calledName)) {
const callNode = captureMap['call'];
const sourceId = findEnclosingFunctionId(callNode, file.path, provider)
|| generateId('File', file.path);

View file

@ -551,9 +551,10 @@ export const closeLbug = async (repoId?: string): Promise<void> => {
export const isLbugReady = (repoId: string): boolean => pool.has(repoId);
/** Regex to detect write operations in user-supplied Cypher queries */
export const CYPHER_WRITE_RE = /\b(CREATE|DELETE|SET|MERGE|REMOVE|DROP|ALTER|COPY|DETACH)\b/i;
export const CYPHER_WRITE_RE = /(?<!:)\b(CREATE|DELETE|SET|MERGE|REMOVE|DROP|ALTER|COPY|DETACH|FOREACH)\b/i;
/** Check if a Cypher query contains write operations */
export function isWriteQuery(query: string): boolean {
return CYPHER_WRITE_RE.test(query);
}

View file

@ -8,7 +8,8 @@
import fs from 'fs/promises';
import path from 'path';
import { initLbug, executeQuery, executeParameterized, closeLbug, isLbugReady } from '../core/lbug-adapter.js';
import { initLbug, executeQuery, executeParameterized, closeLbug, isLbugReady, isWriteQuery } from '../core/lbug-adapter.js';
export { isWriteQuery };
// Embedding imports are lazy (dynamic import) to avoid loading onnxruntime-node
// at MCP server startup — crashes on unsupported Node ABI versions (#89)
// git utilities available if needed
@ -89,13 +90,6 @@ export const IMPACT_RELATION_CONFIDENCE: Readonly<Record<string, number>> = {
const confidenceForRelType = (relType: string | undefined): number =>
IMPACT_RELATION_CONFIDENCE[relType ?? ''] ?? 0.5;
/** Regex to detect write operations in user-supplied Cypher queries */
export const CYPHER_WRITE_RE = /\b(CREATE|DELETE|SET|MERGE|REMOVE|DROP|ALTER|COPY|DETACH)\b/i;
/** Check if a Cypher query contains write operations */
export function isWriteQuery(query: string): boolean {
return CYPHER_WRITE_RE.test(query);
}
/** Structured error logging for query failures — replaces empty catch blocks */
function logQueryError(context: string, err: unknown): void {
@ -777,7 +771,7 @@ export class LocalBackend {
}
// Block write operations (defense-in-depth — DB is already read-only)
if (CYPHER_WRITE_RE.test(params.query)) {
if (isWriteQuery(params.query)) {
return { error: 'Write operations (CREATE, DELETE, SET, MERGE, REMOVE, DROP, ALTER, COPY, DETACH) are not allowed. The knowledge graph is read-only.' };
}
@ -1722,69 +1716,219 @@ export class LocalBackend {
let affectedModules: any[] = [];
if (impacted.length > 0) {
// Cap IN-clause to 100 IDs to prevent oversized queries that crash
// the native DB engine on arm64 macOS (#292)
const cappedImpacted = impacted.slice(0, 100);
const allIds = cappedImpacted.map(i => `'${String(i.id ?? '').replace(/'/g, "''")}'`).join(', ');
const d1Items = (grouped[1] || []).slice(0, 100);
const d1Ids = d1Items.map((i: any) => `'${String(i.id ?? '').replace(/'/g, "''")}'`).join(', ');
const CHUNK_SIZE = 100;
// Max number of chunks to process to avoid unbounded DB round-trips.
// Configurable via env IMPACT_MAX_CHUNKS, default 10 => max items = 1000
const MAX_CHUNKS = parseInt(process.env.IMPACT_MAX_CHUNKS || '10', 10);
// Enrichment queries: sequential on arm64 macOS to avoid SIGSEGV from
// concurrent native DB access (#285, #290, #292); parallel elsewhere
// to preserve performance on unaffected platforms.
const isArm64Mac = process.platform === 'darwin' && process.arch === 'arm64';
// ── Process enrichment: batched chunking (bounded by MAX_CHUNKS) ─
// Uses merged Cypher query (WITH + OPTIONAL MATCH) to fetch
// process + entry point info in 1 round-trip per chunk. Converted to
// parameterized queries to avoid manual string escaping and long query strings.
const entryPointMap = new Map<string, {
name: string; type: string; filePath: string;
affected_process_count: number;
total_hits: number;
earliest_broken_step: number;
}>();
const processQuery = executeQuery(repo.id, `
MATCH (s)-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process)
WHERE s.id IN [${allIds}]
RETURN p.heuristicLabel AS name, COUNT(DISTINCT s.id) AS hits, MIN(r.step) AS minStep, p.stepCount AS stepCount
ORDER BY hits DESC
LIMIT 20
`).catch(() => []);
const moduleQuery = () => executeQuery(repo.id, `
MATCH (s)-[:CodeRelation {type: 'MEMBER_OF'}]->(c:Community)
WHERE s.id IN [${allIds}]
RETURN c.heuristicLabel AS name, COUNT(DISTINCT s.id) AS hits
ORDER BY hits DESC
LIMIT 20
`).catch(() => []);
const directModuleQuery = () => d1Ids
? executeQuery(repo.id, `
MATCH (s)-[:CodeRelation {type: 'MEMBER_OF'}]->(c:Community)
WHERE s.id IN [${d1Ids}]
RETURN DISTINCT c.heuristicLabel AS name
`).catch(() => [])
: Promise.resolve([]);
// Map process id -> entryPointId to allow fixing missing minStep values later
const processToEntryPoint = new Map<string, string>();
// Collect process ids where MIN(r.step) returned null so we can retry in batch
const processesMissingMinStep = new Set<string>();
let processRows: any[], moduleRows: any[], directModuleRows: any[];
if (isArm64Mac) {
// Sequential: avoid concurrent native DB access
processRows = await processQuery;
moduleRows = await moduleQuery();
directModuleRows = await directModuleQuery();
} else {
// Parallel: safe on non-arm64 platforms
processRows = await processQuery;
[moduleRows, directModuleRows] = await Promise.all([moduleQuery(), directModuleQuery()]);
let chunksProcessed = 0;
for (let i = 0; i < impacted.length && chunksProcessed < MAX_CHUNKS; i += CHUNK_SIZE, chunksProcessed++) {
const chunk = impacted.slice(i, i + CHUNK_SIZE);
const ids = chunk.map(item => String(item.id ?? ''));
try {
// Use parameterized list to avoid building long query strings
const rows = await executeParameterized(repo.id, `
MATCH (s)-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process)
WHERE s.id IN $ids
WITH p, COUNT(DISTINCT s.id) AS hits, MIN(r.step) AS minStep
OPTIONAL MATCH (ep {id: p.entryPointId})
RETURN p.id AS pId, p.heuristicLabel AS name, p.processType AS processType,
p.entryPointId AS entryPointId, hits, minStep, p.stepCount AS stepCount,
ep.name AS epName, labels(ep)[0] AS epType, ep.filePath AS epFilePath
`, { ids }).catch(() => []);
for (const row of rows) {
const pId = row.pId ?? row[0];
const epId = row.entryPointId ?? row[3] ?? row.pId ?? row[0];
// Track mapping from process -> entryPoint so we can backfill missing minStep
if (pId) processToEntryPoint.set(String(pId), String(epId));
// Normalize epName: prefer epName, fall back to other columns, and
// ensure we don't keep an empty string (labels(...) can return "").
const epNameRaw = row.epName ?? row[7] ?? row.name ?? row[1] ?? 'unknown';
const epName = (typeof epNameRaw === 'string' && epNameRaw.trim().length > 0) ? epNameRaw.trim() : 'unknown';
// Normalize epType: labels(ep)[0] can return an empty string in
// some DBs (LadybugDB). Using nullish coalescing (??) preserves
// empty strings, which results in empty `type` values being
// propagated. Treat empty-string labels as missing and fall back
// to the next candidate or a sensible default.
const epTypeRaw = row.epType ?? row[8] ?? '';
const epType = (typeof epTypeRaw === 'string' && epTypeRaw.trim().length > 0)
? epTypeRaw.trim()
: 'Function';
const epFilePath = row.epFilePath ?? row[9] ?? '';
const hits = row.hits ?? row[4] ?? 0;
const minStep = row.minStep ?? row[5];
// If the DB returned null for minStep, note the process id so we
// can run a follow-up query using a different aggregation strategy.
if (minStep === null || minStep === undefined) {
if (pId) processesMissingMinStep.add(String(pId));
}
if (!entryPointMap.has(epId)) {
entryPointMap.set(epId, {
name: epName,
type: epType,
filePath: epFilePath,
affected_process_count: 0,
total_hits: 0,
earliest_broken_step: Infinity,
});
}
const ep = entryPointMap.get(epId)!;
ep.affected_process_count += 1;
ep.total_hits += hits;
ep.earliest_broken_step = Math.min(ep.earliest_broken_step, minStep ?? Infinity);
}
} catch (e) {
logQueryError('impact:process-chunk', e);
}
}
// If some processes returned null minStep, try a batched follow-up query
// using the full impacted id set. This handles older indexes or DBs
// where MIN(r.step) can come back null even when step properties exist.
if (processesMissingMinStep.size > 0) {
try {
const pIds = Array.from(processesMissingMinStep);
const allImpactedIds = impacted.map(it => String(it.id ?? ''));
const missingRows = await executeParameterized(repo.id, `
MATCH (s)-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process)
WHERE p.id IN $pIds AND s.id IN $ids
RETURN p.id AS pid, MIN(r.step) AS minStep
`, { pIds, ids: allImpactedIds }).catch(() => []);
for (const mr of missingRows) {
const pid = mr.pid ?? mr[0];
const minStep = mr.minStep ?? mr[1];
const epId = processToEntryPoint.get(String(pid));
if (!epId) continue;
const ep = entryPointMap.get(epId);
if (!ep) continue;
if (typeof minStep === 'number') {
ep.earliest_broken_step = Math.min(ep.earliest_broken_step, minStep);
}
}
} catch (e) {
logQueryError('impact:process-chunk-backfill', e);
}
}
affectedProcesses = processRows.map((r: any) => ({
name: r.name || r[0],
hits: r.hits || r[1],
broken_at_step: r.minStep ?? r[2],
step_count: r.stepCount ?? r[3],
}));
// If we capped chunks, mark traversal incomplete so caller knows results are partial
if (chunksProcessed * CHUNK_SIZE < impacted.length) {
traversalComplete = false;
}
const directModuleSet = new Set(directModuleRows.map((r: any) => r.name || r[0]));
affectedProcesses = Array.from(entryPointMap.values())
.map(ep => ({
...ep,
earliest_broken_step: ep.earliest_broken_step === Infinity ? null : ep.earliest_broken_step,
}))
.sort((a, b) => b.total_hits - a.total_hits);
// ── Module enrichment: use same cap as process enrichment and parameterized queries
const maxItems = Math.min(impacted.length, MAX_CHUNKS * CHUNK_SIZE);
const cappedImpacted = impacted.slice(0, maxItems);
const allIdsArr = cappedImpacted.map((i: any) => String(i.id ?? ''));
const d1Items = (grouped[1] || []).slice(0, maxItems);
const d1IdsArr = d1Items.map((i: any) => String(i.id ?? ''));
// Chunked module enrichment: run the MEMBER_OF queries in chunks
// to avoid large single queries or concurrent Kuzu calls that can
// crash (SIGSEGV) on arm64 macOS; behavior preserves existing maxItems cap and returns equivalent aggregated results.
const moduleHitsMap = new Map<string, number>();
const directModuleSet = new Set<string>();
// Helper to run a single module chunk and accumulate hits by name
const runModuleChunk = async (idsChunk: string[]) => {
if (!idsChunk || idsChunk.length === 0) return;
try {
const rows = await executeParameterized(repo.id, `
MATCH (s)-[:CodeRelation {type: 'MEMBER_OF'}]->(c:Community)
WHERE s.id IN $ids
RETURN c.heuristicLabel AS name, COUNT(DISTINCT s.id) AS hits
ORDER BY hits DESC
LIMIT 20
`, { ids: idsChunk }).catch(() => []);
for (const r of rows) {
const name = r.name ?? r[0] ?? null;
const hits = (r.hits ?? r[1]) || 0;
if (!name) continue;
moduleHitsMap.set(name, (moduleHitsMap.get(name) || 0) + hits);
}
} catch (e) {
logQueryError('impact:module-chunk', e);
}
};
// Run module query chunks sequentially (safe on arm64 macOS)
for (let i = 0; i < allIdsArr.length; i += CHUNK_SIZE) {
const chunkIds = allIdsArr.slice(i, i + CHUNK_SIZE);
await runModuleChunk(chunkIds);
}
// Run direct module query similarly (distinct heuristic labels for depth-1 items)
const runDirectModuleChunk = async (idsChunk: string[]) => {
if (!idsChunk || idsChunk.length === 0) return;
try {
const rows = await executeParameterized(repo.id, `
MATCH (s)-[:CodeRelation {type: 'MEMBER_OF'}]->(c:Community)
WHERE s.id IN $ids
RETURN DISTINCT c.heuristicLabel AS name
`, { ids: idsChunk }).catch(() => []);
for (const r of rows) {
const name = r.name ?? r[0] ?? null;
if (name) directModuleSet.add(name);
}
} catch (e) {
logQueryError('impact:direct-module-chunk', e);
}
};
for (let i = 0; i < d1IdsArr.length; i += CHUNK_SIZE) {
const chunkIds = d1IdsArr.slice(i, i + CHUNK_SIZE);
await runDirectModuleChunk(chunkIds);
}
// Build final moduleRows array from aggregated hits map, sorted & limited
const moduleRows = Array.from(moduleHitsMap.entries())
.map(([name, hits]) => ({ name, hits }))
.sort((a, b) => b.hits - a.hits)
.slice(0, 20);
const directModuleRows = Array.from(directModuleSet).map(name => ({ name }));
// Build affectedModules in the same shape as original implementation
const directModuleNameSet = new Set(directModuleRows.map((r: any) => r.name || r[0]));
affectedModules = moduleRows.map((r: any) => {
const name = r.name || r[0];
const name = r.name ?? r[0];
const hits = r.hits ?? r[1] ?? 0;
return {
name,
hits: r.hits || r[1],
impact: directModuleSet.has(name) ? 'direct' : 'indirect',
hits,
impact: directModuleNameSet.has(name) ? 'direct' : 'indirect',
};
});
}
}
// Risk scoring
const processCount = affectedProcesses.length;

View file

@ -0,0 +1,25 @@
IDENTIFICATION DIVISION.
PROGRAM-ID. AUDITLOG.
DATA DIVISION.
WORKING-STORAGE SECTION.
01 WS-LOG-MESSAGE PIC X(80).
01 WS-TIMESTAMP PIC X(26).
LINKAGE SECTION.
01 LS-CUST-ID PIC 9(8).
01 LS-AMOUNT PIC 9(7)V99.
PROCEDURE DIVISION USING LS-CUST-ID LS-AMOUNT.
MAIN-PARAGRAPH.
PERFORM WRITE-LOG
GOBACK.
WRITE-LOG.
STRING 'Customer ' LS-CUST-ID ' amount ' LS-AMOUNT
DELIMITED BY SIZE INTO WS-LOG-MESSAGE
DISPLAY WS-LOG-MESSAGE.
ENTRY "AUDITLOG-BATCH" USING LS-CUST-ID.
DISPLAY 'Batch audit for ' LS-CUST-ID
GOBACK.

View file

@ -0,0 +1,3 @@
01 PREFIX-RECORD.
05 PREFIX-CODE PIC X(10).
05 PREFIX-NAME PIC X(30).

View file

@ -0,0 +1,6 @@
01 WS-CUSTOMER-DATA.
05 WS-CUST-CODE PIC X(10).
05 WS-CUST-TYPE PIC X(3).
88 PREMIUM-CUSTOMER VALUE 'PRM'.
88 REGULAR-CUSTOMER VALUE 'REG'.
05 WS-CUST-ADDR PIC X(50).

View file

@ -0,0 +1,74 @@
IDENTIFICATION DIVISION.
PROGRAM-ID. CUSTUPDT.
AUTHOR. TEST.
ENVIRONMENT DIVISION.
INPUT-OUTPUT SECTION.
FILE-CONTROL.
SELECT CUSTOMER-FILE ASSIGN TO 'CUSTFILE'
ORGANIZATION IS INDEXED
ACCESS IS DYNAMIC
RECORD KEY IS CUST-ID
FILE STATUS IS WS-FILE-STATUS.
DATA DIVISION.
FILE SECTION.
FD CUSTOMER-FILE.
01 CUSTOMER-RECORD.
05 CUST-ID PIC 9(8).
05 CUST-NAME PIC X(30).
05 CUST-BALANCE PIC 9(7)V99.
WORKING-STORAGE SECTION.
01 WS-FILE-STATUS PIC XX.
01 WS-CUSTOMER-NAME PIC X(30).
01 WS-AMOUNT PIC 9(7)V99.
01 WS-EOF PIC 9 VALUE 0.
88 END-OF-FILE VALUE 1.
01 WS-AMT PIC 9(5)V99.
01 WS-PROG-NAME PIC X(8).
01 FIELD-A PIC 9(5)V99.
01 FIELD-B PIC 9(5)V99.
COPY COPYLIB REPLACING ==PREFIX-== BY ==WS-==.
LINKAGE SECTION.
01 LS-PARAM PIC X(20).
PROCEDURE DIVISION.
INIT-SECTION SECTION.
MAIN-PARAGRAPH.
PERFORM INIT-PARAGRAPH
PERFORM PROCESS-PARAGRAPH
PERFORM CLEANUP-PARAGRAPH
STOP RUN.
INIT-PARAGRAPH.
OPEN I-O CUSTOMER-FILE
MOVE SPACES TO WS-CUSTOMER-NAME.
PROCESSING-SECTION SECTION.
PROCESS-PARAGRAPH.
PERFORM READ-CUSTOMER THRU WRITE-CUSTOMER
CALL "AUDITLOG" USING CUST-ID WS-AMOUNT
CALL WS-PROG-NAME.
READ-CUSTOMER.
READ CUSTOMER-FILE
NOT AT END
MOVE CUST-NAME TO WS-CUSTOMER-NAME
END-READ.
UPDATE-BALANCE.
ADD WS-AMOUNT TO CUST-BALANCE
MOVE WS-AMOUNT TO CUST-BALANCE
MOVE WS-AMT TO FIELD-A FIELD-B.
WRITE-CUSTOMER.
REWRITE CUSTOMER-RECORD.
CLEANUP-PARAGRAPH.
CLOSE CUSTOMER-FILE.
ENTRY 'ALTENTRY' USING LS-PARAM.
DISPLAY 'ALTERNATE ENTRY POINT'
GOBACK.

View file

@ -0,0 +1,33 @@
IDENTIFICATION DIVISION.
PROGRAM-ID. OUTER-PROG.
DATA DIVISION.
WORKING-STORAGE SECTION.
01 WS-OUTER-FLAG PIC 9 VALUE 0.
PROCEDURE DIVISION.
OUTER-MAIN.
PERFORM OUTER-PROCESS
CALL "INNER-PROG"
STOP RUN.
OUTER-PROCESS.
DISPLAY 'OUTER PROCESSING'.
IDENTIFICATION DIVISION.
PROGRAM-ID. INNER-PROG.
DATA DIVISION.
WORKING-STORAGE SECTION.
01 WS-INNER-CODE PIC X(5).
PROCEDURE DIVISION.
INNER-MAIN.
PERFORM INNER-PROCESS
GOBACK.
INNER-PROCESS.
DISPLAY 'INNER PROCESSING'.
END PROGRAM INNER-PROG.
END PROGRAM OUTER-PROG.

View file

@ -0,0 +1,94 @@
IDENTIFICATION DIVISION.
PROGRAM-ID. RPTGEN.
DATA DIVISION.
WORKING-STORAGE SECTION.
COPY CUSTDAT.
01 WS-REPORT-LINE PIC X(132).
01 WS-SQL-CODE PIC S9(9) COMP.
01 WS-COUNT PIC 9(4).
01 WS-MAP-NAME PIC X(8).
01 WS-SORT-FILE PIC X(8).
01 WS-QUEUE-NAME PIC X(16).
01 WS-NEXT-PGM PIC X(8).
PROCEDURE DIVISION.
MAIN-PARAGRAPH.
PERFORM FETCH-DATA
PERFORM FORMAT-REPORT
PERFORM SEND-SCREEN
CALL "CUSTUPDT"
GO TO EXIT-PARAGRAPH.
FETCH-DATA.
EXEC SQL
SELECT CUST_NAME, CUST_BALANCE
FROM CUSTOMER
WHERE CUST_ID = :WS-CUST-CODE
END-EXEC.
FORMAT-REPORT.
PERFORM WS-COUNT TIMES
MOVE WS-CUST-CODE TO WS-REPORT-LINE
END-PERFORM
PERFORM MAIN-PARAGRAPH THRU FORMAT-REPORT
IF WS-COUNT > 0 PERFORM FETCH-DATA
ELSE PERFORM SEND-SCREEN
END-IF
SORT WS-SORT-FILE USING CUSTOMER-DATA
GIVING WS-REPORT-LINE.
SORT WS-SORT-FILE ON ASCENDING KEY WS-COUNT
INPUT PROCEDURE IS BUILD-SORT-INPUT
OUTPUT PROCEDURE IS WRITE-SORTED.
MOVE CORR WS-CUSTOMER-DATA TO WS-REPORT-LINE
SEARCH WS-CUSTOMER-DATA
GO TO FETCH-DATA FORMAT-REPORT SEND-SCREEN
DEPENDING ON WS-COUNT.
SEND-SCREEN.
EXEC CICS
SEND MAP(WS-MAP-NAME) MAPSET('CUSTSET')
FROM(WS-REPORT-LINE)
END-EXEC.
EXEC CICS
LINK PROGRAM('AUDITLOG')
END-EXEC.
EXEC CICS
XCTL PROGRAM('CUSTUPDT')
END-EXEC.
EXEC CICS
READ FILE('CUSTFILE')
INTO(WS-CUSTOMER-DATA)
END-EXEC.
EXEC CICS
WRITEQ TS QUEUE('RPTQUEUE')
FROM(WS-REPORT-LINE)
END-EXEC.
EXEC CICS
HANDLE ABEND LABEL(ABEND-HANDLER)
END-EXEC.
EXEC CICS
RETURN TRANSID('RPTG')
END-EXEC.
EXEC CICS
XCTL PROGRAM(WS-NEXT-PGM)
END-EXEC.
BUILD-SORT-INPUT.
DISPLAY 'BUILDING SORT INPUT'.
WRITE-SORTED.
DISPLAY 'WRITING SORTED OUTPUT'.
ABEND-HANDLER.
DISPLAY 'ABEND OCCURRED'.
EXIT-PARAGRAPH.
STOP RUN.

View file

@ -0,0 +1,5 @@
//CUSTJOB JOB (ACCT),'CUSTOMER UPDATE',CLASS=A,MSGCLASS=X
//STEP1 EXEC PGM=CUSTUPDT
//CUSTFILE DD DSN=PROD.CUSTOMER.MASTER,DISP=SHR
//STEP2 EXEC PGM=RPTGEN
//SYSOUT DD SYSOUT=*

View file

@ -0,0 +1,6 @@
import 'models.dart';
void processUser() {
var user = getUser('alice');
user.save();
}

View file

@ -0,0 +1,11 @@
class User {
String name = '';
bool save() {
return true;
}
}
User getUser(String name) {
return User();
}

View file

@ -0,0 +1,5 @@
import 'models.dart';
void processUser(User user) {
user.address.save();
}

View file

@ -0,0 +1,16 @@
class Address {
String city = '';
void save() {
// persist address
}
}
class User {
String name = '';
Address address = Address();
String greet() {
return name;
}
}

View file

@ -0,0 +1,608 @@
/**
* COBOL: Exhaustive strict integration test.
*
* Every single node and edge produced by the COBOL/JCL pipeline is asserted
* with exact counts AND exact sorted edge-pair lists. No fuzzy assertions.
*
* Ground truth captured from the cobol-app fixture:
* CUSTUPDT.cbl, AUDITLOG.cbl, RPTGEN.cbl, NESTED.cbl,
* CUSTDAT.cpy, COPYLIB.cpy, RUNJOBS.jcl
*/
import { describe, it, expect, beforeAll } from 'vitest';
import path from 'path';
import {
FIXTURES, getRelationships, getNodesByLabel, edgeSet,
runPipelineFromRepo, type PipelineResult,
} from './helpers.js';
describe('COBOL full system extraction', () => {
let result: PipelineResult;
beforeAll(async () => {
result = await runPipelineFromRepo(
path.join(FIXTURES, 'cobol-app'),
() => {},
{ skipGraphPhases: true },
);
}, 60000);
// =====================================================================
// NODE COMPLETENESS — exact count + exact sorted name list per label
// =====================================================================
describe('node completeness', () => {
it('produces exactly 5 Module nodes', () => {
const nodes = getNodesByLabel(result, 'Module');
expect(nodes.length).toBe(5);
expect(nodes).toEqual(['AUDITLOG', 'CUSTUPDT', 'INNER-PROG', 'OUTER-PROG', 'RPTGEN']);
});
it('produces exactly 21 Function nodes', () => {
const nodes = getNodesByLabel(result, 'Function');
expect(nodes.length).toBe(21);
expect(nodes).toEqual([
'ABEND-HANDLER', 'BUILD-SORT-INPUT', 'CLEANUP-PARAGRAPH',
'EXIT-PARAGRAPH', 'FETCH-DATA', 'FORMAT-REPORT', 'INIT-PARAGRAPH',
'INNER-MAIN', 'INNER-PROCESS',
'MAIN-PARAGRAPH', 'MAIN-PARAGRAPH', 'MAIN-PARAGRAPH',
'OUTER-MAIN', 'OUTER-PROCESS',
'PROCESS-PARAGRAPH', 'READ-CUSTOMER', 'SEND-SCREEN',
'UPDATE-BALANCE', 'WRITE-CUSTOMER', 'WRITE-LOG', 'WRITE-SORTED',
]);
});
it('produces exactly 2 Namespace nodes', () => {
expect(getNodesByLabel(result, 'Namespace')).toEqual(['INIT-SECTION', 'PROCESSING-SECTION']);
});
it('produces exactly 36 Property nodes', () => {
const nodes = getNodesByLabel(result, 'Property');
expect(nodes.length).toBe(36);
expect(nodes).toEqual([
'CUST-BALANCE', 'CUST-ID', 'CUST-NAME', 'CUSTOMER-RECORD',
'END-OF-FILE', 'FIELD-A', 'FIELD-B',
'LS-AMOUNT', 'LS-CUST-ID', 'LS-PARAM',
'PREMIUM-CUSTOMER', 'REGULAR-CUSTOMER',
'WS-AMOUNT', 'WS-AMT', 'WS-CODE', 'WS-COUNT',
'WS-CUST-ADDR', 'WS-CUST-CODE', 'WS-CUST-TYPE',
'WS-CUSTOMER-DATA', 'WS-CUSTOMER-NAME', 'WS-EOF',
'WS-FILE-STATUS', 'WS-INNER-CODE', 'WS-LOG-MESSAGE',
'WS-MAP-NAME', 'WS-NAME', 'WS-NEXT-PGM', 'WS-OUTER-FLAG',
'WS-PROG-NAME', 'WS-QUEUE-NAME', 'WS-RECORD',
'WS-REPORT-LINE', 'WS-SORT-FILE', 'WS-SQL-CODE', 'WS-TIMESTAMP',
]);
});
it('produces exactly 1 Record node', () => {
expect(getNodesByLabel(result, 'Record')).toEqual(['CUSTOMER-FILE']);
});
it('produces exactly 15 CodeElement nodes', () => {
const nodes = getNodesByLabel(result, 'CodeElement');
expect(nodes.length).toBe(15);
expect(nodes).toEqual([
'CALL WS-PROG-NAME', 'CICS XCTL WS-NEXT-PGM', 'CUSTJOB',
'EXEC CICS HANDLE ABEND', 'EXEC CICS LINK', 'EXEC CICS READ',
'EXEC CICS RETURN', 'EXEC CICS SEND MAP', 'EXEC CICS WRITEQ TS',
'EXEC CICS XCTL', 'EXEC CICS XCTL', 'EXEC SQL SELECT',
'PROD.CUSTOMER.MASTER', 'STEP1', 'STEP2',
]);
});
it('produces exactly 2 Constructor nodes', () => {
expect(getNodesByLabel(result, 'Constructor')).toEqual(['ALTENTRY', 'AUDITLOG-BATCH']);
});
});
// =====================================================================
// CALLS EDGES — exact count + exact sorted pairs per reason
// =====================================================================
describe('CALLS edge completeness', () => {
it('produces exactly 15 CALLS edges with reason cobol-perform', () => {
const edges = getRelationships(result, 'CALLS').filter(e => e.rel.reason === 'cobol-perform');
expect(edges.length).toBe(15);
expect(edgeSet(edges)).toEqual([
'FORMAT-REPORT \u2192 BUILD-SORT-INPUT',
'FORMAT-REPORT \u2192 FETCH-DATA',
'FORMAT-REPORT \u2192 MAIN-PARAGRAPH',
'FORMAT-REPORT \u2192 SEND-SCREEN',
'FORMAT-REPORT \u2192 WRITE-SORTED',
'INNER-MAIN \u2192 INNER-PROCESS',
'MAIN-PARAGRAPH \u2192 CLEANUP-PARAGRAPH',
'MAIN-PARAGRAPH \u2192 FETCH-DATA',
'MAIN-PARAGRAPH \u2192 FORMAT-REPORT',
'MAIN-PARAGRAPH \u2192 INIT-PARAGRAPH',
'MAIN-PARAGRAPH \u2192 PROCESS-PARAGRAPH',
'MAIN-PARAGRAPH \u2192 SEND-SCREEN',
'MAIN-PARAGRAPH \u2192 WRITE-LOG',
'OUTER-MAIN \u2192 OUTER-PROCESS',
'PROCESS-PARAGRAPH \u2192 READ-CUSTOMER',
]);
});
it('produces exactly 2 CALLS edges with reason cobol-perform-thru', () => {
const edges = getRelationships(result, 'CALLS').filter(e => e.rel.reason === 'cobol-perform-thru');
expect(edges.length).toBe(2);
expect(edgeSet(edges)).toEqual([
'FORMAT-REPORT \u2192 FORMAT-REPORT',
'PROCESS-PARAGRAPH \u2192 WRITE-CUSTOMER',
]);
});
it('produces exactly 3 CALLS edges with reason cobol-call', () => {
const edges = getRelationships(result, 'CALLS').filter(e => e.rel.reason === 'cobol-call');
expect(edges.length).toBe(3);
expect(edgeSet(edges)).toEqual([
'CUSTUPDT \u2192 AUDITLOG',
'OUTER-PROG \u2192 INNER-PROG',
'RPTGEN \u2192 CUSTUPDT',
]);
});
it('produces exactly 4 CALLS edges with reason cobol-goto', () => {
const edges = getRelationships(result, 'CALLS').filter(e => e.rel.reason === 'cobol-goto');
expect(edges.length).toBe(4);
expect(edgeSet(edges)).toEqual([
'FORMAT-REPORT \u2192 FETCH-DATA',
'FORMAT-REPORT \u2192 FORMAT-REPORT',
'FORMAT-REPORT \u2192 SEND-SCREEN',
'MAIN-PARAGRAPH \u2192 EXIT-PARAGRAPH',
]);
});
it('produces exactly 1 CALLS edge with reason cics-link', () => {
const edges = getRelationships(result, 'CALLS').filter(e => e.rel.reason === 'cics-link');
expect(edges.length).toBe(1);
expect(edgeSet(edges)).toEqual(['RPTGEN \u2192 AUDITLOG']);
});
it('produces exactly 1 CALLS edge with reason cics-xctl', () => {
const edges = getRelationships(result, 'CALLS').filter(e => e.rel.reason === 'cics-xctl');
expect(edges.length).toBe(1);
expect(edgeSet(edges)).toEqual(['RPTGEN \u2192 CUSTUPDT']);
});
it('produces exactly 1 CALLS edge with reason cics-handle-abend', () => {
const edges = getRelationships(result, 'CALLS').filter(e => e.rel.reason === 'cics-handle-abend');
expect(edges.length).toBe(1);
expect(edgeSet(edges)).toEqual(['RPTGEN \u2192 ABEND-HANDLER']);
});
it('produces exactly 1 CALLS edge with reason cics-return-transid', () => {
const edges = getRelationships(result, 'CALLS').filter(e => e.rel.reason === 'cics-return-transid');
expect(edges.length).toBe(1);
});
it('produces exactly 2 CALLS edges with reason jcl-exec-pgm', () => {
const edges = getRelationships(result, 'CALLS').filter(e => e.rel.reason === 'jcl-exec-pgm');
expect(edges.length).toBe(2);
expect(edgeSet(edges)).toEqual(['STEP1 \u2192 CUSTUPDT', 'STEP2 \u2192 RPTGEN']);
});
it('produces exactly 1 CALLS edge with reason jcl-dd:CUSTFILE', () => {
const edges = getRelationships(result, 'CALLS').filter(e => e.rel.reason === 'jcl-dd:CUSTFILE');
expect(edges.length).toBe(1);
expect(edgeSet(edges)).toEqual(['STEP1 \u2192 PROD.CUSTOMER.MASTER']);
});
it('produces zero unresolved CALLS edges', () => {
expect(getRelationships(result, 'CALLS').filter(e => e.rel.reason.endsWith('-unresolved')).length).toBe(0);
});
});
// =====================================================================
// CONTAINS EDGES — exact count + exact sorted pairs per reason
// =====================================================================
describe('CONTAINS edge completeness', () => {
it('produces exactly 4 CONTAINS edges with reason cobol-program-id', () => {
const edges = getRelationships(result, 'CONTAINS').filter(e => e.rel.reason === 'cobol-program-id');
expect(edges.length).toBe(4);
expect(edgeSet(edges)).toEqual([
'AUDITLOG.cbl \u2192 AUDITLOG',
'CUSTUPDT.cbl \u2192 CUSTUPDT',
'NESTED.cbl \u2192 OUTER-PROG',
'RPTGEN.cbl \u2192 RPTGEN',
]);
});
it('produces exactly 1 CONTAINS edge with reason cobol-nested-program', () => {
const edges = getRelationships(result, 'CONTAINS').filter(e => e.rel.reason === 'cobol-nested-program');
expect(edges.length).toBe(1);
expect(edgeSet(edges)).toEqual(['OUTER-PROG \u2192 INNER-PROG']);
});
it('produces exactly 2 CONTAINS edges with reason cobol-section', () => {
const edges = getRelationships(result, 'CONTAINS').filter(e => e.rel.reason === 'cobol-section');
expect(edges.length).toBe(2);
expect(edgeSet(edges)).toEqual([
'CUSTUPDT \u2192 INIT-SECTION',
'CUSTUPDT \u2192 PROCESSING-SECTION',
]);
});
it('produces exactly 21 CONTAINS edges with reason cobol-paragraph', () => {
const edges = getRelationships(result, 'CONTAINS').filter(e => e.rel.reason === 'cobol-paragraph');
expect(edges.length).toBe(21);
expect(edgeSet(edges)).toEqual([
'AUDITLOG \u2192 MAIN-PARAGRAPH',
'AUDITLOG \u2192 WRITE-LOG',
'INIT-SECTION \u2192 INIT-PARAGRAPH',
'INIT-SECTION \u2192 MAIN-PARAGRAPH',
'INNER-PROG \u2192 INNER-MAIN',
'INNER-PROG \u2192 INNER-PROCESS',
'OUTER-PROG \u2192 OUTER-MAIN',
'OUTER-PROG \u2192 OUTER-PROCESS',
'PROCESSING-SECTION \u2192 CLEANUP-PARAGRAPH',
'PROCESSING-SECTION \u2192 PROCESS-PARAGRAPH',
'PROCESSING-SECTION \u2192 READ-CUSTOMER',
'PROCESSING-SECTION \u2192 UPDATE-BALANCE',
'PROCESSING-SECTION \u2192 WRITE-CUSTOMER',
'RPTGEN \u2192 ABEND-HANDLER',
'RPTGEN \u2192 BUILD-SORT-INPUT',
'RPTGEN \u2192 EXIT-PARAGRAPH',
'RPTGEN \u2192 FETCH-DATA',
'RPTGEN \u2192 FORMAT-REPORT',
'RPTGEN \u2192 MAIN-PARAGRAPH',
'RPTGEN \u2192 SEND-SCREEN',
'RPTGEN \u2192 WRITE-SORTED',
]);
});
it('produces exactly 36 CONTAINS edges with reason cobol-data-item', () => {
const edges = getRelationships(result, 'CONTAINS').filter(e => e.rel.reason === 'cobol-data-item');
expect(edges.length).toBe(36);
expect(edgeSet(edges)).toEqual([
'AUDITLOG \u2192 LS-AMOUNT',
'AUDITLOG \u2192 LS-CUST-ID',
'AUDITLOG \u2192 WS-LOG-MESSAGE',
'AUDITLOG \u2192 WS-TIMESTAMP',
'CUSTUPDT \u2192 CUST-BALANCE',
'CUSTUPDT \u2192 CUST-ID',
'CUSTUPDT \u2192 CUST-NAME',
'CUSTUPDT \u2192 CUSTOMER-RECORD',
'CUSTUPDT \u2192 END-OF-FILE',
'CUSTUPDT \u2192 FIELD-A',
'CUSTUPDT \u2192 FIELD-B',
'CUSTUPDT \u2192 LS-PARAM',
'CUSTUPDT \u2192 WS-AMOUNT',
'CUSTUPDT \u2192 WS-AMT',
'CUSTUPDT \u2192 WS-CODE',
'CUSTUPDT \u2192 WS-CUSTOMER-NAME',
'CUSTUPDT \u2192 WS-EOF',
'CUSTUPDT \u2192 WS-FILE-STATUS',
'CUSTUPDT \u2192 WS-NAME',
'CUSTUPDT \u2192 WS-PROG-NAME',
'CUSTUPDT \u2192 WS-RECORD',
'INNER-PROG \u2192 WS-INNER-CODE',
'OUTER-PROG \u2192 WS-OUTER-FLAG',
'RPTGEN \u2192 PREMIUM-CUSTOMER',
'RPTGEN \u2192 REGULAR-CUSTOMER',
'RPTGEN \u2192 WS-COUNT',
'RPTGEN \u2192 WS-CUST-ADDR',
'RPTGEN \u2192 WS-CUST-CODE',
'RPTGEN \u2192 WS-CUST-TYPE',
'RPTGEN \u2192 WS-CUSTOMER-DATA',
'RPTGEN \u2192 WS-MAP-NAME',
'RPTGEN \u2192 WS-NEXT-PGM',
'RPTGEN \u2192 WS-QUEUE-NAME',
'RPTGEN \u2192 WS-REPORT-LINE',
'RPTGEN \u2192 WS-SORT-FILE',
'RPTGEN \u2192 WS-SQL-CODE',
]);
});
it('produces exactly 8 CONTAINS edges with reason cobol-exec-cics', () => {
const edges = getRelationships(result, 'CONTAINS').filter(e => e.rel.reason === 'cobol-exec-cics');
expect(edges.length).toBe(8);
expect(edgeSet(edges)).toEqual([
'RPTGEN \u2192 EXEC CICS HANDLE ABEND',
'RPTGEN \u2192 EXEC CICS LINK',
'RPTGEN \u2192 EXEC CICS READ',
'RPTGEN \u2192 EXEC CICS RETURN',
'RPTGEN \u2192 EXEC CICS SEND MAP',
'RPTGEN \u2192 EXEC CICS WRITEQ TS',
'RPTGEN \u2192 EXEC CICS XCTL',
'RPTGEN \u2192 EXEC CICS XCTL',
]);
});
it('produces exactly 1 CONTAINS edge with reason cobol-exec-sql', () => {
expect(edgeSet(getRelationships(result, 'CONTAINS').filter(e => e.rel.reason === 'cobol-exec-sql')))
.toEqual(['RPTGEN \u2192 EXEC SQL SELECT']);
});
it('produces exactly 1 CONTAINS edge with reason cics-dynamic-program', () => {
expect(edgeSet(getRelationships(result, 'CONTAINS').filter(e => e.rel.reason === 'cics-dynamic-program')))
.toEqual(['RPTGEN \u2192 CICS XCTL WS-NEXT-PGM']);
});
it('produces exactly 1 CONTAINS edge with reason cobol-dynamic-call', () => {
expect(edgeSet(getRelationships(result, 'CONTAINS').filter(e => e.rel.reason === 'cobol-dynamic-call')))
.toEqual(['CUSTUPDT \u2192 CALL WS-PROG-NAME']);
});
it('produces exactly 2 CONTAINS edges with reason cobol-entry-point', () => {
expect(edgeSet(getRelationships(result, 'CONTAINS').filter(e => e.rel.reason === 'cobol-entry-point')))
.toEqual(['AUDITLOG \u2192 AUDITLOG-BATCH', 'CUSTUPDT \u2192 ALTENTRY']);
});
it('produces exactly 1 CONTAINS edge with reason cobol-file-declaration', () => {
expect(edgeSet(getRelationships(result, 'CONTAINS').filter(e => e.rel.reason === 'cobol-file-declaration')))
.toEqual(['CUSTUPDT \u2192 CUSTOMER-FILE']);
});
it('produces exactly 1 CONTAINS edge with reason jcl-job', () => {
expect(edgeSet(getRelationships(result, 'CONTAINS').filter(e => e.rel.reason === 'jcl-job')))
.toEqual(['RUNJOBS.jcl \u2192 CUSTJOB']);
});
it('produces exactly 2 CONTAINS edges with reason jcl-step', () => {
expect(edgeSet(getRelationships(result, 'CONTAINS').filter(e => e.rel.reason === 'jcl-step')))
.toEqual(['CUSTJOB \u2192 STEP1', 'CUSTJOB \u2192 STEP2']);
});
});
// =====================================================================
// ACCESSES EDGES — exact count + exact sorted pairs per reason
// =====================================================================
describe('ACCESSES edge completeness', () => {
it('produces exactly 4 ACCESSES edges with reason cobol-move-read', () => {
const edges = getRelationships(result, 'ACCESSES').filter(e => e.rel.reason === 'cobol-move-read');
expect(edges.length).toBe(4);
expect(edgeSet(edges)).toEqual([
'FORMAT-REPORT \u2192 WS-CUST-CODE',
'READ-CUSTOMER \u2192 CUST-NAME',
'UPDATE-BALANCE \u2192 WS-AMOUNT',
'UPDATE-BALANCE \u2192 WS-AMT',
]);
});
it('produces exactly 5 ACCESSES edges with reason cobol-move-write', () => {
const edges = getRelationships(result, 'ACCESSES').filter(e => e.rel.reason === 'cobol-move-write');
expect(edges.length).toBe(5);
expect(edgeSet(edges)).toEqual([
'FORMAT-REPORT \u2192 WS-REPORT-LINE',
'READ-CUSTOMER \u2192 WS-CUSTOMER-NAME',
'UPDATE-BALANCE \u2192 CUST-BALANCE',
'UPDATE-BALANCE \u2192 FIELD-A',
'UPDATE-BALANCE \u2192 FIELD-B',
]);
});
it('produces exactly 1 ACCESSES edge with reason cics-file-read', () => {
expect(getRelationships(result, 'ACCESSES').filter(e => e.rel.reason === 'cics-file-read').length).toBe(1);
});
it('produces exactly 1 ACCESSES edge with reason cics-map', () => {
expect(getRelationships(result, 'ACCESSES').filter(e => e.rel.reason === 'cics-map').length).toBe(1);
});
it('produces exactly 1 ACCESSES edge with reason cics-queue-write', () => {
expect(getRelationships(result, 'ACCESSES').filter(e => e.rel.reason === 'cics-queue-write').length).toBe(1);
});
it('produces exactly 1 ACCESSES edge with reason cics-receive-into', () => {
const edges = getRelationships(result, 'ACCESSES').filter(e => e.rel.reason === 'cics-receive-into');
expect(edges.length).toBe(1);
expect(edges[0].target).toBe('WS-CUSTOMER-DATA');
});
it('produces exactly 2 ACCESSES edges with reason cics-send-from', () => {
const edges = getRelationships(result, 'ACCESSES').filter(e => e.rel.reason === 'cics-send-from');
expect(edges.length).toBe(2);
expect(edgeSet(edges)).toEqual([
'EXEC CICS SEND MAP \u2192 WS-REPORT-LINE',
'EXEC CICS WRITEQ TS \u2192 WS-REPORT-LINE',
]);
});
it('produces exactly 1 ACCESSES edge with reason cobol-search', () => {
const edges = getRelationships(result, 'ACCESSES').filter(e => e.rel.reason === 'cobol-search');
expect(edges.length).toBe(1);
expect(edgeSet(edges)).toEqual(['RPTGEN \u2192 WS-CUSTOMER-DATA']);
});
it('produces exactly 1 ACCESSES edge with reason sort-using', () => {
expect(getRelationships(result, 'ACCESSES').filter(e => e.rel.reason === 'sort-using').length).toBe(1);
});
it('produces exactly 1 ACCESSES edge with reason sort-giving (multi-line SORT)', () => {
expect(getRelationships(result, 'ACCESSES').filter(e => e.rel.reason === 'sort-giving').length).toBe(1);
});
it('produces exactly 2 ACCESSES edges with reason cobol-procedure-using', () => {
const edges = getRelationships(result, 'ACCESSES').filter(e => e.rel.reason === 'cobol-procedure-using');
expect(edges.length).toBe(2);
expect(edgeSet(edges)).toEqual([
'AUDITLOG \u2192 LS-AMOUNT',
'AUDITLOG \u2192 LS-CUST-ID',
]);
});
it('produces exactly 1 ACCESSES edge with reason sql-select', () => {
expect(getRelationships(result, 'ACCESSES').filter(e => e.rel.reason === 'sql-select').length).toBe(1);
});
});
// =====================================================================
// IMPORTS EDGES — exact pairs
// =====================================================================
describe('IMPORTS edge completeness', () => {
it('produces exactly 2 IMPORTS edges with reason cobol-copy', () => {
const edges = getRelationships(result, 'IMPORTS').filter(e => e.rel.reason === 'cobol-copy');
expect(edges.length).toBe(2);
});
});
// =====================================================================
// FEATURE-SPECIFIC ASSERTIONS — validates all review findings resolved
// =====================================================================
describe('multi-PERFORM on same line (Finding #III)', () => {
it('captures both PERFORMs in IF/ELSE on a single logical line', () => {
// IF WS-COUNT > 0 PERFORM FETCH-DATA ELSE PERFORM SEND-SCREEN
const edges = getRelationships(result, 'CALLS').filter(
e => e.rel.reason === 'cobol-perform' && e.source === 'FORMAT-REPORT',
);
const targets = edges.map(e => e.target).sort();
expect(targets).toContain('FETCH-DATA');
expect(targets).toContain('SEND-SCREEN');
});
});
describe('INPUT/OUTPUT PROCEDURE IS in SORT (Finding #iii)', () => {
it('creates CALLS edges for INPUT PROCEDURE and OUTPUT PROCEDURE targets', () => {
const edges = getRelationships(result, 'CALLS').filter(
e => e.rel.reason === 'cobol-perform' && e.source === 'FORMAT-REPORT',
);
const targets = edges.map(e => e.target).sort();
expect(targets).toContain('BUILD-SORT-INPUT');
expect(targets).toContain('WRITE-SORTED');
});
it('creates paragraph nodes for INPUT/OUTPUT PROCEDURE targets', () => {
const nodes = getNodesByLabel(result, 'Function');
expect(nodes).toContain('BUILD-SORT-INPUT');
expect(nodes).toContain('WRITE-SORTED');
});
});
describe('GO TO DEPENDING ON multi-target (Finding #iv)', () => {
it('captures all three targets from GO TO ... DEPENDING ON', () => {
// GO TO FETCH-DATA FORMAT-REPORT SEND-SCREEN DEPENDING ON WS-COUNT
const edges = getRelationships(result, 'CALLS').filter(
e => e.rel.reason === 'cobol-goto' && e.source === 'FORMAT-REPORT',
);
expect(edges.length).toBe(3);
expect(edgeSet(edges)).toEqual([
'FORMAT-REPORT \u2192 FETCH-DATA',
'FORMAT-REPORT \u2192 FORMAT-REPORT',
'FORMAT-REPORT \u2192 SEND-SCREEN',
]);
});
});
describe('MOVE CORR abbreviation (Finding #IV)', () => {
it('produces ACCESSES edges for MOVE CORR with corresponding reason', () => {
const readEdges = getRelationships(result, 'ACCESSES').filter(
e => e.rel.reason === 'cobol-move-corresponding-read',
);
expect(readEdges.length).toBe(1);
expect(edgeSet(readEdges)).toEqual(['FORMAT-REPORT \u2192 WS-CUSTOMER-DATA']);
const writeEdges = getRelationships(result, 'ACCESSES').filter(
e => e.rel.reason === 'cobol-move-corresponding-write',
);
expect(writeEdges.length).toBe(1);
expect(edgeSet(writeEdges)).toEqual(['FORMAT-REPORT \u2192 WS-REPORT-LINE']);
});
});
describe('nested program CONTAINS attribution (Finding #I, #II)', () => {
it('attributes INNER-PROG paragraphs to INNER-PROG, not OUTER-PROG', () => {
const edges = getRelationships(result, 'CONTAINS').filter(
e => e.rel.reason === 'cobol-paragraph' && e.target === 'INNER-MAIN',
);
expect(edges.length).toBe(1);
expect(edges[0].source).toBe('INNER-PROG');
});
it('attributes INNER-PROG data items to INNER-PROG, not OUTER-PROG', () => {
const edges = getRelationships(result, 'CONTAINS').filter(
e => e.rel.reason === 'cobol-data-item' && e.target === 'WS-INNER-CODE',
);
expect(edges.length).toBe(1);
expect(edges[0].source).toBe('INNER-PROG');
});
it('attributes OUTER-PROG data items to OUTER-PROG', () => {
const edges = getRelationships(result, 'CONTAINS').filter(
e => e.rel.reason === 'cobol-data-item' && e.target === 'WS-OUTER-FLAG',
);
expect(edges.length).toBe(1);
expect(edges[0].source).toBe('OUTER-PROG');
});
});
describe('per-program PROCEDURE DIVISION USING (Finding #III partial)', () => {
it('creates ACCESSES edges from AUDITLOG, not from wrong program', () => {
const edges = getRelationships(result, 'ACCESSES').filter(
e => e.rel.reason === 'cobol-procedure-using',
);
expect(edges.length).toBe(2);
// Both edges should source from AUDITLOG (the program that declares USING)
for (const e of edges) {
expect(e.source).toBe('AUDITLOG');
}
});
});
describe('PERFORM THRU edge correctness', () => {
it('captures FORMAT-REPORT PERFORM THRU from MAIN-PARAGRAPH', () => {
const edges = getRelationships(result, 'CALLS').filter(
e => e.rel.reason === 'cobol-perform-thru',
);
expect(edgeSet(edges)).toContain('FORMAT-REPORT \u2192 FORMAT-REPORT');
});
});
describe('nested program CALLS attribution', () => {
it('attributes INNER-PROG PERFORM edges to INNER-PROG paragraphs', () => {
const edges = getRelationships(result, 'CALLS').filter(
e => e.rel.reason === 'cobol-perform' && e.source === 'INNER-MAIN',
);
expect(edges.length).toBe(1);
expect(edges[0].target).toBe('INNER-PROCESS');
});
});
// =====================================================================
// GRAND TOTALS — catch any unexpected edge leakage
// =====================================================================
describe('grand totals', () => {
it('produces exactly 31 total CALLS edges', () => {
// 15 perform + 2 perform-thru + 3 call + 4 goto + 1 link + 1 xctl
// + 1 handle-abend + 1 return-transid + 2 jcl-exec-pgm + 1 jcl-dd
expect(getRelationships(result, 'CALLS').length).toBe(31);
});
it('produces exactly 81 total CONTAINS edges', () => {
// 4 program-id + 1 nested-program + 2 section + 21 paragraph
// + 36 data-item + 8 exec-cics + 1 exec-sql + 1 dynamic-call
// + 1 cics-dynamic-program + 2 entry-point + 1 file-declaration
// + 1 jcl-job + 2 jcl-step
expect(getRelationships(result, 'CONTAINS').length).toBe(81);
});
it('produces exactly 2 total IMPORTS edges', () => {
expect(getRelationships(result, 'IMPORTS').length).toBe(2);
});
it('produces exactly 25 total ACCESSES edges', () => {
// 4 move-read + 5 move-write + 1 move-corresponding-read + 1 move-corresponding-write
// + 1 file-read + 1 map + 1 queue-write
// + 1 receive-into + 2 send-from + 1 search + 1 sort-using + 1 sort-giving
// + 2 procedure-using + 1 sql-select + 2 call-using
expect(getRelationships(result, 'ACCESSES').length).toBe(25);
});
});
});

View file

@ -0,0 +1,147 @@
/**
* Dart: field-type resolution and call-result binding.
* Verifies that class fields are captured as Property nodes with HAS_PROPERTY
* edges, and that calls (including chained and call-result-bound) are resolved.
*
* Remaining known Dart gaps (field-chain ACCESSES) are documented as
* it.todo() tests to be filled when the pipeline is extended.
*/
import { describe, it, expect, beforeAll } from 'vitest';
import path from 'path';
import {
FIXTURES, getRelationships, getNodesByLabel, edgeSet,
runPipelineFromRepo, type PipelineResult,
} from './helpers.js';
import { isLanguageAvailable } from '../../../src/core/tree-sitter/parser-loader.js';
import { SupportedLanguages } from '../../../src/config/supported-languages.js';
const dartAvailable = isLanguageAvailable(SupportedLanguages.Dart);
// ── Phase 8: Field-type resolution ──────────────────────────────────────
describe.skipIf(!dartAvailable)('Dart field-type resolution', () => {
let result: PipelineResult;
beforeAll(async () => {
result = await runPipelineFromRepo(
path.join(FIXTURES, 'dart-field-types'),
() => {},
);
}, 60000);
it('detects classes and their properties', () => {
expect(getNodesByLabel(result, 'Class')).toEqual(
expect.arrayContaining(['Address', 'User']),
);
const properties = getNodesByLabel(result, 'Property');
expect(properties).toContain('address');
expect(properties).toContain('city');
expect(properties).toContain('name');
});
it('emits HAS_PROPERTY edges from class to field', () => {
const propEdges = getRelationships(result, 'HAS_PROPERTY');
expect(edgeSet(propEdges)).toEqual(
expect.arrayContaining([
'User → address',
'User → name',
'Address → city',
]),
);
});
it('resolves save() call from field-chain user.address.save()', () => {
const calls = getRelationships(result, 'CALLS');
// Dart attributes calls to the enclosing Function
const saveCalls = calls.filter(
(c) => c.target === 'save' && c.sourceFilePath.includes('app.dart'),
);
expect(saveCalls.length).toBe(1);
expect(saveCalls[0]!.targetFilePath).toContain('models.dart');
});
it('attributes save() call source to processUser, not File', () => {
const calls = getRelationships(result, 'CALLS');
const saveCalls = calls.filter(
(c) => c.target === 'save' && c.sourceFilePath.includes('app.dart'),
);
expect(saveCalls.length).toBe(1);
expect(saveCalls[0]!.source).toBe('processUser');
expect(saveCalls[0]!.sourceLabel).toBe('Function');
});
it('creates IMPORTS edge between app.dart and models.dart', () => {
const imports = getRelationships(result, 'IMPORTS');
const appImports = imports.filter(
(e) => e.sourceFilePath.includes('app.dart') && e.targetFilePath.includes('models.dart'),
);
expect(appImports.length).toBe(1);
});
// Dart field-chain ACCESSES edges require the call-processor's chain-resolution
// tier (Step 1c) to fire. This needs the type-env's scoped parameter binding
// (processUser's `user: User`) to propagate to processCallsFromExtracted so
// walkMixedChain can resolve User → address → Address and emit ACCESSES.
// The chain extraction (extractMixedChain) and member detection
// (MEMBER_ACCESS_NODE_TYPES) are wired, but the base receiver type lookup
// from the type-env currently returns undefined for Dart function parameters
// in the call-processor context. Tracked for follow-up.
it.skip('emits ACCESSES edges for field reads in chains', () => {
const accesses = getRelationships(result, 'ACCESSES');
const addressReads = accesses.filter(
(e) => e.target === 'address' && e.rel.reason === 'read',
);
expect(addressReads.length).toBe(1);
expect(addressReads[0]!.source).toBe('processUser');
expect(addressReads[0]!.targetLabel).toBe('Property');
});
});
// ── Phase 9: Call-result binding ────────────────────────────────────────
describe.skipIf(!dartAvailable)('Dart call-result binding', () => {
let result: PipelineResult;
beforeAll(async () => {
result = await runPipelineFromRepo(
path.join(FIXTURES, 'dart-call-result-binding'),
() => {},
);
}, 60000);
it('detects classes, methods, and functions', () => {
expect(getNodesByLabel(result, 'Class')).toContain('User');
expect(getNodesByLabel(result, 'Function')).toEqual(
expect.arrayContaining(['getUser', 'processUser']),
);
});
it('resolves save() call via call-result binding', () => {
const calls = getRelationships(result, 'CALLS');
// Dart attributes calls to the enclosing Function
const saveCalls = calls.filter(
(c) => c.target === 'save' && c.sourceFilePath.includes('app.dart'),
);
expect(saveCalls.length).toBe(1);
expect(saveCalls[0]!.targetFilePath).toContain('models.dart');
});
it('resolves getUser() call', () => {
const calls = getRelationships(result, 'CALLS');
const getUserCalls = calls.filter(
(c) => c.target === 'getUser' && c.sourceFilePath.includes('app.dart'),
);
expect(getUserCalls.length).toBe(1);
});
it('attributes calls to processUser, not File', () => {
const calls = getRelationships(result, 'CALLS');
const appCalls = calls.filter(
(c) => c.sourceFilePath.includes('app.dart'),
);
for (const call of appCalls) {
expect(call.source).toBe('processUser');
expect(call.sourceLabel).toBe('Function');
}
});
});

View file

@ -60,11 +60,12 @@ describe('Java heritage resolution', () => {
expect(extends_.some(e => e.target === 'Validatable')).toBe(false);
});
it('emits exactly 2 CALLS edges', () => {
it('emits exactly 3 CALLS edges', () => {
const calls = getRelationships(result, 'CALLS');
expect(calls.length).toBe(2);
expect(calls.length).toBe(3);
expect(edgeSet(calls)).toEqual([
'processUser → save',
'processUser → serialize',
'processUser → validate',
]);
});

View file

@ -10,13 +10,17 @@
import { describe, it, expect, vi, beforeEach } from 'vitest';
// We need to mock the LadybugDB adapter and repo-manager BEFORE importing LocalBackend
vi.mock('../../src/mcp/core/lbug-adapter.js', () => ({
initLbug: vi.fn().mockResolvedValue(undefined),
executeQuery: vi.fn().mockResolvedValue([]),
executeParameterized: vi.fn().mockResolvedValue([]),
closeLbug: vi.fn().mockResolvedValue(undefined),
isLbugReady: vi.fn().mockReturnValue(true),
}));
vi.mock('../../src/mcp/core/lbug-adapter.js', async (importOriginal) => {
const actual = await importOriginal();
return {
...actual,
initLbug: vi.fn().mockResolvedValue(undefined),
executeQuery: vi.fn().mockResolvedValue([]),
executeParameterized: vi.fn().mockResolvedValue([]),
closeLbug: vi.fn().mockResolvedValue(undefined),
isLbugReady: vi.fn().mockReturnValue(true),
};
});
vi.mock('../../src/storage/repo-manager.js', () => ({
listRegisteredRepos: vi.fn().mockResolvedValue([]),

View file

@ -0,0 +1,69 @@
/**
* Unit Tests: COBOL Copy Expander — pseudotext REPLACING support
*/
import { describe, it, expect } from 'vitest';
import { parseReplacingClause } from '../../src/core/ingestion/cobol/cobol-copy-expander.js';
describe('parseReplacingClause', () => {
// Existing quoted-string behavior preserved
it('parses quoted EXACT replacement', () => {
const result = parseReplacingClause(' "OLD-NAME" BY "NEW-NAME" ');
expect(result).toEqual([{ type: 'EXACT', from: 'OLD-NAME', to: 'NEW-NAME' }]);
});
it('parses LEADING replacement', () => {
const result = parseReplacingClause(' LEADING "ESP-" BY "LK-ESP-" ');
expect(result).toEqual([{ type: 'LEADING', from: 'ESP-', to: 'LK-ESP-' }]);
});
it('parses TRAILING replacement', () => {
const result = parseReplacingClause(' TRAILING "-IN" BY "-OUT" ');
expect(result).toEqual([{ type: 'TRAILING', from: '-IN', to: '-OUT' }]);
});
// Pseudotext ==...== support (isPseudotext flag propagated)
it('parses basic pseudotext: ==OLD== BY ==NEW==', () => {
const result = parseReplacingClause(' ==WS-OLD== BY ==WS-NEW== ');
expect(result).toEqual([{ type: 'EXACT', from: 'WS-OLD', to: 'WS-NEW', isPseudotext: true }]);
});
it('parses empty pseudotext (deletion): ==TEXT== BY ====', () => {
const result = parseReplacingClause(' ==REMOVE-ME== BY ==== ');
expect(result).toEqual([{ type: 'EXACT', from: 'REMOVE-ME', to: '', isPseudotext: true }]);
});
it('parses pseudotext with spaces: ==SOME TEXT== BY ==OTHER TEXT==', () => {
const result = parseReplacingClause(' ==WORKING STORAGE== BY ==LOCAL STORAGE== ');
expect(result).toEqual([{ type: 'EXACT', from: 'WORKING STORAGE', to: 'LOCAL STORAGE', isPseudotext: true }]);
});
it('parses pseudotext with single = inside: ==A=B== BY ==C=D==', () => {
const result = parseReplacingClause(' ==A=B== BY ==C=D== ');
expect(result).toEqual([{ type: 'EXACT', from: 'A=B', to: 'C=D', isPseudotext: true }]);
});
it('parses mixed quoted + pseudotext in one clause', () => {
const result = parseReplacingClause(
' "OLD-NAME" BY "NEW-NAME" ==DEL-PREFIX== BY ==== ',
);
expect(result).toEqual([
{ type: 'EXACT', from: 'OLD-NAME', to: 'NEW-NAME' },
{ type: 'EXACT', from: 'DEL-PREFIX', to: '', isPseudotext: true },
]);
});
it('LEADING modifier works alongside pseudotext', () => {
const result = parseReplacingClause(
' LEADING "ESP-" BY "LK-ESP-" ==OLD-EXACT== BY ==NEW-EXACT== ',
);
expect(result).toEqual([
{ type: 'LEADING', from: 'ESP-', to: 'LK-ESP-' },
{ type: 'EXACT', from: 'OLD-EXACT', to: 'NEW-EXACT', isPseudotext: true },
]);
});
it('returns empty array for empty input', () => {
expect(parseReplacingClause('')).toEqual([]);
expect(parseReplacingClause(' ')).toEqual([]);
});
});

File diff suppressed because it is too large Load diff

View file

@ -186,4 +186,61 @@ describe('createKnowledgeGraph', () => {
g.forEachRelationship(r => types.push(r.type));
expect(types).toEqual(['CALLS']);
});
// ─── removeRelationship ─────────────────────────────────────────────
it('removes a relationship by id', () => {
const g = createKnowledgeGraph();
g.addNode(makeNode('fn:a', 'a'));
g.addNode(makeNode('fn:b', 'b'));
g.addRelationship(makeRel('fn:a', 'fn:b'));
expect(g.relationshipCount).toBe(1);
const removed = g.removeRelationship('fn:a-CALLS-fn:b');
expect(removed).toBe(true);
expect(g.relationshipCount).toBe(0);
});
it('removeRelationship returns false for unknown id', () => {
const g = createKnowledgeGraph();
expect(g.removeRelationship('nonexistent')).toBe(false);
});
it('removeRelationship returns false on second call with same id', () => {
const g = createKnowledgeGraph();
g.addNode(makeNode('fn:a', 'a'));
g.addNode(makeNode('fn:b', 'b'));
g.addRelationship(makeRel('fn:a', 'fn:b'));
expect(g.removeRelationship('fn:a-CALLS-fn:b')).toBe(true);
expect(g.removeRelationship('fn:a-CALLS-fn:b')).toBe(false);
});
it('removeRelationship does not affect nodes', () => {
const g = createKnowledgeGraph();
g.addNode(makeNode('fn:a', 'a'));
g.addNode(makeNode('fn:b', 'b'));
g.addRelationship(makeRel('fn:a', 'fn:b'));
g.removeRelationship('fn:a-CALLS-fn:b');
expect(g.nodeCount).toBe(2);
expect(g.getNode('fn:a')).toBeDefined();
expect(g.getNode('fn:b')).toBeDefined();
});
it('removeRelationship leaves other relationships intact', () => {
const g = createKnowledgeGraph();
g.addNode(makeNode('fn:a', 'a'));
g.addNode(makeNode('fn:b', 'b'));
g.addNode(makeNode('fn:c', 'c'));
g.addRelationship(makeRel('fn:a', 'fn:b'));
g.addRelationship(makeRel('fn:b', 'fn:c'));
expect(g.relationshipCount).toBe(2);
g.removeRelationship('fn:a-CALLS-fn:b');
expect(g.relationshipCount).toBe(1);
const remaining = [...g.iterRelationships()];
expect(remaining[0].sourceId).toBe('fn:b');
expect(remaining[0].targetId).toBe('fn:c');
});
});

View file

@ -0,0 +1,230 @@
import { describe, it, expect, vi, beforeEach } from 'vitest';
// Mock the lbug-adapter module before importing LocalBackend so the class
// uses the mocked implementations of executeQuery / executeParameterized.
const executeQueryMock = vi.fn();
const executeParameterizedMock = vi.fn();
// Use the exact import specifier including .js to match runtime imports
vi.mock('../../src/mcp/core/lbug-adapter.js', async (importOriginal) => {
const actual = await importOriginal();
return {
...actual,
initLbug: vi.fn(),
executeQuery: (...args: any[]) => executeQueryMock(...args),
executeParameterized: (...args: any[]) => executeParameterizedMock(...args),
closeLbug: vi.fn(),
isLbugReady: vi.fn().mockReturnValue(true),
};
});
import { LocalBackend } from '../../src/mcp/local/local-backend';
describe('impact: batching and grouping', () => {
beforeEach(() => {
vi.clearAllMocks();
});
it('batches 250 IDs into 3 chunked STEP_IN_PROCESS queries', async () => {
// Prepare backend and a fake repo handle
const backend = new LocalBackend();
const repoHandle = {
id: 'repo1', name: 'repo1', repoPath: '/tmp/repo', storagePath: '/tmp/repo/.gitnexus',
lbugPath: '/tmp/repo/.gitnexus/lbug', indexedAt: 'now', lastCommit: 'c', stats: {},
} as any;
(backend as any).repos.set(repoHandle.id, repoHandle);
(backend as any).ensureInitialized = vi.fn().mockResolvedValue(undefined);
// executeParameterized: resolve target -> return a symbol row (default)
executeParameterizedMock.mockImplementation(async (...args: any[]) => {
const query = typeof args[1] === 'string' ? args[1] : String(args[0] ?? '');
// The initial target-resolution call will not contain STEP_IN_PROCESS
if (!query.includes('STEP_IN_PROCESS')) return [{ id: 'sym1', name: 'Target', filePath: 'f' }];
// For STEP_IN_PROCESS calls, fall through to test's executeQueryMock logic by returning [] here.
return [];
});
// Track chunk sizes
const chunkSizes: number[] = [];
let chunkCallIndex = 0;
executeQueryMock.mockImplementation(async (...args: any[]) => {
const query = typeof args[1] === 'string' ? args[1] : String(args[0] ?? '');
// Depth traversal query (find related nodes) -- return 250 impacted ids
if (query.includes("r.type IN") && !query.includes('STEP_IN_PROCESS')) {
const res: any[] = [];
for (let i = 0; i < 250; i++) {
res.push({ id: `node-${i}`, name: `n${i}`, filePath: `file-${i}.js`, relType: 'CALLS', confidence: null });
}
return res;
}
// NOTE: process-chunk enrichment previously used executeQuery; our
// implementation now calls executeParameterized for those chunks. We
// still keep this branch to support any legacy calls, but primary
// chunk tracking will be handled via executeParameterizedMock below.
return [];
});
// Handle parameterized calls (including chunked STEP_IN_PROCESS queries)
executeParameterizedMock.mockImplementation(async (...args: any[]) => {
const query = typeof args[1] === 'string' ? args[1] : String(args[0] ?? '');
const params = args[2] || {};
if (query.includes('STEP_IN_PROCESS')) {
// Count ids passed in as params.ids
const ids = Array.isArray(params.ids) ? params.ids : [];
const cnt = ids.length;
chunkSizes.push(cnt);
const idx = chunkCallIndex++;
return [{ entryPointId: `ep-${Math.floor(idx)}`, epName: `epName-${idx}`, epType: 'Function', epFilePath: `/path/${idx}`, hits: cnt, minStep: 1 }];
}
// Default target resolution
return [{ id: 'sym1', name: 'Target', filePath: 'f' }];
});
const params = { target: 'Target', direction: 'downstream', maxDepth: 1 } as any;
const res = await (backend as any)._impactImpl(repoHandle, params);
// Expect 3 chunk calls: 100 + 100 + 50
expect(chunkSizes.length).toBe(3);
const total = chunkSizes.reduce((s, v) => s + v, 0);
expect(total).toBe(250);
// Result impacted count should be 250
expect(res.impactedCount).toBe(250);
});
it('groups entry points across chunks and deduplicates correctly', async () => {
const backend = new LocalBackend();
const repoHandle = {
id: 'repo2', name: 'repo2', repoPath: '/tmp/repo2', storagePath: '/tmp/repo2/.gitnexus',
lbugPath: '/tmp/repo2/.gitnexus/lbug', indexedAt: 'now', lastCommit: 'c', stats: {},
} as any;
(backend as any).repos.set(repoHandle.id, repoHandle);
(backend as any).ensureInitialized = vi.fn().mockResolvedValue(undefined);
executeParameterizedMock.mockImplementation(async (...args: any[]) => {
const query = typeof args[1] === 'string' ? args[1] : String(args[0] ?? '');
if (!query.includes('STEP_IN_PROCESS')) return [{ id: 'symA', name: 'TargetA', filePath: 'f' }];
// For STEP_IN_PROCESS in this test, return grouping rows
return [
{ entryPointId: 'ep-1', epName: 'EP1', epType: 'Function', epFilePath: '/p/1', hits: 2, minStep: 1 },
{ entryPointId: 'ep-2', epName: 'EP2', epType: 'Function', epFilePath: '/p/2', hits: 2, minStep: 2 },
{ entryPointId: 'ep-1', epName: 'EP1', epType: 'Function', epFilePath: '/p/1', hits: 1, minStep: 3 },
{ entryPointId: 'ep-3', epName: 'EP3', epType: 'Function', epFilePath: '/p/3', hits: 1, minStep: 1 },
];
});
// Prepare impacted nodes: smaller set for clarity (6 nodes -> chunk size default 100 so single chunk)
executeQueryMock.mockImplementation(async (...args: any[]) => {
const query = typeof args[1] === 'string' ? args[1] : String(args[0] ?? '');
if (query.includes("r.type IN") && !query.includes('STEP_IN_PROCESS')) {
// return 6 nodes
const res: any[] = [];
for (let i = 0; i < 6; i++) res.push({ id: `node-${i}`, name: `n${i}`, filePath: `file-${i}.js`, relType: 'CALLS', confidence: null });
return res;
}
return [];
});
const params = { target: 'TargetA', direction: 'downstream', maxDepth: 1 } as any;
const res = await (backend as any)._impactImpl(repoHandle, params);
// affected_processes should be grouped by entryPointId: ep-1, ep-2, ep-3 => 3 unique
expect(Array.isArray(res.affected_processes)).toBe(true);
const names = res.affected_processes.map((p: any) => p.name);
expect(names.sort()).toEqual(['EP1', 'EP2', 'EP3'].sort());
const ep1 = res.affected_processes.find((p: any) => p.name === 'EP1');
expect(ep1.total_hits).toBe(3);
const ep2 = res.affected_processes.find((p: any) => p.name === 'EP2');
expect(ep2.total_hits).toBe(2);
});
it('caps enrichment to MAX_CHUNKS and sets partial when capped', async () => {
// Temporarily set MAX_CHUNKS small for deterministic test
process.env.IMPACT_MAX_CHUNKS = '3'; // CHUNK_SIZE 100 => maxItems = 300
const backend = new LocalBackend();
const repoHandle = {
id: 'repo3', name: 'repo3', repoPath: '/tmp/repo3', storagePath: '/tmp/repo3/.gitnexus',
lbugPath: '/tmp/repo3/.gitnexus/lbug', indexedAt: 'now', lastCommit: 'c', stats: {},
} as any;
(backend as any).repos.set(repoHandle.id, repoHandle);
(backend as any).ensureInitialized = vi.fn().mockResolvedValue(undefined);
// Depth traversal returns 500 impacted nodes
executeQueryMock.mockImplementation(async (...args: any[]) => {
const query = typeof args[1] === 'string' ? args[1] : String(args[0] ?? '');
if (query.includes("r.type IN") && !query.includes('STEP_IN_PROCESS')) {
const res: any[] = [];
for (let i = 0; i < 500; i++) res.push({ id: `node-${i}`, name: `n${i}`, filePath: `file-${i}.js`, relType: 'CALLS', confidence: null });
return res;
}
return [];
});
const chunkSizes: number[] = [];
executeParameterizedMock.mockImplementation(async (...args: any[]) => {
const query = typeof args[1] === 'string' ? args[1] : String(args[0] ?? '');
const params = args[2] || {};
if (query.includes('STEP_IN_PROCESS')) {
const ids = Array.isArray(params.ids) ? params.ids : [];
chunkSizes.push(ids.length);
return [{ entryPointId: 'ep-x', epName: 'EPX', epType: 'Function', epFilePath: '/p/x', hits: ids.length, minStep: 1 }];
}
if (query.includes('COUNT(DISTINCT s.id)')) {
// moduleQuery: return a module row
return [{ name: 'ModuleA', hits: 42 }];
}
if (query.includes('RETURN DISTINCT c.heuristicLabel')) {
// directModuleQuery
return [{ name: 'ModuleA' }];
}
// Default: target resolution
return [{ id: 'symX', name: 'TargetX', filePath: 'f' }];
});
const params = { target: 'TargetX', direction: 'downstream', maxDepth: 1 } as any;
const res = await (backend as any)._impactImpl(repoHandle, params);
// Expect we processed only MAX_CHUNKS chunks (3) -> total ids handled = 300
expect(chunkSizes.length).toBe(3);
const totalHandled = chunkSizes.reduce((s, v) => s + v, 0);
expect(totalHandled).toBe(300);
// Because we capped enrichment, the result should include partial: true
expect(res.partial).toBe(true);
// Module enrichment should have been called in chunks (3 calls, totaling 300 ids)
const memberCalls = (executeParameterizedMock.mock.calls || []).filter((c: any[]) => {
const q = typeof c[1] === 'string' ? c[1] : String(c[0] ?? '');
// Only count the module-hits query (which returns COUNT(DISTINCT s.id)).
// The process-chunk query also uses COUNT(DISTINCT s.id), so require MEMBER_OF
// to avoid double-counting process-chunk calls.
return q.includes('COUNT(DISTINCT s.id)') && q.includes('MEMBER_OF');
});
// MAX_CHUNKS = 3 in this test, so expect 3 module-enrichment chunk calls
// DEBUG: print memberCalls and their ids lengths
expect(memberCalls.length).toBe(3);
const totalModuleIds = memberCalls.reduce((sum: number, call: any[]) => sum + ((Array.isArray(call[2]?.ids) ? call[2].ids.length : 0)), 0);
// eslint-disable-next-line no-console
expect(totalModuleIds).toBe(300);
// Affected modules should include ModuleA
expect(Array.isArray(res.affected_modules)).toBe(true);
const modNames = res.affected_modules.map((m: any) => m.name);
expect(modNames).toContain('ModuleA');
// Cleanup env
delete process.env.IMPACT_MAX_CHUNKS;
});
});

View file

@ -1,9 +1,9 @@
import { describe, it, expect } from 'vitest';
import { getLanguageFromFilename } from '../../src/core/ingestion/utils/language-detection.js';
import { isBuiltInOrNoise } from '../../src/core/ingestion/utils/noise-filter.js';
import { getProvider } from '../../src/core/ingestion/languages/index.js';
import { SupportedLanguages } from '../../src/config/supported-languages.js';
import { extractFunctionName } from '../../src/core/ingestion/utils/ast-helpers.js';
import { getTreeSitterBufferSize, TREE_SITTER_BUFFER_SIZE, TREE_SITTER_MAX_BUFFER } from '../../src/core/ingestion/constants.js';
import { SupportedLanguages } from '../../src/config/supported-languages.js';
import Parser from 'tree-sitter';
import C from 'tree-sitter-c';
import CPP from 'tree-sitter-cpp';
@ -140,203 +140,212 @@ describe('getLanguageFromFilename', () => {
});
describe('isBuiltInOrNoise', () => {
const js = getProvider(SupportedLanguages.JavaScript);
const py = getProvider(SupportedLanguages.Python);
const php = getProvider(SupportedLanguages.PHP);
const c = getProvider(SupportedLanguages.C);
const kt = getProvider(SupportedLanguages.Kotlin);
const swift = getProvider(SupportedLanguages.Swift);
const rust = getProvider(SupportedLanguages.Rust);
const cs = getProvider(SupportedLanguages.CSharp);
describe('JavaScript/TypeScript', () => {
it('filters console methods', () => {
expect(isBuiltInOrNoise('console')).toBe(true);
expect(isBuiltInOrNoise('log')).toBe(true);
expect(isBuiltInOrNoise('warn')).toBe(true);
expect(js.isBuiltInName('console')).toBe(true);
expect(js.isBuiltInName('log')).toBe(true);
expect(js.isBuiltInName('warn')).toBe(true);
});
it('filters React hooks', () => {
expect(isBuiltInOrNoise('useState')).toBe(true);
expect(isBuiltInOrNoise('useEffect')).toBe(true);
expect(isBuiltInOrNoise('useCallback')).toBe(true);
expect(js.isBuiltInName('useState')).toBe(true);
expect(js.isBuiltInName('useEffect')).toBe(true);
expect(js.isBuiltInName('useCallback')).toBe(true);
});
it('filters array methods', () => {
expect(isBuiltInOrNoise('map')).toBe(true);
expect(isBuiltInOrNoise('filter')).toBe(true);
expect(isBuiltInOrNoise('reduce')).toBe(true);
expect(js.isBuiltInName('map')).toBe(true);
expect(js.isBuiltInName('filter')).toBe(true);
expect(js.isBuiltInName('reduce')).toBe(true);
});
});
describe('Python', () => {
it('filters built-in functions', () => {
expect(isBuiltInOrNoise('print')).toBe(true);
expect(isBuiltInOrNoise('len')).toBe(true);
expect(isBuiltInOrNoise('range')).toBe(true);
expect(py.isBuiltInName('print')).toBe(true);
expect(py.isBuiltInName('len')).toBe(true);
expect(py.isBuiltInName('range')).toBe(true);
});
});
describe('PHP', () => {
it('filters PHP built-in functions', () => {
expect(isBuiltInOrNoise('echo')).toBe(true);
expect(isBuiltInOrNoise('isset')).toBe(true);
expect(isBuiltInOrNoise('date')).toBe(true);
expect(isBuiltInOrNoise('json_encode')).toBe(true);
expect(isBuiltInOrNoise('array_map')).toBe(true);
expect(php.isBuiltInName('echo')).toBe(true);
expect(php.isBuiltInName('isset')).toBe(true);
expect(php.isBuiltInName('date')).toBe(true);
expect(php.isBuiltInName('json_encode')).toBe(true);
expect(php.isBuiltInName('array_map')).toBe(true);
});
it('filters PHP string functions', () => {
expect(isBuiltInOrNoise('strlen')).toBe(true);
expect(isBuiltInOrNoise('substr')).toBe(true);
expect(isBuiltInOrNoise('str_replace')).toBe(true);
expect(php.isBuiltInName('strlen')).toBe(true);
expect(php.isBuiltInName('substr')).toBe(true);
expect(php.isBuiltInName('str_replace')).toBe(true);
});
});
describe('C/C++', () => {
it('filters standard library functions', () => {
expect(isBuiltInOrNoise('printf')).toBe(true);
expect(isBuiltInOrNoise('malloc')).toBe(true);
expect(isBuiltInOrNoise('free')).toBe(true);
expect(c.isBuiltInName('printf')).toBe(true);
expect(c.isBuiltInName('malloc')).toBe(true);
expect(c.isBuiltInName('free')).toBe(true);
});
it('filters Linux kernel macros', () => {
expect(isBuiltInOrNoise('container_of')).toBe(true);
expect(isBuiltInOrNoise('ARRAY_SIZE')).toBe(true);
expect(isBuiltInOrNoise('pr_info')).toBe(true);
expect(c.isBuiltInName('container_of')).toBe(true);
expect(c.isBuiltInName('ARRAY_SIZE')).toBe(true);
expect(c.isBuiltInName('pr_info')).toBe(true);
});
});
describe('Kotlin', () => {
it('filters stdlib functions', () => {
expect(isBuiltInOrNoise('println')).toBe(true);
expect(isBuiltInOrNoise('listOf')).toBe(true);
expect(isBuiltInOrNoise('TODO')).toBe(true);
expect(kt.isBuiltInName('println')).toBe(true);
expect(kt.isBuiltInName('listOf')).toBe(true);
expect(kt.isBuiltInName('TODO')).toBe(true);
});
it('filters coroutine functions', () => {
expect(isBuiltInOrNoise('launch')).toBe(true);
expect(isBuiltInOrNoise('async')).toBe(true);
expect(kt.isBuiltInName('launch')).toBe(true);
expect(kt.isBuiltInName('async')).toBe(true);
});
});
describe('Swift', () => {
it('filters built-in functions', () => {
expect(isBuiltInOrNoise('print')).toBe(true);
expect(isBuiltInOrNoise('fatalError')).toBe(true);
expect(swift.isBuiltInName('print')).toBe(true);
expect(swift.isBuiltInName('fatalError')).toBe(true);
});
it('filters UIKit methods', () => {
expect(isBuiltInOrNoise('addSubview')).toBe(true);
expect(isBuiltInOrNoise('reloadData')).toBe(true);
expect(swift.isBuiltInName('addSubview')).toBe(true);
expect(swift.isBuiltInName('reloadData')).toBe(true);
});
});
describe('Rust', () => {
it('filters Result/Option methods', () => {
expect(isBuiltInOrNoise('unwrap')).toBe(true);
expect(isBuiltInOrNoise('expect')).toBe(true);
expect(isBuiltInOrNoise('unwrap_or')).toBe(true);
expect(isBuiltInOrNoise('unwrap_or_else')).toBe(true);
expect(isBuiltInOrNoise('unwrap_or_default')).toBe(true);
expect(isBuiltInOrNoise('ok')).toBe(true);
expect(isBuiltInOrNoise('err')).toBe(true);
expect(isBuiltInOrNoise('is_ok')).toBe(true);
expect(isBuiltInOrNoise('is_err')).toBe(true);
expect(isBuiltInOrNoise('map_err')).toBe(true);
expect(isBuiltInOrNoise('and_then')).toBe(true);
expect(isBuiltInOrNoise('or_else')).toBe(true);
expect(rust.isBuiltInName('unwrap')).toBe(true);
expect(rust.isBuiltInName('expect')).toBe(true);
expect(rust.isBuiltInName('unwrap_or')).toBe(true);
expect(rust.isBuiltInName('unwrap_or_else')).toBe(true);
expect(rust.isBuiltInName('unwrap_or_default')).toBe(true);
expect(rust.isBuiltInName('ok')).toBe(true);
expect(rust.isBuiltInName('err')).toBe(true);
expect(rust.isBuiltInName('is_ok')).toBe(true);
expect(rust.isBuiltInName('is_err')).toBe(true);
expect(rust.isBuiltInName('map_err')).toBe(true);
expect(rust.isBuiltInName('and_then')).toBe(true);
expect(rust.isBuiltInName('or_else')).toBe(true);
});
it('filters trait conversion methods', () => {
expect(isBuiltInOrNoise('clone')).toBe(true);
expect(isBuiltInOrNoise('to_string')).toBe(true);
expect(isBuiltInOrNoise('to_owned')).toBe(true);
expect(isBuiltInOrNoise('into')).toBe(true);
expect(isBuiltInOrNoise('from')).toBe(true);
expect(isBuiltInOrNoise('as_ref')).toBe(true);
expect(isBuiltInOrNoise('as_mut')).toBe(true);
expect(rust.isBuiltInName('clone')).toBe(true);
expect(rust.isBuiltInName('to_string')).toBe(true);
expect(rust.isBuiltInName('to_owned')).toBe(true);
expect(rust.isBuiltInName('into')).toBe(true);
expect(rust.isBuiltInName('from')).toBe(true);
expect(rust.isBuiltInName('as_ref')).toBe(true);
expect(rust.isBuiltInName('as_mut')).toBe(true);
});
it('filters iterator methods', () => {
expect(isBuiltInOrNoise('iter')).toBe(true);
expect(isBuiltInOrNoise('into_iter')).toBe(true);
expect(isBuiltInOrNoise('collect')).toBe(true);
expect(isBuiltInOrNoise('fold')).toBe(true);
expect(isBuiltInOrNoise('for_each')).toBe(true);
expect(rust.isBuiltInName('iter')).toBe(true);
expect(rust.isBuiltInName('into_iter')).toBe(true);
expect(rust.isBuiltInName('collect')).toBe(true);
expect(rust.isBuiltInName('fold')).toBe(true);
expect(rust.isBuiltInName('for_each')).toBe(true);
});
it('filters collection methods', () => {
expect(isBuiltInOrNoise('len')).toBe(true);
expect(isBuiltInOrNoise('is_empty')).toBe(true);
expect(isBuiltInOrNoise('push')).toBe(true);
expect(isBuiltInOrNoise('pop')).toBe(true);
expect(isBuiltInOrNoise('insert')).toBe(true);
expect(isBuiltInOrNoise('remove')).toBe(true);
expect(isBuiltInOrNoise('contains')).toBe(true);
expect(rust.isBuiltInName('len')).toBe(true);
expect(rust.isBuiltInName('is_empty')).toBe(true);
expect(rust.isBuiltInName('push')).toBe(true);
expect(rust.isBuiltInName('pop')).toBe(true);
expect(rust.isBuiltInName('insert')).toBe(true);
expect(rust.isBuiltInName('remove')).toBe(true);
expect(rust.isBuiltInName('contains')).toBe(true);
});
it('filters macro-like and panic functions', () => {
expect(isBuiltInOrNoise('format')).toBe(true);
expect(isBuiltInOrNoise('panic')).toBe(true);
expect(isBuiltInOrNoise('unreachable')).toBe(true);
expect(isBuiltInOrNoise('todo')).toBe(true);
expect(isBuiltInOrNoise('unimplemented')).toBe(true);
expect(isBuiltInOrNoise('vec')).toBe(true);
expect(isBuiltInOrNoise('println')).toBe(true);
expect(isBuiltInOrNoise('eprintln')).toBe(true);
expect(isBuiltInOrNoise('dbg')).toBe(true);
expect(rust.isBuiltInName('format')).toBe(true);
expect(rust.isBuiltInName('panic')).toBe(true);
expect(rust.isBuiltInName('unreachable')).toBe(true);
expect(rust.isBuiltInName('todo')).toBe(true);
expect(rust.isBuiltInName('unimplemented')).toBe(true);
expect(rust.isBuiltInName('vec')).toBe(true);
expect(rust.isBuiltInName('println')).toBe(true);
expect(rust.isBuiltInName('eprintln')).toBe(true);
expect(rust.isBuiltInName('dbg')).toBe(true);
});
it('filters sync primitives', () => {
expect(isBuiltInOrNoise('lock')).toBe(true);
expect(isBuiltInOrNoise('try_lock')).toBe(true);
expect(isBuiltInOrNoise('spawn')).toBe(true);
expect(isBuiltInOrNoise('join')).toBe(true);
expect(isBuiltInOrNoise('sleep')).toBe(true);
expect(rust.isBuiltInName('lock')).toBe(true);
expect(rust.isBuiltInName('try_lock')).toBe(true);
expect(rust.isBuiltInName('spawn')).toBe(true);
expect(rust.isBuiltInName('join')).toBe(true);
expect(rust.isBuiltInName('sleep')).toBe(true);
});
it('filters enum constructors', () => {
expect(isBuiltInOrNoise('Some')).toBe(true);
expect(isBuiltInOrNoise('None')).toBe(true);
expect(isBuiltInOrNoise('Ok')).toBe(true);
expect(isBuiltInOrNoise('Err')).toBe(true);
expect(rust.isBuiltInName('Some')).toBe(true);
expect(rust.isBuiltInName('None')).toBe(true);
expect(rust.isBuiltInName('Ok')).toBe(true);
expect(rust.isBuiltInName('Err')).toBe(true);
});
it('does not filter user-defined Rust functions', () => {
expect(isBuiltInOrNoise('process_request')).toBe(false);
expect(isBuiltInOrNoise('handle_connection')).toBe(false);
expect(isBuiltInOrNoise('build_response')).toBe(false);
expect(rust.isBuiltInName('process_request')).toBe(false);
expect(rust.isBuiltInName('handle_connection')).toBe(false);
expect(rust.isBuiltInName('build_response')).toBe(false);
});
});
describe('C#/.NET', () => {
it('filters Console I/O', () => {
expect(isBuiltInOrNoise('Console')).toBe(true);
expect(isBuiltInOrNoise('WriteLine')).toBe(true);
expect(isBuiltInOrNoise('ReadLine')).toBe(true);
expect(cs.isBuiltInName('Console')).toBe(true);
expect(cs.isBuiltInName('WriteLine')).toBe(true);
expect(cs.isBuiltInName('ReadLine')).toBe(true);
});
it('filters LINQ methods', () => {
expect(isBuiltInOrNoise('Where')).toBe(true);
expect(isBuiltInOrNoise('Select')).toBe(true);
expect(isBuiltInOrNoise('GroupBy')).toBe(true);
expect(isBuiltInOrNoise('OrderBy')).toBe(true);
expect(isBuiltInOrNoise('FirstOrDefault')).toBe(true);
expect(isBuiltInOrNoise('ToList')).toBe(true);
expect(cs.isBuiltInName('Where')).toBe(true);
expect(cs.isBuiltInName('Select')).toBe(true);
expect(cs.isBuiltInName('GroupBy')).toBe(true);
expect(cs.isBuiltInName('OrderBy')).toBe(true);
expect(cs.isBuiltInName('FirstOrDefault')).toBe(true);
expect(cs.isBuiltInName('ToList')).toBe(true);
});
it('filters Task async methods', () => {
expect(isBuiltInOrNoise('Task')).toBe(true);
expect(isBuiltInOrNoise('Run')).toBe(true);
expect(isBuiltInOrNoise('WhenAll')).toBe(true);
expect(isBuiltInOrNoise('ConfigureAwait')).toBe(true);
expect(cs.isBuiltInName('Task')).toBe(true);
expect(cs.isBuiltInName('Run')).toBe(true);
expect(cs.isBuiltInName('WhenAll')).toBe(true);
expect(cs.isBuiltInName('ConfigureAwait')).toBe(true);
});
it('filters Object base methods', () => {
expect(isBuiltInOrNoise('ToString')).toBe(true);
expect(isBuiltInOrNoise('GetType')).toBe(true);
expect(isBuiltInOrNoise('Equals')).toBe(true);
expect(isBuiltInOrNoise('GetHashCode')).toBe(true);
expect(cs.isBuiltInName('ToString')).toBe(true);
expect(cs.isBuiltInName('GetType')).toBe(true);
expect(cs.isBuiltInName('Equals')).toBe(true);
expect(cs.isBuiltInName('GetHashCode')).toBe(true);
});
});
describe('user-defined functions', () => {
it('does not filter custom function names', () => {
expect(isBuiltInOrNoise('myCustomFunction')).toBe(false);
expect(isBuiltInOrNoise('processData')).toBe(false);
expect(isBuiltInOrNoise('handleUserRequest')).toBe(false);
expect(js.isBuiltInName('myCustomFunction')).toBe(false);
expect(py.isBuiltInName('processData')).toBe(false);
expect(rust.isBuiltInName('handleUserRequest')).toBe(false);
});
});
});

View file

@ -0,0 +1,52 @@
// ...existing code...
import { describe, it, expect } from 'vitest';
import { isWriteQuery as isWriteQueryAdapter } from '../../src/mcp/core/lbug-adapter';
import { isWriteQuery as isWriteQueryBackend } from '../../src/mcp/local/local-backend';
describe('isWriteQuery regex tests', () => {
const writeQueries = [
'CREATE (n:Test {name: "x"})',
'MATCH (n) SET n.x = 1',
'MERGE (n:Foo {id: 1})',
'DELETE n',
'DROP INDEX ON :Foo(prop)',
'ALTER TABLE Something',
'COPY TO something',
'DETACH DELETE n',
];
const readQueries = [
'MATCH (n:CreateHelpers) RETURN n',
'MATCH (a)-[:CALLS]->(b) RETURN a, b',
'MATCH (f:File)-[r:DEFINES]->(n) RETURN n',
"MATCH (n) WHERE n.name = 'MERGEHelper' RETURN n", // word present as data
'MATCH (n) RETURN n',
'MATCH (n) WHERE n.content CONTAINS ":CREATE" RETURN n',
'MATCH (n:SomethingWithSET) RETURN n',
];
it('adapter isWriteQuery should detect real write queries', () => {
for (const q of writeQueries) {
expect(isWriteQueryAdapter(q), `adapter should detect write for: ${q}`).toBe(true);
}
});
it('adapter isWriteQuery should not false-positive on label/rel or data', () => {
for (const q of readQueries) {
expect(isWriteQueryAdapter(q), `adapter false-positive on: ${q}`).toBe(false);
}
});
it('backend isWriteQuery should detect real write queries', () => {
for (const q of writeQueries) {
expect(isWriteQueryBackend(q), `backend should detect write for: ${q}`).toBe(true);
}
});
it('backend isWriteQuery should not false-positive on label/rel or data', () => {
for (const q of readQueries) {
expect(isWriteQueryBackend(q), `backend false-positive on: ${q}`).toBe(false);
}
});
});

View file

@ -0,0 +1,338 @@
import { describe, it, expect } from 'vitest';
import { parseJcl } from '../../src/core/ingestion/cobol/jcl-parser.js';
import type { JclParseResults } from '../../src/core/ingestion/cobol/jcl-parser.js';
describe('parseJcl', () => {
// ── JOB statements ──────────────────────────────────────────────────
describe('JOB statements', () => {
it('extracts job name', () => {
const jcl = `//MYJOB JOB (ACCT),'MY JOB'`;
const r = parseJcl(jcl, 'test.jcl');
expect(r.jobs).toHaveLength(1);
expect(r.jobs[0].name).toBe('MYJOB');
expect(r.jobs[0].line).toBe(1);
});
it('extracts CLASS and MSGCLASS parameters', () => {
const jcl = `//PAYJOB JOB (ACCT),'PAYROLL',CLASS=A,MSGCLASS=X`;
const r = parseJcl(jcl, 'test.jcl');
expect(r.jobs).toHaveLength(1);
expect(r.jobs[0].name).toBe('PAYJOB');
expect(r.jobs[0].class).toBe('A');
expect(r.jobs[0].msgclass).toBe('X');
});
it('handles job with no CLASS or MSGCLASS', () => {
const jcl = `//BAREJOB JOB (ACCT),'BARE'`;
const r = parseJcl(jcl, 'test.jcl');
expect(r.jobs).toHaveLength(1);
expect(r.jobs[0].name).toBe('BAREJOB');
expect(r.jobs[0].class).toBeUndefined();
expect(r.jobs[0].msgclass).toBeUndefined();
});
});
// ── EXEC statements ─────────────────────────────────────────────────
describe('EXEC statements', () => {
it('extracts step with PGM=program', () => {
const jcl = [
'//MYJOB JOB (ACCT)',
'//STEP1 EXEC PGM=IEFBR14',
].join('\n');
const r = parseJcl(jcl, 'test.jcl');
expect(r.steps).toHaveLength(1);
expect(r.steps[0].name).toBe('STEP1');
expect(r.steps[0].program).toBe('IEFBR14');
expect(r.steps[0].proc).toBeUndefined();
});
it('extracts step with proc name (no PGM= keyword)', () => {
const jcl = [
'//MYJOB JOB (ACCT)',
'//STEP1 EXEC MYPROC',
].join('\n');
const r = parseJcl(jcl, 'test.jcl');
expect(r.steps).toHaveLength(1);
expect(r.steps[0].name).toBe('STEP1');
expect(r.steps[0].program).toBeUndefined();
expect(r.steps[0].proc).toBe('MYPROC');
});
it('associates step with current job', () => {
const jcl = [
'//JOB1 JOB (ACCT)',
'//STEPA EXEC PGM=PROG1',
'//JOB2 JOB (ACCT)',
'//STEPB EXEC PGM=PROG2',
].join('\n');
const r = parseJcl(jcl, 'test.jcl');
expect(r.steps).toHaveLength(2);
expect(r.steps[0].jobName).toBe('JOB1');
expect(r.steps[1].jobName).toBe('JOB2');
});
});
// ── DD statements ───────────────────────────────────────────────────
describe('DD statements', () => {
it('extracts DD name and dataset (DSN=)', () => {
const jcl = [
'//MYJOB JOB (ACCT)',
'//STEP1 EXEC PGM=IEFBR14',
'//INPUT DD DSN=MY.DATA.SET,DISP=SHR',
].join('\n');
const r = parseJcl(jcl, 'test.jcl');
expect(r.ddStatements).toHaveLength(1);
expect(r.ddStatements[0].ddName).toBe('INPUT');
expect(r.ddStatements[0].dataset).toBe('MY.DATA.SET');
});
it('extracts DISP parameter', () => {
const jcl = [
'//MYJOB JOB (ACCT)',
'//STEP1 EXEC PGM=IEFBR14',
'//OUTPUT DD DSN=MY.OUT,DISP=(NEW,CATLG,DELETE)',
].join('\n');
const r = parseJcl(jcl, 'test.jcl');
expect(r.ddStatements).toHaveLength(1);
expect(r.ddStatements[0].disp).toBe('NEW');
});
it('associates DD with current step', () => {
const jcl = [
'//MYJOB JOB (ACCT)',
'//STEP1 EXEC PGM=PROG1',
'//DD1 DD DSN=DS1,DISP=SHR',
'//STEP2 EXEC PGM=PROG2',
'//DD2 DD DSN=DS2,DISP=SHR',
].join('\n');
const r = parseJcl(jcl, 'test.jcl');
expect(r.ddStatements).toHaveLength(2);
expect(r.ddStatements[0].stepName).toBe('STEP1');
expect(r.ddStatements[1].stepName).toBe('STEP2');
});
});
// ── PROC definitions ────────────────────────────────────────────────
describe('PROC definitions', () => {
it('extracts in-stream PROC with name', () => {
const jcl = [
'//MYPROC PROC',
'//STEP1 EXEC PGM=IEFBR14',
'// PEND',
].join('\n');
const r = parseJcl(jcl, 'test.jcl');
expect(r.procs).toHaveLength(1);
expect(r.procs[0].name).toBe('MYPROC');
expect(r.procs[0].isInStream).toBe(true);
});
it('handles PROC/PEND pairs', () => {
const jcl = [
'//PROC1 PROC',
'//S1 EXEC PGM=PROG1',
'// PEND',
'//PROC2 PROC',
'//S2 EXEC PGM=PROG2',
'// PEND',
].join('\n');
const r = parseJcl(jcl, 'test.jcl');
expect(r.procs).toHaveLength(2);
expect(r.procs[0].name).toBe('PROC1');
expect(r.procs[1].name).toBe('PROC2');
});
});
// ── INCLUDE / SET ───────────────────────────────────────────────────
describe('INCLUDE and SET', () => {
it('extracts INCLUDE MEMBER=name', () => {
const jcl = `// INCLUDE MEMBER=MYINCL`;
const r = parseJcl(jcl, 'test.jcl');
expect(r.includes).toHaveLength(1);
expect(r.includes[0].member).toBe('MYINCL');
expect(r.includes[0].line).toBe(1);
});
it('extracts SET variable=value', () => {
const jcl = `// SET ENV=PROD`;
const r = parseJcl(jcl, 'test.jcl');
expect(r.sets).toHaveLength(1);
expect(r.sets[0].variable).toBe('ENV');
expect(r.sets[0].value).toBe('PROD');
});
});
// ── Conditionals ────────────────────────────────────────────────────
describe('Conditionals', () => {
it('extracts IF condition THEN', () => {
const jcl = `// IF STEP1.RC = 0 THEN`;
const r = parseJcl(jcl, 'test.jcl');
expect(r.conditionals).toHaveLength(1);
expect(r.conditionals[0].type).toBe('IF');
expect(r.conditionals[0].condition).toBe('STEP1.RC = 0');
});
it('extracts ELSE and ENDIF', () => {
const jcl = [
'// IF STEP1.RC = 0 THEN',
'//GOOD EXEC PGM=GOODPGM',
'// ELSE',
'//BAD EXEC PGM=BADPGM',
'// ENDIF',
].join('\n');
const r = parseJcl(jcl, 'test.jcl');
expect(r.conditionals).toHaveLength(3);
expect(r.conditionals[0].type).toBe('IF');
expect(r.conditionals[1].type).toBe('ELSE');
expect(r.conditionals[1].condition).toBeUndefined();
expect(r.conditionals[2].type).toBe('ENDIF');
expect(r.conditionals[2].condition).toBeUndefined();
});
});
// ── JCLLIB ──────────────────────────────────────────────────────────
describe('JCLLIB', () => {
it('extracts JCLLIB ORDER=(lib1,lib2)', () => {
const jcl = `// JCLLIB ORDER=(SYS1.PROCLIB,USER.PROCLIB)`;
const r = parseJcl(jcl, 'test.jcl');
expect(r.jcllib).toHaveLength(1);
expect(r.jcllib[0].order).toEqual(['SYS1.PROCLIB', 'USER.PROCLIB']);
expect(r.jcllib[0].line).toBe(1);
});
});
// ── Continuation lines ──────────────────────────────────────────────
describe('Continuation lines', () => {
it('joins continuation lines (col 72 non-blank + next line starts with //)', () => {
// Build a DD line that is exactly 72 chars with non-blank at col 72 (index 71).
// The continuation line provides the DISP parameter.
// "//DD1 DD DSN=MY.VERY.LONG.DATASET.NAME.THAT.KEEPS.GOING," is 60 chars.
// Pad to 71 then add non-blank at col 72.
const base = '//DD1 DD DSN=MY.VERY.LONG.DATASET.NAME.THAT.KEEPS.GOING,';
const padding = ' '.repeat(71 - base.length);
const line1 = base + padding + 'X'; // col 72 is 'X' (non-blank) -> continuation
const line2 = '// DISP=SHR';
const jcl = [
'//MYJOB JOB (ACCT)',
'//STEP1 EXEC PGM=IEFBR14',
line1,
line2,
].join('\n');
const r = parseJcl(jcl, 'test.jcl');
// The continuation should join the DD line so both DSN and DISP are parsed
expect(r.ddStatements).toHaveLength(1);
expect(r.ddStatements[0].ddName).toBe('DD1');
expect(r.ddStatements[0].dataset).toBe('MY.VERY.LONG.DATASET.NAME.THAT.KEEPS.GOING');
expect(r.ddStatements[0].disp).toBe('SHR');
});
});
// ── Edge cases ──────────────────────────────────────────────────────
describe('Edge cases', () => {
it('skips JCL comments (//*)', () => {
const jcl = [
'//* This is a comment',
'//MYJOB JOB (ACCT)',
'//* Another comment',
].join('\n');
const r = parseJcl(jcl, 'test.jcl');
expect(r.jobs).toHaveLength(1);
expect(r.jobs[0].name).toBe('MYJOB');
});
it('skips non-JCL lines', () => {
const jcl = [
'This is not a JCL line',
'//MYJOB JOB (ACCT)',
' Some data',
'//STEP1 EXEC PGM=IEFBR14',
].join('\n');
const r = parseJcl(jcl, 'test.jcl');
expect(r.jobs).toHaveLength(1);
expect(r.steps).toHaveLength(1);
});
it('empty input returns empty results', () => {
const r = parseJcl('', 'test.jcl');
expect(r.jobs).toEqual([]);
expect(r.steps).toEqual([]);
expect(r.ddStatements).toEqual([]);
expect(r.procs).toEqual([]);
expect(r.includes).toEqual([]);
expect(r.sets).toEqual([]);
expect(r.jcllib).toEqual([]);
expect(r.conditionals).toEqual([]);
});
it('complete JCL job with multiple steps and DDs', () => {
const jcl = [
'//* Complete payroll job',
'//PAYJOB JOB (ACCT123),\'PAYROLL RUN\',CLASS=A,MSGCLASS=X',
'// JCLLIB ORDER=(PAY.PROCLIB,SYS1.PROCLIB)',
'// SET ENV=PROD',
'// INCLUDE MEMBER=STDPARMS',
'//*',
'// IF 1 = 1 THEN',
'//STEP01 EXEC PGM=PAYEXT',
'//INPUT DD DSN=PAY.MASTER,DISP=SHR',
'//OUTPUT DD DSN=PAY.EXTRACT,DISP=(NEW,CATLG,DELETE)',
'//SYSPRINT DD SYSOUT=*',
'//*',
'//STEP02 EXEC PAYCALC',
'//INFILE DD DSN=PAY.EXTRACT,DISP=SHR',
'// ELSE',
'//STEP03 EXEC PGM=IEFBR14',
'// ENDIF',
].join('\n');
const r = parseJcl(jcl, 'payroll.jcl');
// Jobs
expect(r.jobs).toHaveLength(1);
expect(r.jobs[0]).toEqual({
name: 'PAYJOB',
line: 2,
class: 'A',
msgclass: 'X',
});
// JCLLIB
expect(r.jcllib).toHaveLength(1);
expect(r.jcllib[0].order).toEqual(['PAY.PROCLIB', 'SYS1.PROCLIB']);
// SET
expect(r.sets).toHaveLength(1);
expect(r.sets[0]).toEqual({ variable: 'ENV', value: 'PROD', line: 4 });
// INCLUDE
expect(r.includes).toHaveLength(1);
expect(r.includes[0].member).toBe('STDPARMS');
// Conditionals
expect(r.conditionals).toHaveLength(3);
expect(r.conditionals[0].type).toBe('IF');
expect(r.conditionals[1].type).toBe('ELSE');
expect(r.conditionals[2].type).toBe('ENDIF');
// Steps
expect(r.steps).toHaveLength(3);
expect(r.steps[0]).toMatchObject({ name: 'STEP01', program: 'PAYEXT', jobName: 'PAYJOB' });
expect(r.steps[1]).toMatchObject({ name: 'STEP02', proc: 'PAYCALC', jobName: 'PAYJOB' });
expect(r.steps[2]).toMatchObject({ name: 'STEP03', program: 'IEFBR14', jobName: 'PAYJOB' });
// DD statements
expect(r.ddStatements).toHaveLength(4);
expect(r.ddStatements[0]).toMatchObject({ ddName: 'INPUT', stepName: 'STEP01', dataset: 'PAY.MASTER', disp: 'SHR' });
expect(r.ddStatements[1]).toMatchObject({ ddName: 'OUTPUT', stepName: 'STEP01', disp: 'NEW' });
expect(r.ddStatements[2]).toMatchObject({ ddName: 'SYSPRINT', stepName: 'STEP01' });
expect(r.ddStatements[3]).toMatchObject({ ddName: 'INFILE', stepName: 'STEP02', dataset: 'PAY.EXTRACT' });
});
});
});

View file

@ -0,0 +1,92 @@
import { describe, it, expect } from 'vitest';
import { getProvider } from '../../src/core/ingestion/languages/index.js';
import { SupportedLanguages } from '../../src/config/supported-languages.js';
const isBuiltIn = (name: string, lang: SupportedLanguages) => getProvider(lang).isBuiltInName(name);
describe('isBuiltInOrNoise (per-language)', () => {
describe('language-specific filtering', () => {
it('filters console for JS but not Python', () => {
expect(isBuiltIn('console', SupportedLanguages.JavaScript)).toBe(true);
expect(isBuiltIn('console', SupportedLanguages.Python)).toBe(false);
});
it('filters println for Kotlin but not Java', () => {
expect(isBuiltIn('println', SupportedLanguages.Kotlin)).toBe(true);
expect(isBuiltIn('println', SupportedLanguages.Java)).toBe(false);
});
it('filters malloc for C but not JavaScript', () => {
expect(isBuiltIn('malloc', SupportedLanguages.C)).toBe(true);
expect(isBuiltIn('malloc', SupportedLanguages.JavaScript)).toBe(false);
});
it('filters setState for Dart but not TypeScript', () => {
expect(isBuiltIn('setState', SupportedLanguages.Dart)).toBe(true);
expect(isBuiltIn('setState', SupportedLanguages.TypeScript)).toBe(false);
});
it('filters unwrap for Rust but not Go', () => {
expect(isBuiltIn('unwrap', SupportedLanguages.Rust)).toBe(true);
expect(isBuiltIn('unwrap', SupportedLanguages.Go)).toBe(false);
});
it('filters puts for Ruby but not PHP', () => {
expect(isBuiltIn('puts', SupportedLanguages.Ruby)).toBe(true);
expect(isBuiltIn('puts', SupportedLanguages.PHP)).toBe(false);
});
it('filters echo for PHP but not Python', () => {
expect(isBuiltIn('echo', SupportedLanguages.PHP)).toBe(true);
expect(isBuiltIn('echo', SupportedLanguages.Python)).toBe(false);
});
it('filters NSLog for Swift but not C', () => {
expect(isBuiltIn('NSLog', SupportedLanguages.Swift)).toBe(true);
expect(isBuiltIn('NSLog', SupportedLanguages.C)).toBe(false);
});
it('filters ToString for C# but not Rust', () => {
expect(isBuiltIn('ToString', SupportedLanguages.CSharp)).toBe(true);
expect(isBuiltIn('ToString', SupportedLanguages.Rust)).toBe(false);
});
});
describe('cross-language pollution eliminated', () => {
it('close is filtered for C# but not C (POSIX)', () => {
expect(isBuiltIn('Close', SupportedLanguages.CSharp)).toBe(true);
expect(isBuiltIn('close', SupportedLanguages.C)).toBe(false);
});
it('then/catch are JS-specific, not filtered for Rust', () => {
expect(isBuiltIn('then', SupportedLanguages.JavaScript)).toBe(true);
expect(isBuiltIn('catch', SupportedLanguages.JavaScript)).toBe(true);
expect(isBuiltIn('then', SupportedLanguages.Rust)).toBe(false);
});
it('emit is Kotlin-specific, not filtered for Java', () => {
expect(isBuiltIn('emit', SupportedLanguages.Kotlin)).toBe(true);
expect(isBuiltIn('emit', SupportedLanguages.Java)).toBe(false);
});
});
describe('languages without builtInNames', () => {
it('Java has no language-specific noise', () => {
expect(isBuiltIn('System', SupportedLanguages.Java)).toBe(false);
expect(isBuiltIn('println', SupportedLanguages.Java)).toBe(false);
});
it('Go has no language-specific noise', () => {
expect(isBuiltIn('fmt', SupportedLanguages.Go)).toBe(false);
expect(isBuiltIn('Println', SupportedLanguages.Go)).toBe(false);
});
});
describe('domain names not filtered', () => {
it('does not filter arbitrary names', () => {
expect(isBuiltIn('processOrder', SupportedLanguages.TypeScript)).toBe(false);
expect(isBuiltIn('UserService', SupportedLanguages.Java)).toBe(false);
expect(isBuiltIn('handle_request', SupportedLanguages.Rust)).toBe(false);
});
});
});

View file

@ -375,27 +375,27 @@ So return-type-aware receiver inference already exists in a constrained downstre
## Language Feature Matrix
| Feature | TS | JS | Java | Kotlin | C# | Go | Rust | Python | PHP | Ruby | Swift | C++ | C |
|---------|:--:|:--:|:----:|:------:|:--:|:--:|:----:|:------:|:---:|:----:|:-----:|:---:|:-:|
| Declarations | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| Parameters | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| Initializer / constructor inference | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| Constructor binding scan | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| For-loop element types | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes††† | Yes | Yes |
| Pattern binding | Yes | Yes | Yes | Yes | No | Yes | Yes | No | No | No | Partial‡‡‡ | No | No |
| Assignment chains | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | No | Yes | Yes | Yes |
| Field/property type resolution | Yes | No† | Yes | Yes | Yes | Yes | Yes | Yes* | Yes | YARD | No | Yes | No‡ |
| Comment-based types | JSDoc | JSDoc | No | No | No | No | No | No | PHPDoc | YARD | No | No | No |
| Return type extraction | JSDoc | JSDoc | No | No | No | No | No | No | PHPDoc | YARD | No | No | No |
| Call-result variable binding | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes¶ | Yes††† | Yes | No |
| Field access binding | Yes | No† | Yes | Yes | Yes | Yes | Yes | No‖ | Yes | N/A | Yes††† | Yes | No |
| Method-call-result binding | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes¶ | Yes††† | Yes | No |
| Write access (ACCESSES write) | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes§ | Yes | Yes | Yes | No |
| Parameter types extracted | Yes** | No | Yes | Yes | Yes | Yes | Yes | Partial†† | No | No | No | Yes | No |
| Method overload disambiguation | Yes** | No | Yes | Yes | Yes | No | No | No | No | No | No | Yes | No |
| Constructor-visible virtual dispatch | Yes | No | Yes | Yes‡‡ | Yes | No | No | No | No | No | No | Yes§§ | No |
| Optional parameter arity resolution | Yes | No | No | Yes | Yes | No | No | Yes | Yes | Yes | No | Yes | No |
| Cross-file binding propagation | Yes | Yes | Yes‖‖ | Yes | Yes¶¶ | Yes*** | Yes | Yes | Partial | Yes*** | Yes*** | Yes*** | Yes*** |
| Feature | TS | JS | Java | Kotlin | C# | Go | Rust | Python | PHP | Ruby | Swift | C++ | C | Dart |
|---------|:--:|:--:|:----:|:------:|:--:|:--:|:----:|:------:|:---:|:----:|:-----:|:---:|:-:|:----:|
| Declarations | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| Parameters | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| Initializer / constructor inference | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| Constructor binding scan | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| For-loop element types | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes††† | Yes | Yes | Yes |
| Pattern binding | Yes | Yes | Yes | Yes | No | Yes | Yes | No | No | No | Partial‡‡‡ | No | No | No |
| Assignment chains | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | No | Yes | Yes | Yes | Yes |
| Field/property type resolution | Yes | No† | Yes | Yes | Yes | Yes | Yes | Yes* | Yes | YARD | No | Yes | No‡ | No |
| Comment-based types | JSDoc | JSDoc | No | No | No | No | No | No | PHPDoc | YARD | No | No | No | No |
| Return type extraction | JSDoc | JSDoc | No | No | No | No | No | No | PHPDoc | YARD | No | No | No | No |
| Call-result variable binding | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes¶ | Yes††† | Yes | No | Yes |
| Field access binding | Yes | No† | Yes | Yes | Yes | Yes | Yes | No‖ | Yes | N/A | Yes††† | Yes | No | Yes |
| Method-call-result binding | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes¶ | Yes††† | Yes | No | Yes |
| Write access (ACCESSES write) | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes§ | Yes | Yes | Yes | No | Yes |
| Parameter types extracted | Yes** | No | Yes | Yes | Yes | Yes | Yes | Partial†† | No | No | No | Yes | No | No |
| Method overload disambiguation | Yes** | No | Yes | Yes | Yes | No | No | No | No | No | No | Yes | No | No |
| Constructor-visible virtual dispatch | Yes | No | Yes | Yes‡‡ | Yes | No | No | No | No | No | No | Yes§§ | No | Yes |
| Optional parameter arity resolution | Yes | No | No | Yes | Yes | No | No | Yes | Yes | Yes | No | Yes | No | No |
| Cross-file binding propagation | Yes | Yes | Yes‖‖ | Yes | Yes¶¶ | Yes*** | Yes | Yes | Partial | Yes*** | Yes*** | Yes*** | Yes*** | Yes*** |
\* Python class-level annotated attributes (`address: Address`) now resolve `declaredType` correctly. The `self.x` instance attribute pattern is not yet supported.