mirror of
https://github.com/abhigyanpatwari/GitNexus.git
synced 2026-10-09 03:17:54 +00:00
Merge remote-tracking branch 'origin/main' into feat/phase8-9-implementation
# Conflicts: # gitnexus/src/core/ingestion/languages/csharp.ts # gitnexus/src/core/ingestion/languages/dart.ts # gitnexus/src/core/ingestion/languages/kotlin.ts # gitnexus/src/core/ingestion/languages/python.ts # gitnexus/src/core/ingestion/languages/ruby.ts # gitnexus/src/core/ingestion/languages/rust.ts # gitnexus/src/core/ingestion/languages/typescript.ts
This commit is contained in:
commit
dddfd0d789
69 changed files with 11315 additions and 391 deletions
|
|
@ -375,6 +375,7 @@ GitNexus builds a complete knowledge graph of your codebase through a multi-phas
|
|||
| Swift | — | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| C | — | — | ✓ | — | ✓ | ✓ | — | ✓ | ✓ |
|
||||
| C++ | — | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
|
||||
| Dart | ✓ | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
|
||||
|
||||
**Imports** — cross-file import resolution · **Named Bindings** — `import { X as Y }` / re-export tracking · **Exports** — public/exported symbol detection · **Heritage** — class inheritance, interfaces, mixins · **Type Annotations** — explicit type extraction for receiver resolution · **Constructor Inference** — infer receiver type from constructor calls (`self`/`this` resolution included for all languages) · **Config** — language toolchain config parsing (tsconfig, go.mod, etc.) · **Frameworks** — AST-based framework pattern detection · **Entry Points** — entry point scoring heuristics
|
||||
|
||||
|
|
@ -539,7 +540,7 @@ The wiki generator reads the indexed graph structure, groups files into modules
|
|||
- [X] Constructor-Inferred Type Resolution, `self`/`this` Receiver Mapping
|
||||
- [X] Wiki Generation, Multi-File Rename, Git-Diff Impact Analysis
|
||||
- [X] Process-Grouped Search, 360-Degree Context, Claude Code Hooks
|
||||
- [X] Multi-Repo MCP, Zero-Config Setup, 13 Language Support
|
||||
- [X] Multi-Repo MCP, Zero-Config Setup, 14 Language Support
|
||||
- [X] Community Detection, Process Detection, Confidence Scoring
|
||||
- [X] Hybrid Search, Vector Index
|
||||
|
||||
|
|
|
|||
100
docs/code-indexing/cobol/README.md
Normal file
100
docs/code-indexing/cobol/README.md
Normal file
|
|
@ -0,0 +1,100 @@
|
|||
# COBOL Code Indexing
|
||||
|
||||
GitNexus indexes COBOL codebases using a **regex-only extraction** strategy, bypassing tree-sitter entirely. This document explains why, how the pipeline works, and links to detailed sub-documents.
|
||||
|
||||
## Why Regex-Only?
|
||||
|
||||
The tree-sitter-cobol grammar (v0.0.1) has three critical limitations that make it unusable for production indexing:
|
||||
|
||||
| Issue | Impact | Severity |
|
||||
|-------|--------|----------|
|
||||
| External scanner hangs on ~5% of files | No timeout mechanism exists for the C scanner; the process blocks indefinitely | **Blocking** |
|
||||
| Only ~15% of paragraph headers detected | Most procedure-division paragraphs are invisible to the grammar | High |
|
||||
| Patch markers in cols 1-6 cause parse errors | Enterprise COBOL uses non-standard sequence area content (e.g., `mzADD`, `estero`, `#FIX`) | High |
|
||||
|
||||
Because the external scanner hang cannot be interrupted (there is no `setTimeoutMicros` equivalent for tree-sitter), using tree-sitter-cobol would hang the indexing pipeline on a non-trivial fraction of real-world files.
|
||||
|
||||
The regex-only approach provides:
|
||||
|
||||
- **Speed**: ~1ms per file average extraction time
|
||||
- **Reliability**: zero hangs, zero crashes across 13,000+ files
|
||||
- **Coverage**: captures all critical symbols -- program name, paragraphs, sections, CALL, PERFORM, COPY, data items (01-77, 88-level), file declarations, FD entries, EXEC SQL/CICS blocks, ENTRY points, and MOVE statements
|
||||
|
||||
## Architecture
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
A[Repository Scan] --> B{File Detection}
|
||||
B -->|Extension match| C[COBOL file]
|
||||
B -->|GITNEXUS_COBOL_DIRS match| C
|
||||
B -->|No match| Z[Skip]
|
||||
|
||||
C --> D{Copybook?}
|
||||
D -->|Yes| E[Add to Copybook Map]
|
||||
D -->|No| F[Source Program]
|
||||
|
||||
E --> G[COPY Expansion Engine]
|
||||
F --> G
|
||||
|
||||
G -->|Inline copybook content| H[Expanded Source]
|
||||
H --> I[Patch Marker Cleanup]
|
||||
I --> J[Regex State Machine]
|
||||
|
||||
J --> K[Extracted Symbols]
|
||||
K --> L[Graph Model Builder]
|
||||
L --> M[Knowledge Graph]
|
||||
|
||||
subgraph "Per-Chunk Processing"
|
||||
G
|
||||
H
|
||||
I
|
||||
J
|
||||
K
|
||||
L
|
||||
end
|
||||
|
||||
subgraph "Post-Processing"
|
||||
M --> N[Community Detection]
|
||||
M --> O[Process Detection]
|
||||
M --> P[Contract Detection]
|
||||
end
|
||||
|
||||
style J fill:#e8f5e9,stroke:#2e7d32
|
||||
style G fill:#e3f2fd,stroke:#1565c0
|
||||
```
|
||||
|
||||
## COBOL vs Tree-Sitter Languages
|
||||
|
||||
| Feature | COBOL (Regex) | Tree-Sitter Languages |
|
||||
|---------|--------------|----------------------|
|
||||
| Parser | Single-pass regex state machine | tree-sitter grammar + queries |
|
||||
| Speed | ~1ms/file | ~5ms/file |
|
||||
| AST available | No | Yes |
|
||||
| COPY expansion | Yes (pre-processing step) | N/A |
|
||||
| Deep indexing | Data items, SQL, CICS, FD, ENTRY | Type annotations, generics, etc. |
|
||||
| Call extraction | PERFORM (intra-file) + CALL (cross-program) | AST-based call site detection |
|
||||
| Import extraction | COPY statements | `import`/`require`/`use`/`#include` |
|
||||
| Coverage | All critical symbols | Language-dependent query coverage |
|
||||
| Failure mode | Never hangs | External scanner can hang (COBOL only) |
|
||||
|
||||
## Sub-Documents
|
||||
|
||||
| Document | Description |
|
||||
|----------|-------------|
|
||||
| [File Detection](./file-detection.md) | Extension mapping, `GITNEXUS_COBOL_DIRS`, copybook classification |
|
||||
| [COPY Expansion](./copy-expansion.md) | Copybook inlining, REPLACING transformations, cycle detection |
|
||||
| [Regex Extraction](./regex-extraction.md) | State machine, regex patterns, line processing |
|
||||
| [Deep Indexing](./deep-indexing.md) | Data items, EXEC SQL/CICS, file declarations, FD, ENTRY, MOVE |
|
||||
| [Graph Model](./graph-model.md) | COBOL-specific node types, edge types, full annotated example |
|
||||
| [Performance](./performance.md) | Benchmarks, worker pool tuning, caps, troubleshooting |
|
||||
|
||||
## Key Source Files
|
||||
|
||||
| File | Purpose |
|
||||
|------|---------|
|
||||
| `gitnexus/src/core/ingestion/cobol-preprocessor.ts` | Patch marker cleanup + regex extraction engine |
|
||||
| `gitnexus/src/core/ingestion/cobol-copy-expander.ts` | COPY statement expansion with REPLACING |
|
||||
| `gitnexus/src/core/ingestion/utils.ts` | `getLanguageFromPath`, `getLanguageFromFilename` |
|
||||
| `gitnexus/src/core/ingestion/pipeline.ts` | `isCobolCopybook`, `expandCobolCopies`, `detectCrossProgamContracts` |
|
||||
| `gitnexus/src/core/ingestion/workers/parse-worker.ts` | `processCobolRegexOnly` -- graph model builder |
|
||||
| `gitnexus/src/core/ingestion/workers/worker-pool.ts` | Configurable sub-batch size for COBOL |
|
||||
157
docs/code-indexing/cobol/copy-expansion.md
Normal file
157
docs/code-indexing/cobol/copy-expansion.md
Normal file
|
|
@ -0,0 +1,157 @@
|
|||
# COBOL COPY Expansion
|
||||
|
||||
The COPY statement is COBOL's include mechanism -- analogous to `#include` in C or `import` in modern languages. GitNexus expands COPY statements **before** regex extraction so that symbols defined inside copybooks (data items, paragraphs, etc.) are visible in the program's extracted graph.
|
||||
|
||||
## Supported Syntax
|
||||
|
||||
### Basic COPY
|
||||
|
||||
```cobol
|
||||
COPY CPSESP.
|
||||
COPY "WORKGRID.CPY".
|
||||
```
|
||||
|
||||
Inlines the content of the named copybook, replacing the COPY line(s).
|
||||
|
||||
### COPY with REPLACING
|
||||
|
||||
```cobol
|
||||
COPY CPSESP REPLACING "ANAZI-KEY" BY "LK-KEY".
|
||||
COPY CPSESP REPLACING LEADING "ESP-" BY "LK-ESP-"
|
||||
LEADING "KPSESPL" BY "LK-KPSESPL".
|
||||
COPY LINKAGE REPLACING TRAILING "-IN" BY "-OUT".
|
||||
```
|
||||
|
||||
Three REPLACING types are supported:
|
||||
|
||||
| Type | Syntax | Behavior | Example |
|
||||
| ------------ | ------------------------------------ | --------------------------------------- | -------------------------------- |
|
||||
| **EXACT** | `REPLACING "OLD" BY "NEW"` | Replace exact identifier matches | `ANAZI-KEY` becomes `LK-KEY` |
|
||||
| **LEADING** | `REPLACING LEADING "PFX-" BY "NEW-"` | Replace prefix on all COBOL identifiers | `ESP-NAME` becomes `LK-ESP-NAME` |
|
||||
| **TRAILING** | `REPLACING TRAILING "-IN" BY "-OUT"` | Replace suffix on all COBOL identifiers | `DATA-IN` becomes `DATA-OUT` |
|
||||
|
||||
Multiple REPLACING clauses can appear in a single COPY statement. They are applied in order to each COBOL identifier in the copybook content.
|
||||
|
||||
### Multi-Line COPY
|
||||
|
||||
COPY statements can span multiple lines (standard COBOL continuation rules apply):
|
||||
|
||||
```cobol
|
||||
COPY CPSESP REPLACING
|
||||
- LEADING "ESP-" BY "LK-ESP-"
|
||||
- LEADING "KPSESPL" BY "LK-KPSESPL".
|
||||
```
|
||||
|
||||
Continuation lines (indicator `-` in column 7) are merged before COPY statement scanning.
|
||||
|
||||
## Expansion Flow
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant Pipeline
|
||||
participant Expander as COPY Expander
|
||||
participant Resolver
|
||||
participant Reader
|
||||
|
||||
Pipeline->>Pipeline: Identify all COBOL files
|
||||
Pipeline->>Pipeline: Classify copybooks vs programs
|
||||
Pipeline->>Reader: Read all copybook content upfront
|
||||
Reader-->>Pipeline: Copybook content map (name -> content)
|
||||
|
||||
loop For each source file in chunk
|
||||
Pipeline->>Expander: expandCopies(content, filePath, resolveFile, readFile)
|
||||
Expander->>Expander: Merge continuation lines
|
||||
Expander->>Expander: Detect COPY statements via regex
|
||||
|
||||
loop For each COPY statement (reverse order)
|
||||
Expander->>Resolver: resolveFile(copyTarget)
|
||||
Resolver-->>Expander: Copybook key or null
|
||||
|
||||
alt Resolved successfully
|
||||
Expander->>Reader: readFile(resolvedKey)
|
||||
Reader-->>Expander: Copybook content
|
||||
|
||||
Expander->>Expander: Apply REPLACING transformations
|
||||
Expander->>Expander: Recurse for nested COPYs (depth + 1)
|
||||
Expander->>Expander: Splice expanded content into output
|
||||
else Not resolved
|
||||
Expander->>Expander: Keep original COPY line
|
||||
end
|
||||
end
|
||||
|
||||
Expander-->>Pipeline: Expanded content + resolution metadata
|
||||
Pipeline->>Pipeline: Replace file content with expanded content
|
||||
end
|
||||
```
|
||||
|
||||
The return type `CopyExpansionResult` contains `expandedContent` and `copyResolutions`. The `expansionDepth` field has been removed from the return type (it was unused by callers).
|
||||
|
||||
COPY statement line numbers in `CopyResolution` are 1-based (consistent with the preprocessor's line numbering). The splice operation that replaces COPY lines with expanded content adjusts for 0-based array indexing internally.
|
||||
|
||||
## Cycle Detection
|
||||
|
||||
Circular COPY references (e.g., copybook A includes copybook B which includes copybook A) are detected and handled:
|
||||
|
||||
1. Each expansion chain maintains a `visited` set of resolved copybook paths
|
||||
2. If a copybook path is already in the visited set, the expansion is skipped
|
||||
3. A `warnedCircular` set (internal to `expandCopies()`, not a parameter) deduplicates warning messages within a single file expansion
|
||||
|
||||
Known circular copybooks in PROJECT-NAME: `ANAZI`, `ANDIP`, `QDIPE` (self-referential includes).
|
||||
|
||||
## Max Depth
|
||||
|
||||
Nested COPY expansion is limited to **10 levels** (`DEFAULT_MAX_DEPTH`). If a COPY chain exceeds this depth, a warning is logged and the remaining COPY statements are left unexpanded.
|
||||
|
||||
## Max Total Expansions
|
||||
|
||||
A breadth amplification guard caps the total number of COPY expansions across all branches within a single file to **500** (`MAX_TOTAL_EXPANSIONS`). This prevents exponential blowup from diamond-shaped COPY graphs where N copybooks each include N other copybooks. Once the limit is reached, further COPY statements in that file are left unexpanded and a single warning is logged.
|
||||
|
||||
## REPLACING Application Detail
|
||||
|
||||
The REPLACING engine works by scanning all COBOL identifiers (matching `\b[A-Z][A-Z0-9-]*\b`) in the copybook content and applying each replacement rule:
|
||||
|
||||
```
|
||||
Original copybook content:
|
||||
05 ESP-NAME PIC X(30).
|
||||
05 ESP-CODE PIC X(10).
|
||||
05 KPSESPL-FLAG PIC X(01).
|
||||
|
||||
After REPLACING LEADING "ESP-" BY "LK-ESP-" LEADING "KPSESPL" BY "LK-KPSESPL":
|
||||
05 LK-ESP-NAME PIC X(30).
|
||||
05 LK-ESP-CODE PIC X(10).
|
||||
05 LK-KPSESPL-FLAG PIC X(01).
|
||||
```
|
||||
|
||||
For LEADING replacements, the engine checks if each identifier starts with the `from` prefix (case-insensitive) and replaces only the prefix portion, preserving the rest of the identifier.
|
||||
|
||||
For TRAILING replacements, the same logic applies to suffixes.
|
||||
|
||||
For EXACT replacements, only identifiers that match the `from` value exactly (case-insensitive) are replaced.
|
||||
|
||||
## Copybook Resolution
|
||||
|
||||
The resolver tries multiple strategies to match a COPY target name to a copybook file:
|
||||
|
||||
1. **Exact match**: `COPY CPSESP` resolves to copybook named `CPSESP`
|
||||
2. **Strip extension**: `COPY WORKGRID.CPY` strips `.CPY` and resolves to `WORKGRID`
|
||||
3. **Add extension**: `COPY CPSESP` tries `CPSESP.CPY` and `CPSESP.COPY`
|
||||
|
||||
If no match is found, the COPY statement is left in place (unexpanded) and a resolution record with `resolvedPath: null` is created.
|
||||
|
||||
## Pipeline Integration
|
||||
|
||||
The expansion runs **per chunk**, after file content is read but before dispatch to worker threads:
|
||||
|
||||
1. All copybook files are read upfront (they are typically small, collectively under 100MB)
|
||||
2. Per chunk, the copybook map is merged with chunk content (in case a chunk contains copybooks)
|
||||
3. Only programs (not copybooks themselves) undergo expansion
|
||||
4. The expanded content replaces the original content in-place before worker dispatch
|
||||
|
||||
## Inline Comment Handling
|
||||
|
||||
The copy expander's `stripInlineComment()` helper is quote-aware: pipe characters (`|`) inside single- or double-quoted strings are preserved. This matches the same quote-aware logic used by the preprocessor.
|
||||
|
||||
## Source Files
|
||||
|
||||
- `gitnexus/src/core/ingestion/cobol-copy-expander.ts` -- `expandCopies()`, `parseReplacingClause()`, `applyReplacing()`
|
||||
- `gitnexus/src/core/ingestion/pipeline.ts` -- `expandCobolCopies()`, copybook map construction, chunk integration
|
||||
312
docs/code-indexing/cobol/deep-indexing.md
Normal file
312
docs/code-indexing/cobol/deep-indexing.md
Normal file
|
|
@ -0,0 +1,312 @@
|
|||
# COBOL Deep Indexing
|
||||
|
||||
Beyond basic symbol extraction (program name, paragraphs, CALL, PERFORM, COPY), GitNexus performs deep indexing of COBOL-specific constructs: data items, EXEC SQL/CICS blocks, file declarations, FD entries, ENTRY points, and MOVE statements.
|
||||
|
||||
## Data Items
|
||||
|
||||
### Level Numbers
|
||||
|
||||
| Level Range | Meaning | Graph Node Type |
|
||||
|-------------|---------|-----------------|
|
||||
| 01 | Record (group item) | `Record` |
|
||||
| 02-49 | Elementary/group items | `Property` |
|
||||
| 66 | RENAMES | `Property` |
|
||||
| 77 | Independent item | `Property` |
|
||||
| 88 | Condition name | `Const` |
|
||||
|
||||
FILLER items are skipped (no useful name for the graph).
|
||||
|
||||
### Clauses Parsed
|
||||
|
||||
The `parseDataItemClauses()` function extracts these clauses from the trailing text of a data item declaration:
|
||||
|
||||
| Clause | Pattern | Example |
|
||||
|--------|---------|---------|
|
||||
| `PIC` / `PICTURE` | `\bPIC(?:TURE)?\s+(?:IS\s+)?(\S+)` | `PIC X(30)`, `PICTURE IS 9(5)V99` |
|
||||
| `USAGE` | `\bUSAGE\s+(?:IS\s+)?(COMP\|BINARY\|...)` | `USAGE IS COMP-3`, `BINARY` |
|
||||
| `REDEFINES` | `\bREDEFINES\s+([A-Z][A-Z0-9-]+)` | `REDEFINES WK-DATE-NUM` |
|
||||
| `OCCURS` | `\bOCCURS\s+(\d+)` | `OCCURS 12 TIMES` |
|
||||
|
||||
Standalone COMP variants (without the `USAGE` keyword) are also detected: `COMP`, `COMP-1` through `COMP-6`, `COMP-X`, `BINARY`, `PACKED-DECIMAL`.
|
||||
|
||||
### Data Hierarchy
|
||||
|
||||
Data items form a hierarchical structure based on level numbers. The extractor uses a **stack algorithm**:
|
||||
|
||||
```
|
||||
Processing order:
|
||||
01 WK-RECORD -> push {01, WK-RECORD} -> parent: Module
|
||||
05 WK-NAME -> push {05, WK-NAME} -> parent: WK-RECORD (01 < 05)
|
||||
10 WK-FIRST -> push {10, WK-FIRST} -> parent: WK-NAME (05 < 10)
|
||||
10 WK-LAST -> pop WK-FIRST, push -> parent: WK-NAME (05 < 10)
|
||||
05 WK-CODE -> pop WK-LAST, WK-NAME -> parent: WK-RECORD (01 < 05)
|
||||
88 WK-ACTIVE -> (88 handled separately) -> parent: WK-CODE
|
||||
```
|
||||
|
||||
The stack maintains items where each entry's level is strictly less than the next. When a new item arrives with a level <= the top of stack, items are popped until the stack top has a smaller level. A `CONTAINS` edge is created from the stack top to the new item.
|
||||
|
||||
For 88-level condition names, the parent is the immediately preceding non-88 data item (found by scanning backwards).
|
||||
|
||||
### Annotated Example
|
||||
|
||||
```cobol
|
||||
01 WK-EMPLOYEE.
|
||||
05 WK-EMP-ID PIC 9(6).
|
||||
05 WK-EMP-NAME PIC X(30).
|
||||
05 WK-EMP-STATUS PIC X(01).
|
||||
88 WK-ACTIVE VALUE "A".
|
||||
88 WK-INACTIVE VALUE "I".
|
||||
05 WK-SALARY PIC 9(7)V99 COMP-3.
|
||||
05 WK-DEPT PIC X(04) OCCURS 3 TIMES.
|
||||
```
|
||||
|
||||
Produces:
|
||||
- `Record` node: `WK-EMPLOYEE` (level 01, section: working-storage)
|
||||
- `Property` nodes: `WK-EMP-ID`, `WK-EMP-NAME`, `WK-EMP-STATUS`, `WK-SALARY`, `WK-DEPT`
|
||||
- `Const` nodes: `WK-ACTIVE` (values: `A`), `WK-INACTIVE` (values: `I`)
|
||||
- `CONTAINS` edges: `WK-EMPLOYEE -> WK-EMP-ID`, `WK-EMPLOYEE -> WK-EMP-NAME`, etc.
|
||||
- `CONTAINS` edges: `WK-EMP-STATUS -> WK-ACTIVE`, `WK-EMP-STATUS -> WK-INACTIVE`
|
||||
|
||||
### Data Item Cap
|
||||
|
||||
A maximum of **500 data items per file** (`MAX_DATA_ITEMS_PER_FILE`) are processed. Some COBOL programs (especially after COPY expansion) can have 10,000+ data items, which would cause graph bloat and push the V8 relationship Map past its 16.7M entry limit across thousands of files.
|
||||
|
||||
The cap applies after extraction: the first 500 items in source order are kept. Since 01-level records appear first, critical top-level structure is preserved.
|
||||
|
||||
## EXEC SQL
|
||||
|
||||
EXEC SQL blocks are accumulated across lines between `EXEC SQL` and `END-EXEC`, then parsed as a unit.
|
||||
|
||||
### Operation Classification
|
||||
|
||||
The first SQL keyword determines the operation:
|
||||
|
||||
| First Keyword | Operation |
|
||||
|---------------|-----------|
|
||||
| `SELECT` | SELECT |
|
||||
| `INSERT` | INSERT |
|
||||
| `UPDATE` | UPDATE |
|
||||
| `DELETE` | DELETE |
|
||||
| `DECLARE` | DECLARE |
|
||||
| `OPEN` | OPEN |
|
||||
| `CLOSE` | CLOSE |
|
||||
| `FETCH` | FETCH |
|
||||
| *(anything else)* | OTHER |
|
||||
|
||||
### Table Extraction
|
||||
|
||||
Tables are extracted from SQL clauses:
|
||||
|
||||
| Clause Pattern | Example |
|
||||
|----------------|---------|
|
||||
| `FROM <table>` | `SELECT * FROM EMPLOYEES` |
|
||||
| `INSERT INTO <table>` | `INSERT INTO EMPLOYEES` |
|
||||
| `UPDATE <table>` | `UPDATE EMPLOYEES SET ...` |
|
||||
| `JOIN <table>` | `LEFT JOIN DEPARTMENTS ON ...` |
|
||||
|
||||
Note: The `INTO` pattern is restricted to `INSERT INTO` to avoid false positives from `FETCH ... INTO :host-var` and `SELECT ... INTO :host-var` statements, where `INTO` introduces host variables rather than table names.
|
||||
|
||||
### Cursor Detection
|
||||
|
||||
```cobol
|
||||
EXEC SQL
|
||||
DECLARE C-EMPLOYEES CURSOR FOR
|
||||
SELECT EMP-ID, EMP-NAME FROM EMPLOYEES
|
||||
WHERE DEPT = :WK-DEPT
|
||||
END-EXEC
|
||||
```
|
||||
|
||||
Extracts: cursor `C-EMPLOYEES`, table `EMPLOYEES`, host variable `WK-DEPT`.
|
||||
|
||||
### Host Variables
|
||||
|
||||
Host variables are COBOL variables referenced in SQL with a `:` prefix. The colon is stripped:
|
||||
|
||||
```sql
|
||||
WHERE EMP-ID = :WK-EMP-ID AND DEPT = :WK-DEPT
|
||||
```
|
||||
|
||||
Extracts: `WK-EMP-ID`, `WK-DEPT`.
|
||||
|
||||
### Graph Output
|
||||
|
||||
- `CodeElement` node per table, with description `sql-table op:{OP}`
|
||||
- `CodeElement` node per cursor, with description `sql-cursor`
|
||||
- `ACCESSES` edge from Module to each CodeElement
|
||||
- Deduplication: if the same table appears in multiple SQL blocks, only one node is created
|
||||
|
||||
## EXEC CICS
|
||||
|
||||
EXEC CICS blocks are accumulated and parsed similarly to SQL blocks.
|
||||
|
||||
### Command Detection
|
||||
|
||||
Two-word commands are detected first (matched against the block start):
|
||||
|
||||
```
|
||||
SEND MAP, RECEIVE MAP, SEND TEXT, SEND CONTROL, READ NEXT, READ PREV
|
||||
```
|
||||
|
||||
If no two-word command matches, the first word is used (e.g., `LINK`, `XCTL`, `RETURN`, `READ`, `WRITE`).
|
||||
|
||||
### Extraction
|
||||
|
||||
| Element | Pattern | Example |
|
||||
|---------|---------|---------|
|
||||
| MAP name | `MAP('name')` or `MAP("name")` | `EXEC CICS SEND MAP('EMPMENU')` |
|
||||
| PROGRAM name | `PROGRAM('name')` or `PROGRAM("name")` | `EXEC CICS LINK PROGRAM('BGTABUP')` |
|
||||
| TRANSID | `TRANSID('name')` or `TRANSID("name")` | `EXEC CICS START TRANSID('EMP1')` |
|
||||
|
||||
### Graph Output
|
||||
|
||||
- MAP: `CodeElement` node with description `cics-map cmd:{CMD}` + `ACCESSES` edge from Module
|
||||
- PROGRAM: `CALLS` edge (cross-program call via CICS LINK/XCTL)
|
||||
- TRANSID: `CodeElement` node with description `cics-transid cmd:{CMD}` + `ACCESSES` edge from Module
|
||||
|
||||
### Annotated Example
|
||||
|
||||
```cobol
|
||||
EXEC CICS
|
||||
SEND MAP('EMPMENU')
|
||||
MAPSET('EMPSET')
|
||||
FROM(WK-MAP-DATA)
|
||||
ERASE
|
||||
END-EXEC
|
||||
```
|
||||
|
||||
Produces:
|
||||
- `CodeElement` node: `EMPMENU` (description: `cics-map cmd:SEND MAP`)
|
||||
- `ACCESSES` edge: Module -> `EMPMENU`
|
||||
|
||||
## File Declarations
|
||||
|
||||
SELECT statements in the INPUT-OUTPUT SECTION are accumulated across multiple lines (until a period terminator) and parsed for:
|
||||
|
||||
| Clause | Pattern | Example |
|
||||
|--------|---------|---------|
|
||||
| SELECT | `SELECT <name>` | `SELECT MASTER-FILE` |
|
||||
| ASSIGN | `ASSIGN TO <file>` | `ASSIGN TO "MASTER.DAT"` |
|
||||
| ORGANIZATION | `ORGANIZATION IS <type>` | `ORGANIZATION IS INDEXED` |
|
||||
| ACCESS | `ACCESS MODE IS <mode>` | `ACCESS MODE IS DYNAMIC` |
|
||||
| RECORD KEY | `RECORD KEY IS <field>` | `RECORD KEY IS WK-EMP-ID` |
|
||||
| FILE STATUS | `FILE STATUS IS <field>` | `FILE STATUS IS WK-FILE-STATUS` |
|
||||
|
||||
### Graph Output
|
||||
|
||||
- `CodeElement` node with description containing all parsed clauses (e.g., `select org:INDEXED access:DYNAMIC key:WK-EMP-ID status:WK-FILE-STATUS assign:MASTER.DAT`)
|
||||
- `RECORD_KEY_OF` edge: from Property node to CodeElement (confidence 0.8)
|
||||
- `FILE_STATUS_OF` edge: from Property node to CodeElement (confidence 0.8)
|
||||
|
||||
## FD Entries
|
||||
|
||||
FD (File Description) entries associate a file name with its record layout:
|
||||
|
||||
```cobol
|
||||
FD MASTER-FILE.
|
||||
01 MASTER-RECORD.
|
||||
05 MR-EMP-ID PIC 9(6).
|
||||
05 MR-EMP-NAME PIC X(30).
|
||||
```
|
||||
|
||||
The extractor tracks `pendingFdName` state: when an `FD` line is seen, the next 01-level data item becomes its record.
|
||||
|
||||
### Graph Output
|
||||
|
||||
- `CodeElement` node with description `fd record:{recordName}`
|
||||
- `CONTAINS` edge: FD CodeElement -> Record node
|
||||
- `CONTAINS` edge: SELECT CodeElement -> FD CodeElement (linking file declaration to file description)
|
||||
|
||||
## ENTRY Points
|
||||
|
||||
The `ENTRY` statement defines additional entry points into a COBOL program (in addition to the main program entry):
|
||||
|
||||
```cobol
|
||||
ENTRY "SUBPROG" USING WK-PARAM-1 WK-PARAM-2.
|
||||
```
|
||||
|
||||
### Graph Output
|
||||
|
||||
- `Constructor` node with description `entry params:{param1},{param2}` (or just `entry` if no parameters)
|
||||
- `CONTAINS` edge: Module -> Constructor
|
||||
- Symbol table entry (so the entry point is discoverable by name)
|
||||
|
||||
## PROCEDURE DIVISION USING
|
||||
|
||||
```cobol
|
||||
PROCEDURE DIVISION USING WK-INPUT-REC WK-OUTPUT-REC.
|
||||
```
|
||||
|
||||
The USING clause identifies parameters received by the program from its caller.
|
||||
|
||||
### Graph Output
|
||||
|
||||
- `RECEIVES` edge: Module -> Property (for each parameter name, confidence 0.8)
|
||||
|
||||
## MOVE Statements
|
||||
|
||||
MOVE statements produce `ACCESSES` edges in the graph:
|
||||
|
||||
```cobol
|
||||
MOVE WK-NAME TO OUT-NAME.
|
||||
MOVE CORRESPONDING WK-INPUT TO WK-OUTPUT.
|
||||
MOVE CORR WK-IN TO WK-OUT.
|
||||
```
|
||||
|
||||
### Extraction Details
|
||||
|
||||
- Source and target identifiers are captured
|
||||
- `CORRESPONDING` and its abbreviation `CORR` are both recognized (bulk field-by-field move)
|
||||
- Figurative constants (SPACES, ZEROS, LOW-VALUES, HIGH-VALUES, QUOTES, ALL) are skipped
|
||||
- The enclosing paragraph (`caller`) is tracked for context
|
||||
|
||||
### MOVE CORRESPONDING / CORR Edge Reasons
|
||||
|
||||
MOVE CORRESPONDING (and CORR) produces distinct edge reasons to differentiate from simple MOVE:
|
||||
|
||||
| Edge | Reason (simple MOVE) | Reason (CORRESPONDING/CORR) |
|
||||
|------|---------------------|-----------------------------|
|
||||
| Read (source) | `cobol-move-read` | `cobol-move-corresponding-read` |
|
||||
| Write (target) | `cobol-move-write` | `cobol-move-corresponding-write` |
|
||||
|
||||
This distinction allows queries to find bulk field-by-field moves separately from simple variable assignments.
|
||||
|
||||
## GO TO DEPENDING ON
|
||||
|
||||
The `GO TO` statement with multiple targets and a `DEPENDING ON` clause is a computed branch:
|
||||
|
||||
```cobol
|
||||
GO TO PARA-1 PARA-2 PARA-3
|
||||
DEPENDING ON WK-SELECTOR.
|
||||
```
|
||||
|
||||
All target paragraph names are extracted and emitted as separate `gotos` entries. Each target produces a `CALLS` edge in the graph (same semantics as PERFORM). The `DEPENDING ON` variable is not currently tracked as a data-flow dependency.
|
||||
|
||||
## SORT INPUT/OUTPUT PROCEDURE
|
||||
|
||||
SORT and MERGE statements can specify procedural entry points instead of file-based I/O:
|
||||
|
||||
```cobol
|
||||
SORT SORT-FILE ON ASCENDING KEY SORT-KEY
|
||||
INPUT PROCEDURE IS PREPARE-INPUT
|
||||
OUTPUT PROCEDURE IS FORMAT-OUTPUT.
|
||||
```
|
||||
|
||||
`INPUT PROCEDURE IS` and `OUTPUT PROCEDURE IS` targets are extracted as control-flow targets (same as PERFORM). They produce `performs` entries and corresponding `CALLS` edges in the graph.
|
||||
|
||||
## Fixed-Format Literal Continuation
|
||||
|
||||
In fixed-format COBOL, string literals can span multiple lines using the continuation indicator (`-` in column 7). When a continuation line starts with a quote character, the extractor joins it with the predecessor by removing the trailing quote from the previous line and the opening quote from the continuation:
|
||||
|
||||
```
|
||||
Line N: MOVE "THIS IS A LONG STRI
|
||||
Line N+1 (cont): - "NG VALUE" TO WK-FIELD.
|
||||
Merged: MOVE "THIS IS A LONG STRING VALUE" TO WK-FIELD.
|
||||
```
|
||||
|
||||
The trailing `"` on line N and the opening `"` on line N+1 are both removed, producing a seamless literal. If no matching quote is found on the predecessor line, the continuation is appended as-is.
|
||||
|
||||
## Source Files
|
||||
|
||||
- `gitnexus/src/core/ingestion/cobol-preprocessor.ts` -- All extraction logic, clause parsers, EXEC block parsers
|
||||
- `gitnexus/src/core/ingestion/workers/parse-worker.ts` -- `processCobolRegexOnly()`, graph node/edge emission
|
||||
- `gitnexus/src/core/ingestion/parsing-processor.ts` -- Sequential fallback with same `MAX_DATA_ITEMS_PER_FILE` cap
|
||||
126
docs/code-indexing/cobol/file-detection.md
Normal file
126
docs/code-indexing/cobol/file-detection.md
Normal file
|
|
@ -0,0 +1,126 @@
|
|||
# COBOL File Detection
|
||||
|
||||
GitNexus detects COBOL files through two mechanisms: extension-based mapping and directory-based override for extensionless files. This document covers both, plus the copybook/program classification logic.
|
||||
|
||||
## Extension Mapping
|
||||
|
||||
### Program Extensions
|
||||
|
||||
| Extension | Type |
|
||||
|-----------|------|
|
||||
| `.cbl` | COBOL program |
|
||||
| `.cob` | COBOL program |
|
||||
| `.cobol` | COBOL program |
|
||||
|
||||
### Copybook Extensions
|
||||
|
||||
| Extension | Type | Notes |
|
||||
|-----------|------|-------|
|
||||
| `.cpy` | Copybook | Standard |
|
||||
| `.copy` | Copybook | Standard |
|
||||
| `.gnm` / `.GNM` | Copybook | Enterprise (GnuCOBOL naming) |
|
||||
| `.fd` / `.FD` | Copybook | File Description fragment |
|
||||
| `.wrk` / `.WRK` | Copybook | Working-Storage fragment |
|
||||
| `.sel` / `.SEL` | Copybook | SELECT clause fragment |
|
||||
| `.open` / `.OPEN` | Copybook | File OPEN fragment |
|
||||
| `.close` / `.CLOSE` | Copybook | File CLOSE fragment |
|
||||
| `.ini` / `.INI` | Copybook | Initialization fragment |
|
||||
| `.def` / `.DEF` | Copybook | Definition fragment |
|
||||
|
||||
All extension matching is case-sensitive in `getLanguageFromFilename` (the extensions above are matched as written, including uppercase variants like `.GNM`).
|
||||
|
||||
## Extensionless File Detection: `GITNEXUS_COBOL_DIRS`
|
||||
|
||||
Many enterprise COBOL repositories use extensionless files -- the filename alone identifies the program (e.g., `s/BGTABFL` is the source for program `BGTABFL`). GitNexus handles this via the `GITNEXUS_COBOL_DIRS` environment variable.
|
||||
|
||||
### Configuration
|
||||
|
||||
Set `GITNEXUS_COBOL_DIRS` to a comma-separated list of directory names:
|
||||
|
||||
```bash
|
||||
# Files in s/, c/, and wfproc/ directories (at any depth) are treated as COBOL
|
||||
export GITNEXUS_COBOL_DIRS=s,c,wfproc
|
||||
```
|
||||
|
||||
The matching is **case-insensitive** and checks all path segments:
|
||||
|
||||
- `/repo/s/BGTABFL` -- matches segment `s` -- COBOL
|
||||
- `/repo/src/c/CPSESP` -- matches segment `c` -- COBOL
|
||||
- `/repo/wfproc/WF001` -- matches segment `wfproc` -- COBOL
|
||||
- `/repo/docs/README` -- no matching segment -- skipped
|
||||
|
||||
### Decision Tree
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
A[getLanguageFromPath] --> B[getLanguageFromFilename]
|
||||
B --> C{Known extension?}
|
||||
C -->|Yes .cbl/.cob/.cobol/.cpy/...| D[Return COBOL]
|
||||
C -->|Yes .ts/.py/.java/...| E[Return other language]
|
||||
C -->|No match| F{Has extension?}
|
||||
|
||||
F -->|"Has dot in basename"| G[Return null]
|
||||
F -->|"No dot = extensionless"| H{GITNEXUS_COBOL_DIRS set?}
|
||||
|
||||
H -->|No| G
|
||||
H -->|Yes| I{Any path segment<br/>matches a configured dir?}
|
||||
|
||||
I -->|Yes| D
|
||||
I -->|No| G
|
||||
|
||||
style D fill:#e8f5e9,stroke:#2e7d32
|
||||
style G fill:#ffebee,stroke:#c62828
|
||||
```
|
||||
|
||||
### Implementation Detail
|
||||
|
||||
The `GITNEXUS_COBOL_DIRS` value is parsed once (on first call) and cached in a `Set<string>`:
|
||||
|
||||
```typescript
|
||||
// From gitnexus/src/core/ingestion/utils.ts
|
||||
const getCobolDirs = (): Set<string> => {
|
||||
if (_cobolDirs) return _cobolDirs;
|
||||
const raw = process.env.GITNEXUS_COBOL_DIRS;
|
||||
_cobolDirs = raw
|
||||
? new Set(raw.split(',').map(d => d.trim().toLowerCase()))
|
||||
: new Set();
|
||||
return _cobolDirs;
|
||||
};
|
||||
```
|
||||
|
||||
The path segment check splits the full path on `/` and tests each segment against the cached set.
|
||||
|
||||
## Copybook vs Program Classification
|
||||
|
||||
After a file is identified as COBOL, it must be classified as either a **program** (to be parsed for symbols) or a **copybook** (to be loaded into the copybook map for COPY expansion).
|
||||
|
||||
### Classification Rules
|
||||
|
||||
A COBOL file is classified as a **copybook** if ANY of these conditions is true:
|
||||
|
||||
1. It has a recognized copybook extension (`.cpy`, `.copy`, `.gnm`, `.fd`, `.wrk`, `.sel`, `.open`, `.close`, `.ini`, `.def`)
|
||||
2. It is an extensionless file whose path contains a directory segment matching one of: `c`, `copy`, `copybooks`, `copylib`, `cpy`
|
||||
|
||||
A file is classified as a **program** if:
|
||||
|
||||
1. It has a program extension (`.cbl`, `.cob`, `.cobol`), OR
|
||||
2. It is extensionless and does NOT match any copybook directory pattern
|
||||
|
||||
### Copybook Name Resolution
|
||||
|
||||
Copybook names are derived from the filename:
|
||||
|
||||
- Strip the extension (if any)
|
||||
- Convert to uppercase
|
||||
|
||||
Examples:
|
||||
- `c/CPSESP` -- name: `CPSESP`
|
||||
- `copy/workgrid.cpy` -- name: `WORKGRID`
|
||||
- `c/ANAZI.GNM` -- name: `ANAZI`
|
||||
|
||||
This name is used to resolve `COPY CPSESP.` statements during expansion.
|
||||
|
||||
## Source Files
|
||||
|
||||
- `gitnexus/src/core/ingestion/utils.ts` -- `getLanguageFromPath()`, `getLanguageFromFilename()`, `getCobolDirs()`
|
||||
- `gitnexus/src/core/ingestion/pipeline.ts` -- `isCobolCopybook()`, `getCopybookName()`, `COPYBOOK_EXTENSIONS`, `COBOL_PROGRAM_EXTENSIONS`
|
||||
193
docs/code-indexing/cobol/graph-model.md
Normal file
193
docs/code-indexing/cobol/graph-model.md
Normal file
|
|
@ -0,0 +1,193 @@
|
|||
# COBOL Graph Model
|
||||
|
||||
This document describes the graph nodes and edges that GitNexus creates for COBOL codebases. The COBOL graph model is richer than most tree-sitter languages because it captures domain-specific constructs: file declarations, FD entries, data hierarchies, SQL tables, CICS maps, and cross-program contracts.
|
||||
|
||||
## Entity-Relationship Diagram
|
||||
|
||||
```mermaid
|
||||
erDiagram
|
||||
File ||--o{ Module : DEFINES
|
||||
File ||--o{ Function : DEFINES
|
||||
File ||--o{ Namespace : DEFINES
|
||||
File ||--o{ Record : DEFINES
|
||||
File ||--o{ Property : DEFINES
|
||||
File ||--o{ Const : DEFINES
|
||||
File ||--o{ CodeElement : DEFINES
|
||||
File ||--o{ Constructor : DEFINES
|
||||
File }o--o{ File : IMPORTS
|
||||
|
||||
Module ||--o{ Record : CONTAINS
|
||||
Module ||--o{ Constructor : CONTAINS
|
||||
Module }o--o{ CodeElement : ACCESSES
|
||||
Module }o--o{ Module : CALLS
|
||||
Module }o--o{ Module : CONTRACTS
|
||||
Module }o--o{ Property : RECEIVES
|
||||
|
||||
Record ||--o{ Property : CONTAINS
|
||||
Record ||--o{ Const : CONTAINS
|
||||
Record }o--o{ Record : REDEFINES
|
||||
|
||||
Property ||--o{ Property : CONTAINS
|
||||
Property ||--o{ Const : CONTAINS
|
||||
Property }o--o{ Property : REDEFINES
|
||||
Property }o--o{ CodeElement : RECORD_KEY_OF
|
||||
Property }o--o{ CodeElement : FILE_STATUS_OF
|
||||
|
||||
CodeElement ||--o{ CodeElement : CONTAINS
|
||||
CodeElement ||--o{ Record : CONTAINS
|
||||
|
||||
Function }o--o{ Function : CALLS
|
||||
```
|
||||
|
||||
## Node Types
|
||||
|
||||
| Node Type | COBOL Concept | Created From | Example |
|
||||
|-----------|--------------|--------------|---------|
|
||||
| `Module` | PROGRAM-ID | `PROGRAM-ID. BGTABFL` | Name: `BGTABFL`, description may include author and date |
|
||||
| `Function` | Paragraph | `PROCESS-RECORD.` at column 8 | Name: `PROCESS-RECORD` |
|
||||
| `Namespace` | Procedure section | `MAIN-LOGIC SECTION.` at column 8 | Name: `MAIN-LOGIC` |
|
||||
| `Record` | 01-level data item | `01 WK-EMPLOYEE.` | Description: `level:01 section:working-storage` |
|
||||
| `Property` | 02-49/66/77 data item | `05 WK-NAME PIC X(30).` | Description: `level:05 pic:X(30) section:working-storage` |
|
||||
| `Const` | 88-level condition | `88 WK-ACTIVE VALUE "A".` | Description: `level:88 values:A` |
|
||||
| `CodeElement` | SELECT, FD, SQL table, CICS map, cursor, transid | Various | Description varies by subtype |
|
||||
| `Constructor` | ENTRY point | `ENTRY "SUBPROG" USING WK-DATA` | Description: `entry params:WK-DATA` |
|
||||
|
||||
### CodeElement Subtypes
|
||||
|
||||
CodeElement is used for multiple COBOL constructs, distinguished by their description prefix:
|
||||
|
||||
| Subtype | ID Pattern | Description Format | Example |
|
||||
|---------|-----------|-------------------|---------|
|
||||
| File SELECT | `CodeElement:{path}:SELECT:{name}` | `select org:INDEXED access:DYNAMIC ...` | `SELECT MASTER-FILE` |
|
||||
| FD entry | `CodeElement:{path}:FD:{name}` | `fd record:{recordName}` | `FD MASTER-FILE` |
|
||||
| SQL table | `CodeElement:{path}:sql-table:{name}` | `sql-table op:SELECT` | Table `EMPLOYEES` |
|
||||
| SQL cursor | `CodeElement:{path}:sql-cursor:{name}` | `sql-cursor` | Cursor `C-EMPLOYEES` |
|
||||
| CICS map | `CodeElement:{path}:cics-map:{name}` | `cics-map cmd:SEND MAP` | Map `EMPMENU` |
|
||||
| CICS transid | `CodeElement:{path}:cics-transid:{name}` | `cics-transid cmd:START` | Transid `EMP1` |
|
||||
|
||||
## Edge Types
|
||||
|
||||
| Edge Type | Source | Target | Created By | Confidence | Example |
|
||||
|-----------|--------|--------|-----------|------------|---------|
|
||||
| `DEFINES` | File | any node | File defines its symbols | 1.0 | File -> Module `BGTABFL` |
|
||||
| `CALLS` | Function | Function | `PERFORM X [THRU Y]` | (via call-processor) | `PROCESS-RECORD` -> `CALC-TAX` |
|
||||
| `CALLS` | Module | Module | `CALL "BGTABUP"` | (via call-processor) | `BGTABFL` -> `BGTABUP` |
|
||||
| `CALLS` | Module | Module | `EXEC CICS LINK PROGRAM('X')` | (via call-processor) | `BGTABFL` -> `BGTABUP` |
|
||||
| `IMPORTS` | File | File | `COPY copybook` | (via import-processor) | Source file -> Copybook file |
|
||||
| `CONTAINS` | Module | Record | Data hierarchy root | 1.0 | `BGTABFL` -> `WK-EMPLOYEE` |
|
||||
| `CONTAINS` | Record | Property | Data hierarchy | 1.0 | `WK-EMPLOYEE` -> `WK-NAME` |
|
||||
| `CONTAINS` | Property | Property | Nested data items | 1.0 | `WK-ADDRESS` -> `WK-CITY` |
|
||||
| `CONTAINS` | Record/Property | Const | 88-level parent | 1.0 | `WK-STATUS` -> `WK-ACTIVE` |
|
||||
| `CONTAINS` | CodeElement (FD) | Record | FD record link | 1.0 | `FD:MASTER-FILE` -> `MASTER-RECORD` |
|
||||
| `CONTAINS` | CodeElement (SELECT) | CodeElement (FD) | SELECT-FD link | 0.9 | `SELECT:MASTER-FILE` -> `FD:MASTER-FILE` |
|
||||
| `CONTAINS` | Module | Constructor | ENTRY in module | 1.0 | `BGTABFL` -> `SUBPROG` |
|
||||
| `REDEFINES` | Record | Record | `01 X REDEFINES Y` | 1.0 | `WK-DATE-NUM` -> `WK-DATE-ALPHA` |
|
||||
| `REDEFINES` | Property | Property | `05 X REDEFINES Y` | 1.0 | `WK-CODE-NUM` -> `WK-CODE-ALPHA` |
|
||||
| `RECORD_KEY_OF` | Property | CodeElement (SELECT) | `RECORD KEY IS field` | 0.8 | `WK-EMP-ID` -> `SELECT:MASTER-FILE` |
|
||||
| `FILE_STATUS_OF` | Property | CodeElement (SELECT) | `FILE STATUS IS field` | 0.8 | `WK-FS` -> `SELECT:MASTER-FILE` |
|
||||
| `ACCESSES` | Module | CodeElement | EXEC SQL/CICS | 0.9 | `BGTABFL` -> `sql-table:EMPLOYEES` |
|
||||
| `RECEIVES` | Module | Property | `PROCEDURE USING` | 0.8 | `BGTABFL` -> `WK-INPUT-REC` |
|
||||
| `CONTRACTS` | Module | Module | Shared copybook detection | 0.9 | `BGTABFL` -> `BGTABUP` (via `CPSESP`) |
|
||||
|
||||
## Full Annotated Example
|
||||
|
||||
Given this COBOL program:
|
||||
|
||||
```cobol
|
||||
IDENTIFICATION DIVISION.
|
||||
PROGRAM-ID. EMPMAINT.
|
||||
AUTHOR. Development Team.
|
||||
|
||||
ENVIRONMENT DIVISION.
|
||||
INPUT-OUTPUT SECTION.
|
||||
FILE-CONTROL.
|
||||
SELECT EMP-FILE
|
||||
ASSIGN TO "EMPLOYEE.DAT"
|
||||
ORGANIZATION IS INDEXED
|
||||
ACCESS MODE IS DYNAMIC
|
||||
RECORD KEY IS EMP-ID
|
||||
FILE STATUS IS WS-FILE-STATUS.
|
||||
|
||||
DATA DIVISION.
|
||||
FILE SECTION.
|
||||
FD EMP-FILE.
|
||||
01 EMP-RECORD.
|
||||
05 EMP-ID PIC 9(6).
|
||||
05 EMP-NAME PIC X(30).
|
||||
|
||||
WORKING-STORAGE SECTION.
|
||||
01 WS-FLAGS.
|
||||
05 WS-FILE-STATUS PIC X(02).
|
||||
05 WS-EOF-FLAG PIC X(01).
|
||||
88 WS-EOF VALUE "Y".
|
||||
|
||||
LINKAGE SECTION.
|
||||
01 LK-SEARCH-KEY PIC 9(6).
|
||||
|
||||
PROCEDURE DIVISION USING LK-SEARCH-KEY.
|
||||
MAIN-LOGIC SECTION.
|
||||
MAIN-START.
|
||||
PERFORM OPEN-FILE
|
||||
PERFORM PROCESS-RECORDS
|
||||
PERFORM CLOSE-FILE
|
||||
STOP RUN.
|
||||
|
||||
OPEN-FILE.
|
||||
OPEN I-O EMP-FILE.
|
||||
|
||||
PROCESS-RECORDS.
|
||||
MOVE LK-SEARCH-KEY TO EMP-ID
|
||||
EXEC SQL
|
||||
SELECT EMP_SALARY INTO :WS-SALARY
|
||||
FROM EMPLOYEES
|
||||
WHERE EMP_ID = :EMP-ID
|
||||
END-EXEC
|
||||
CALL "EMPREPORT".
|
||||
|
||||
CLOSE-FILE.
|
||||
CLOSE EMP-FILE.
|
||||
```
|
||||
|
||||
The graph produced contains:
|
||||
|
||||
**Nodes:**
|
||||
- `Module`: EMPMAINT (description: `author:Development Team`)
|
||||
- `Namespace`: MAIN-LOGIC
|
||||
- `Function`: MAIN-START, OPEN-FILE, PROCESS-RECORDS, CLOSE-FILE
|
||||
- `Record`: EMP-RECORD, WS-FLAGS, LK-SEARCH-KEY
|
||||
- `Property`: EMP-ID, EMP-NAME, WS-FILE-STATUS, WS-EOF-FLAG
|
||||
- `Const`: WS-EOF (values: Y)
|
||||
- `CodeElement`: SELECT:EMP-FILE, FD:EMP-FILE, sql-table:EMPLOYEES
|
||||
- (COPY imports, if any, would produce File IMPORTS edges)
|
||||
|
||||
**Edges:**
|
||||
- `DEFINES`: File -> all nodes
|
||||
- `CONTAINS`: EMPMAINT -> EMP-RECORD, EMPMAINT -> WS-FLAGS, EMPMAINT -> LK-SEARCH-KEY
|
||||
- `CONTAINS`: EMP-RECORD -> EMP-ID, EMP-RECORD -> EMP-NAME
|
||||
- `CONTAINS`: WS-FLAGS -> WS-FILE-STATUS, WS-FLAGS -> WS-EOF-FLAG
|
||||
- `CONTAINS`: WS-EOF-FLAG -> WS-EOF
|
||||
- `CONTAINS`: FD:EMP-FILE -> EMP-RECORD
|
||||
- `CONTAINS`: SELECT:EMP-FILE -> FD:EMP-FILE
|
||||
- `CALLS`: MAIN-START -> OPEN-FILE, MAIN-START -> PROCESS-RECORDS, MAIN-START -> CLOSE-FILE
|
||||
- `CALLS`: EMPMAINT -> EMPREPORT (external CALL)
|
||||
- `ACCESSES`: EMPMAINT -> sql-table:EMPLOYEES
|
||||
- `RECEIVES`: EMPMAINT -> LK-SEARCH-KEY (PROCEDURE USING)
|
||||
- `RECORD_KEY_OF`: EMP-ID -> SELECT:EMP-FILE
|
||||
- `FILE_STATUS_OF`: WS-FILE-STATUS -> SELECT:EMP-FILE
|
||||
|
||||
## How COBOL Differs from Tree-Sitter Languages
|
||||
|
||||
| Aspect | COBOL | Tree-Sitter Languages |
|
||||
|--------|-------|----------------------|
|
||||
| Node variety | 8 types (Module, Function, Namespace, Record, Property, Const, CodeElement, Constructor) | Typically 4-6 (Function, Class, Method, Interface, Module, Const) |
|
||||
| Domain edges | RECORD_KEY_OF, FILE_STATUS_OF, ACCESSES, RECEIVES, CONTRACTS, REDEFINES | Primarily CALLS, IMPORTS, EXTENDS, IMPLEMENTS |
|
||||
| Data hierarchy | Deep CONTAINS chains (01 -> 05 -> 10 -> 88) | Flat class members |
|
||||
| Cross-program calls | CALL "name" + CICS LINK PROGRAM | Import-based resolution |
|
||||
| Contract detection | Shared COPY copybook between caller/callee | Not applicable |
|
||||
| Metadata | AUTHOR, DATE-WRITTEN on Module | JSDoc/docstring (not indexed) |
|
||||
|
||||
## Source Files
|
||||
|
||||
- `gitnexus/src/core/ingestion/workers/parse-worker.ts` -- `processCobolRegexOnly()`, node/edge emission logic
|
||||
- `gitnexus/src/core/ingestion/pipeline.ts` -- `detectCrossProgamContracts()` for CONTRACTS edges
|
||||
- `gitnexus/src/core/ingestion/cobol-preprocessor.ts` -- `CobolRegexResults` interface (all extracted data)
|
||||
261
docs/code-indexing/cobol/performance.md
Normal file
261
docs/code-indexing/cobol/performance.md
Normal file
|
|
@ -0,0 +1,261 @@
|
|||
# COBOL Performance and Tuning
|
||||
|
||||
This document covers real-world benchmarks, worker pool configuration, memory management, known limitations, and troubleshooting for COBOL indexing.
|
||||
|
||||
## PROJECT-NAME Benchmark
|
||||
|
||||
The PROJECT-NAME project is a large Italian payroll system written in COBOL. It serves as the primary benchmark for COBOL indexing performance.
|
||||
|
||||
### Input
|
||||
|
||||
| Metric | Value |
|
||||
| --------------------------- | ---------------------------------------------------------------------------- |
|
||||
| Paths scanned | 14,217 |
|
||||
| Parseable files | 13,129 |
|
||||
| Total source size | 224 MB |
|
||||
| Chunks | 12 (at 20 MB budget) |
|
||||
| Copybooks loaded | 2,976 |
|
||||
| Copybooks used in expansion | 2,955 |
|
||||
| Key directories | `s/` (7773 programs), `c/` (3036 copybooks), `wfproc/` (1973 workflow files) |
|
||||
|
||||
### Output
|
||||
|
||||
| Metric | Value |
|
||||
| ---------------------- | ------ |
|
||||
| Graph nodes | 2.79M |
|
||||
| Graph edges | 5.67M |
|
||||
| Clusters (communities) | 16,679 |
|
||||
| Execution flows | 300 |
|
||||
|
||||
### Timing
|
||||
|
||||
| Phase | Duration |
|
||||
| ------------------------------- | ----------------- |
|
||||
| Total | ~251s |
|
||||
| KuzuDB write | 132s |
|
||||
| Full-text search indexing | 6.7s |
|
||||
| Regex extraction (avg per file) | ~1ms |
|
||||
| COPY expansion + deep indexing | Remainder (~112s) |
|
||||
|
||||
### Indexing Command
|
||||
|
||||
```bash
|
||||
cd /path/to/PROJECT-NAME
|
||||
GITNEXUS_COBOL_DIRS=s,c,wfproc GITNEXUS_VERBOSE=1 node --max-old-space-size=8192 \
|
||||
/path/to/gitnexus/dist/cli/index.js analyze --force
|
||||
```
|
||||
|
||||
## Open-Source Benchmarks
|
||||
|
||||
### CardDemo (AWS)
|
||||
|
||||
| Metric | Value |
|
||||
| ------ | ----- |
|
||||
| Graph nodes | 12,323 |
|
||||
| Graph edges | 8,893 |
|
||||
| Total time | 7.4s |
|
||||
|
||||
### ACAS
|
||||
|
||||
| Metric | Value |
|
||||
| ------ | ----- |
|
||||
| Graph nodes | 14,016 |
|
||||
| Graph edges | 15,452 |
|
||||
| Total time | 9.3s |
|
||||
|
||||
### Micro-Benchmark (Single-File Extraction)
|
||||
|
||||
| Metric | Value |
|
||||
| ------ | ----- |
|
||||
| Per-iteration | 0.65ms |
|
||||
| Throughput | ~382K lines/sec |
|
||||
|
||||
## Worker Pool Tuning
|
||||
|
||||
### Sub-Batch Size
|
||||
|
||||
The worker pool splits each worker's chunk into sub-batches to bound peak memory per `postMessage` serialization. COBOL repos use a smaller sub-batch size than the default:
|
||||
|
||||
| Parameter | Default | COBOL Mode |
|
||||
| --------------------- | ----------- | ------------------- |
|
||||
| Sub-batch size | 1,500 files | 200 files |
|
||||
| Per sub-batch timeout | 120s | 120s (configurable) |
|
||||
|
||||
**Why 200?** COBOL regex extraction + preprocessing takes ~1ms per file on average, but with COPY expansion and deep indexing the effective time is ~150ms per file. At sub-batch size 1500, that would be ~225s per sub-batch, exceeding the 120s timeout.
|
||||
|
||||
COBOL mode is activated automatically when `GITNEXUS_COBOL_DIRS` is set:
|
||||
|
||||
```typescript
|
||||
// From pipeline.ts
|
||||
const cobolSubBatch = process.env.GITNEXUS_COBOL_DIRS ? 200 : undefined;
|
||||
workerPool = createWorkerPool(workerUrl, undefined, cobolSubBatch);
|
||||
```
|
||||
|
||||
### Worker Count
|
||||
|
||||
Workers default to `min(8, cpus - 1)`. For COBOL repos, this is usually sufficient since regex extraction is CPU-bound but fast. The bottleneck is typically KuzuDB write, not extraction.
|
||||
|
||||
### Timeout Configuration
|
||||
|
||||
| Environment Variable | Default | Purpose |
|
||||
| ------------------------------------ | --------------- | --------------------------------------------------- |
|
||||
| `GITNEXUS_WORKER_TIMEOUT_MS` | 120,000 (2 min) | Per sub-batch processing timeout |
|
||||
| `GITNEXUS_WORKER_STARTUP_TIMEOUT_MS` | 60,000 (1 min) | Worker initialization timeout (tree-sitter loading) |
|
||||
|
||||
For COBOL-only repos, worker startup is faster because tree-sitter native modules are loaded lazily (skipped entirely if only COBOL files are present).
|
||||
|
||||
## Data Item Cap
|
||||
|
||||
### Configuration
|
||||
|
||||
```typescript
|
||||
const MAX_DATA_ITEMS_PER_FILE = 500;
|
||||
```
|
||||
|
||||
This constant appears in both `parse-worker.ts` (worker path) and `parsing-processor.ts` (sequential fallback).
|
||||
|
||||
### Rationale
|
||||
|
||||
Some COBOL programs, especially after COPY expansion, can have 10,000+ data items. At that scale:
|
||||
|
||||
- The in-memory relationship Map (for CONTAINS, REDEFINES, etc.) approaches the V8 16.7M entry limit across thousands of files
|
||||
- KuzuDB write time increases linearly with edge count
|
||||
- Most deep-nested items (level 20+) are rarely queried individually
|
||||
|
||||
### Impact
|
||||
|
||||
The cap truncates data items beyond the 500th in source order. Since 01-level Records appear first in COBOL source, the cap preserves:
|
||||
|
||||
- All 01-level record definitions
|
||||
- The most important 02-49 level items (those closest to the record root)
|
||||
- 88-level conditions associated with early items
|
||||
|
||||
To increase the cap for specific needs, modify the `MAX_DATA_ITEMS_PER_FILE` constant in both files.
|
||||
|
||||
## Memory Management
|
||||
|
||||
### COPY Expansion Breadth Guard
|
||||
|
||||
A per-file `MAX_TOTAL_EXPANSIONS = 500` limit prevents exponential blowup from diamond-shaped COPY graphs (e.g., N copybooks each containing N COPY statements). Once the limit is reached, further COPY statements in that file are left unexpanded. See [copy-expansion.md](copy-expansion.md) for details.
|
||||
|
||||
### COPY Expansion Memory
|
||||
|
||||
All copybook content is loaded upfront into a Map before chunk processing begins. For PROJECT-NAME:
|
||||
|
||||
- 2,976 copybooks, typically under 100MB total
|
||||
- The Map is shared (read-only) across chunk iterations
|
||||
- Per-chunk, the copybook map is merged with chunk file content (in case a chunk contains copybooks not in the pre-loaded set)
|
||||
- After all chunks are processed, the copybook map is freed (`cobolCopybookContents = undefined`)
|
||||
|
||||
### Chunk Budget
|
||||
|
||||
Source files are grouped into chunks of max 20MB (`CHUNK_BYTE_BUDGET`). Each chunk's lifecycle:
|
||||
|
||||
1. Read file content into memory
|
||||
2. Expand COPY statements (mutates content in-place)
|
||||
3. Dispatch to workers for extraction
|
||||
4. Workers return serialized results
|
||||
5. Merge results into graph
|
||||
6. Chunk content goes out of scope (GC reclaims)
|
||||
|
||||
This ensures only ~20MB of source + ~200-400MB of working memory (ASTs, extracted records, serialization) is active at any time.
|
||||
|
||||
### Shared Warning Deduplication
|
||||
|
||||
The `warnedCircular` set (used by the COPY expansion engine) is shared across all files in a chunk. This prevents the same circular copybook warning (e.g., `ANAZI includes itself`) from being logged thousands of times.
|
||||
|
||||
## Known Limitations
|
||||
|
||||
| Limitation | Impact | Workaround |
|
||||
| ---------------------------------------- | --------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------- |
|
||||
| tree-sitter-cobol hangs on ~5% of files | Cannot use tree-sitter for COBOL | Regex-only extraction (current approach) |
|
||||
| Data item cap (500/file) | May miss deeply nested items in large programs | Increase `MAX_DATA_ITEMS_PER_FILE` in source |
|
||||
| Circular copybooks (ANAZI, ANDIP, QDIPE) | Self-referential includes cannot be expanded | Detected and skipped with warning |
|
||||
| wfproc/ files may not be pure COBOL | Workflow files may produce extraction noise | Exclude `wfproc` from `GITNEXUS_COBOL_DIRS` if problematic |
|
||||
| No MOVE DATA_FLOW edges yet | Data flow between variables not in graph | Reserved for future release |
|
||||
| Continuation line handling | Some complex multi-line continuations (especially in string literals spanning 3+ lines) may not merge correctly | Known edge case; affects <0.1% of lines |
|
||||
| Single-line EXEC blocks | `EXEC SQL SELECT ... END-EXEC` on one line is handled, but pathological nesting is not | Extremely rare in practice |
|
||||
| Extension case sensitivity | `.GNM` and `.gnm` are matched differently | Use the exact case from the codebase |
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### "COPY expansion failed"
|
||||
|
||||
```
|
||||
[pipeline] COPY expansion failed for s/BGTABFL: Cannot read properties of null
|
||||
```
|
||||
|
||||
**Cause:** A copybook referenced by a COPY statement cannot be found.
|
||||
|
||||
**Fix:**
|
||||
|
||||
1. Verify `GITNEXUS_COBOL_DIRS` includes the directory containing copybooks (typically `c`)
|
||||
2. Check that copybook filenames match the COPY target (case-insensitive, after stripping extensions)
|
||||
3. Ensure copybook files are not in `.gitignore`
|
||||
|
||||
### Worker sub-batch timeout
|
||||
|
||||
```
|
||||
Worker 3 sub-batch timed out after 120s (chunk: 200 items)
|
||||
```
|
||||
|
||||
**Cause:** A sub-batch took longer than the timeout. Typically happens when one file is extremely large (50,000+ lines after COPY expansion).
|
||||
|
||||
**Fix:** Increase the timeout:
|
||||
|
||||
```bash
|
||||
GITNEXUS_WORKER_TIMEOUT_MS=300000 gitnexus analyze
|
||||
```
|
||||
|
||||
### Memory errors (heap out of memory)
|
||||
|
||||
```
|
||||
FATAL ERROR: CALL_AND_RETRY_LAST Allocation failed - JavaScript heap out of memory
|
||||
```
|
||||
|
||||
**Fix:** Increase Node.js heap size:
|
||||
|
||||
```bash
|
||||
node --max-old-space-size=16384 /path/to/gitnexus/dist/cli/index.js analyze
|
||||
```
|
||||
|
||||
For very large repos (>500MB source), consider `--max-old-space-size=32768`.
|
||||
|
||||
### Concurrent analyze corruption
|
||||
|
||||
**Rule:** Only ONE `gitnexus analyze` process should run at a time per repository. Concurrent writes to KuzuDB corrupt the database.
|
||||
|
||||
If corruption occurs:
|
||||
|
||||
```bash
|
||||
# Remove the KuzuDB directory and re-index
|
||||
rm -rf .gitnexus/kuzu
|
||||
gitnexus analyze --force
|
||||
```
|
||||
|
||||
### Slow KuzuDB write phase
|
||||
|
||||
The KuzuDB write phase (132s for PROJECT-NAME) is the bottleneck for large COBOL repos. This is proportional to the number of nodes and edges being written. Reducing `MAX_DATA_ITEMS_PER_FILE` or excluding non-essential directories from `GITNEXUS_COBOL_DIRS` can help.
|
||||
|
||||
### Verbose output
|
||||
|
||||
Enable verbose logging to see per-phase timing and statistics:
|
||||
|
||||
```bash
|
||||
GITNEXUS_VERBOSE=1 gitnexus analyze
|
||||
```
|
||||
|
||||
This outputs:
|
||||
|
||||
- Scan statistics (paths, parseable files, chunk count)
|
||||
- Worker pool configuration (worker count, sub-batch size)
|
||||
- COPY expansion statistics (copybooks loaded, files expanded)
|
||||
- Community and process detection results
|
||||
- Contract detection results
|
||||
|
||||
## Source Files
|
||||
|
||||
- `gitnexus/src/core/ingestion/workers/worker-pool.ts` -- `DEFAULT_SUB_BATCH_SIZE`, `SUB_BATCH_TIMEOUT_MS`, `WORKER_STARTUP_TIMEOUT_MS`
|
||||
- `gitnexus/src/core/ingestion/pipeline.ts` -- `CHUNK_BYTE_BUDGET`, COBOL sub-batch configuration, chunk lifecycle
|
||||
- `gitnexus/src/core/ingestion/workers/parse-worker.ts` -- `MAX_DATA_ITEMS_PER_FILE`, `processCobolRegexOnly()`
|
||||
- `gitnexus/src/core/ingestion/parsing-processor.ts` -- Sequential fallback `MAX_DATA_ITEMS_PER_FILE`
|
||||
206
docs/code-indexing/cobol/regex-extraction.md
Normal file
206
docs/code-indexing/cobol/regex-extraction.md
Normal file
|
|
@ -0,0 +1,206 @@
|
|||
# COBOL Regex Extraction
|
||||
|
||||
The `extractCobolSymbolsWithRegex()` function in `cobol-preprocessor.ts` performs single-pass, state-machine-driven extraction of all COBOL symbols. This document describes the state machine, line processing flow, and every regex pattern used.
|
||||
|
||||
## State Machine: Division Tracking
|
||||
|
||||
The extractor tracks which COBOL division is currently being processed. Division transitions are detected by the `RE_DIVISION` pattern.
|
||||
|
||||
```mermaid
|
||||
stateDiagram-v2
|
||||
[*] --> null : Start of file
|
||||
null --> identification : IDENTIFICATION DIVISION
|
||||
identification --> environment : ENVIRONMENT DIVISION
|
||||
environment --> data : DATA DIVISION
|
||||
data --> procedure : PROCEDURE DIVISION
|
||||
|
||||
note right of identification
|
||||
Extracts: PROGRAM-ID, AUTHOR, DATE-WRITTEN
|
||||
end note
|
||||
note right of environment
|
||||
Extracts: SELECT ... ASSIGN ... (file declarations)
|
||||
end note
|
||||
note right of data
|
||||
Extracts: FD entries, data items (01-77, 88), COPY
|
||||
end note
|
||||
note right of procedure
|
||||
Extracts: paragraphs, sections, PERFORM, CALL,
|
||||
ENTRY, MOVE, EXEC SQL/CICS
|
||||
end note
|
||||
```
|
||||
|
||||
## State Machine: Data Section Tracking
|
||||
|
||||
Within the DATA DIVISION, a secondary state machine tracks the current section to tag data items with their origin.
|
||||
|
||||
```mermaid
|
||||
stateDiagram-v2
|
||||
[*] --> unknown : DATA DIVISION entered
|
||||
unknown --> working_storage : WORKING-STORAGE SECTION
|
||||
unknown --> linkage : LINKAGE SECTION
|
||||
unknown --> file : FILE SECTION
|
||||
unknown --> local_storage : LOCAL-STORAGE SECTION
|
||||
working_storage --> linkage : LINKAGE SECTION
|
||||
working_storage --> file : FILE SECTION
|
||||
linkage --> working_storage : WORKING-STORAGE SECTION
|
||||
file --> working_storage : WORKING-STORAGE SECTION
|
||||
file --> linkage : LINKAGE SECTION
|
||||
local_storage --> working_storage : WORKING-STORAGE SECTION
|
||||
```
|
||||
|
||||
Within the ENVIRONMENT DIVISION, the `currentEnvSection` tracks whether we are in `INPUT-OUTPUT` or `CONFIGURATION` section. SELECT statement accumulation only occurs in `INPUT-OUTPUT`.
|
||||
|
||||
## Line Processing Flow
|
||||
|
||||
Each raw source line goes through this pipeline:
|
||||
|
||||
```
|
||||
Raw line
|
||||
|
|
||||
v
|
||||
Length < 7? ---------> Skip (flush pending if any)
|
||||
|
|
||||
v
|
||||
Indicator col 7
|
||||
|
|
||||
+-- '*' or '/' -----> Comment: skip entirely
|
||||
|
|
||||
+-- '-' ------------> Continuation: append to pending line
|
||||
|
|
||||
+-- other ----------> Normal: flush pending, strip inline comments (|),
|
||||
buffer as new pending logical line
|
||||
```
|
||||
|
||||
After all lines are processed, the final pending line is flushed, along with any accumulated SELECT statement, SORT/MERGE accumulator, and any open EXEC block (truncated file without `END-EXEC`).
|
||||
|
||||
### Inline Comment Stripping
|
||||
|
||||
Enterprise COBOL (particularly Italian dialect) uses the pipe character `|` as an inline comment marker. The `stripInlineComment()` helper is **quote-aware**: it tracks whether the scan position is inside a single- or double-quoted string and only treats `|` as a comment marker when outside quotes. Pipe characters inside string literals are preserved.
|
||||
|
||||
Free-format `*>` inline comment stripping uses the same quote-aware approach: the scanner walks character by character, toggling quote state, and only recognizes `*>` as a comment marker when not inside a quoted string.
|
||||
|
||||
### Patch Marker Handling
|
||||
|
||||
The `preprocessCobolSource()` function (run before extraction in the worker) replaces non-standard content in columns 1-6. Standard COBOL expects spaces or digit sequence numbers in this area. If any letter or `#` character is found, the entire sequence area is replaced with 6 spaces:
|
||||
|
||||
```
|
||||
Before: mzADD MOVE WK-AMT TO WK-TOTAL
|
||||
After: MOVE WK-AMT TO WK-TOTAL
|
||||
```
|
||||
|
||||
This preserves exact line count for position mapping.
|
||||
|
||||
## Regex Pattern Reference
|
||||
|
||||
All patterns are compiled once as module-level constants and reused across calls.
|
||||
|
||||
### Division and Section Detection
|
||||
|
||||
| Constant | Pattern | Purpose | Example Match |
|
||||
|----------|---------|---------|---------------|
|
||||
| `RE_DIVISION` | `\b(IDENTIFICATION\|ENVIRONMENT\|DATA\|PROCEDURE)\s+DIVISION\b` | Division boundary | `PROCEDURE DIVISION` |
|
||||
| `RE_SECTION` | `\b(WORKING-STORAGE\|LINKAGE\|FILE\|LOCAL-STORAGE\|INPUT-OUTPUT\|CONFIGURATION)\s+SECTION\b` | Section boundary | `WORKING-STORAGE SECTION` |
|
||||
|
||||
### IDENTIFICATION DIVISION
|
||||
|
||||
| Constant | Pattern | Purpose | Example Match |
|
||||
|----------|---------|---------|---------------|
|
||||
| `RE_PROGRAM_ID` | `\bPROGRAM-ID\.\s*([A-Z][A-Z0-9-]*)` | Program name | `PROGRAM-ID. BGTABFL` |
|
||||
| `RE_AUTHOR` | `^\s+AUTHOR\.\s*(.+)` | Author metadata | `AUTHOR. D. Smith` |
|
||||
| `RE_DATE_WRITTEN` | `^\s+DATE-WRITTEN\.\s*(.+)` | Date metadata | `DATE-WRITTEN. 2024-01-15` |
|
||||
|
||||
### ENVIRONMENT DIVISION
|
||||
|
||||
| Constant | Pattern | Purpose | Example Match |
|
||||
|----------|---------|---------|---------------|
|
||||
| `RE_SELECT_START` | `\bSELECT\s+(?:OPTIONAL\s+)?([A-Z][A-Z0-9-]+)` | File SELECT start (with optional `SELECT OPTIONAL` support) | `SELECT MASTER-FILE`, `SELECT OPTIONAL TRANS-FILE` |
|
||||
|
||||
SELECT statements are accumulated across multiple lines until a period terminator is found, then parsed for ASSIGN, ORGANIZATION, ACCESS, RECORD KEY, and FILE STATUS clauses.
|
||||
|
||||
### DATA DIVISION
|
||||
|
||||
| Constant | Pattern | Purpose | Example Match |
|
||||
|----------|---------|---------|---------------|
|
||||
| `RE_FD` | `^\s+FD\s+([A-Z][A-Z0-9-]+)` | File description | `FD MASTER-FILE` |
|
||||
| `RE_DATA_ITEM` | `^\s+(\d{1,2})\s+([A-Z][A-Z0-9-]+)\s*(.*)` | Data item (01-77) | `05 WK-NAME PIC X(30)` |
|
||||
| `RE_ANONYMOUS_REDEFINES` | `^\s+(\d{1,2})\s+REDEFINES\s+([A-Z][A-Z0-9-]+)` | Anonymous REDEFINES | `01 REDEFINES WK-REC` |
|
||||
| `RE_88_LEVEL` | `^\s+88\s+([A-Z][A-Z0-9-]+)\s+VALUES?\s+(?:ARE\s+)?(.+)` | Condition name | `88 WK-ACTIVE VALUE "Y"` |
|
||||
|
||||
The trailing clauses of `RE_DATA_ITEM` are parsed by `parseDataItemClauses()` for PIC, USAGE, OCCURS, and REDEFINES.
|
||||
|
||||
### PROCEDURE DIVISION
|
||||
|
||||
| Constant | Pattern | Purpose | Example Match |
|
||||
|----------|---------|---------|---------------|
|
||||
| `RE_PROC_SECTION` | `^ ([A-Z][A-Z0-9-]+)\s+SECTION\.\s*$` | Procedure section header | ` MAIN-LOGIC SECTION.` |
|
||||
| `RE_PROC_PARAGRAPH` | `^ ([A-Z][A-Z0-9-]+)\.\s*$` | Paragraph header | ` PROCESS-RECORD.` |
|
||||
| `RE_PERFORM` | `\bPERFORM\s+([A-Z][A-Z0-9-]+)(?:\s+THRU\s+([A-Z][A-Z0-9-]+))?` | PERFORM call | `PERFORM CALC-TAX THRU CALC-TAX-EXIT` |
|
||||
| `RE_PROC_USING` | `\bPROCEDURE\s+DIVISION\s+USING\s+([\s\S]*?)(?:\.\|$)` | USING parameters | `PROCEDURE DIVISION USING WK-PARAM` |
|
||||
| `RE_ENTRY` | `\bENTRY\s+"([^"]+)"(?:\s+USING\s+([\s\S]*?))?(?:\.\|$)` | ENTRY point | `ENTRY "SUBPROG" USING WK-DATA` |
|
||||
| `RE_MOVE` | `\bMOVE\s+((?:CORRESPONDING\|CORR)\s+)?([A-Z][A-Z0-9-]+)\s+TO\s+(.+)` | MOVE statement (supports CORR abbreviation and multi-target) | `MOVE WK-NAME TO OUT-NAME`, `MOVE CORR WK-IN TO WK-OUT` |
|
||||
|
||||
The USING parameter list (`RE_PROC_USING`) is split on `\bRETURNING\b` before tokenization -- any RETURNING clause and everything after it is excluded from the parameter list (`.split(/\bRETURNING\b/i)[0]`).
|
||||
|
||||
Note: `RE_PROC_SECTION` and `RE_PROC_PARAGRAPH` require exactly 7 spaces of leading indentation (COBOL area A starting at column 8). This is the standard COBOL paragraph indentation.
|
||||
|
||||
### All-Division Patterns
|
||||
|
||||
These patterns are checked regardless of current division:
|
||||
|
||||
| Constant | Pattern | Purpose | Example Match |
|
||||
|----------|---------|---------|---------------|
|
||||
| `RE_CALL` | `\bCALL\s+"([^"]+)"` | External program call | `CALL "BGTABUP"` |
|
||||
| `RE_COPY_UNQUOTED` | `\bCOPY\s+([A-Z][A-Z0-9-]+)(?:\s\|\.)` | COPY (unquoted) | `COPY CPSESP.` |
|
||||
| `RE_COPY_QUOTED` | `\bCOPY\s+"([^"]+)"(?:\s\|\.)` | COPY (quoted) | `COPY "WORKGRID.CPY".` |
|
||||
|
||||
### SORT/MERGE Support
|
||||
|
||||
| Constant | Purpose |
|
||||
|----------|---------|
|
||||
| `SORT_CLAUSE_NOISE` | Set of SORT/MERGE clause keywords filtered from USING/GIVING file lists: `ON`, `ASCENDING`, `DESCENDING`, `KEY`, `WITH`, `DUPLICATES`, `IN`, `ORDER`, `COLLATING`, `SEQUENCE`, `IS`, `THROUGH`, `THRU`, `INPUT`, `OUTPUT`, `PROCEDURE` |
|
||||
|
||||
SORT and MERGE statements are accumulated across multiple lines (like SELECT) until a period terminator is found, then parsed for USING/GIVING file lists and INPUT/OUTPUT PROCEDURE targets. The `flushSort()` helper encapsulates the flush-and-parse logic, mirroring the existing `flushSelect()` pattern. Both helpers are called at EOF to handle truncated files.
|
||||
|
||||
### GO TO Multi-Target
|
||||
|
||||
`RE_GOTO` captures all paragraph names in a `GO TO` statement, including the multi-target form `GO TO p1 p2 p3 DEPENDING ON x`. The captured group contains all target names (space-separated), which are split into individual targets. Each target produces a separate `gotos` entry.
|
||||
|
||||
### PROGRAM-ID Detection
|
||||
|
||||
PROGRAM-ID is detected regardless of the current division state. This handles sibling programs that appear after `END PROGRAM` and omit the `IDENTIFICATION DIVISION` header -- the extractor will still capture the PROGRAM-ID and push a new program boundary.
|
||||
|
||||
### EXEC Block Patterns
|
||||
|
||||
| Constant | Pattern | Purpose | Example Match |
|
||||
|----------|---------|---------|---------------|
|
||||
| `RE_EXEC_SQL_START` | `\bEXEC\s+SQL\b` | Start of EXEC SQL block | `EXEC SQL` |
|
||||
| `RE_EXEC_CICS_START` | `\bEXEC\s+CICS\b` | Start of EXEC CICS block | `EXEC CICS` |
|
||||
| `RE_END_EXEC` | `\bEND-EXEC\b` | End of EXEC block | `END-EXEC` |
|
||||
|
||||
EXEC blocks accumulate all lines between `EXEC SQL/CICS` and `END-EXEC`, then delegate to `parseExecSqlBlock()` or `parseExecCicsBlock()` for detailed extraction.
|
||||
|
||||
## Excluded Paragraph Names
|
||||
|
||||
The following names are excluded from paragraph detection to avoid false positives from division/section headers:
|
||||
|
||||
```
|
||||
DECLARATIVES, END, PROCEDURE, IDENTIFICATION,
|
||||
ENVIRONMENT, DATA, WORKING-STORAGE, LINKAGE,
|
||||
FILE, LOCAL-STORAGE, COMMUNICATION, REPORT,
|
||||
SCREEN, INPUT-OUTPUT, CONFIGURATION
|
||||
```
|
||||
|
||||
Additionally, paragraph candidates containing `DIVISION` or `SECTION` as substrings are excluded.
|
||||
|
||||
## MOVE Skip List (Figurative Constants)
|
||||
|
||||
MOVE statements where the source is a figurative constant are skipped:
|
||||
|
||||
```
|
||||
SPACES, ZEROS, ZEROES, LOW-VALUES, LOW-VALUE,
|
||||
HIGH-VALUES, HIGH-VALUE, QUOTES, QUOTE, ALL
|
||||
```
|
||||
|
||||
## Source Files
|
||||
|
||||
- `gitnexus/src/core/ingestion/cobol-preprocessor.ts` -- `preprocessCobolSource()`, `extractCobolSymbolsWithRegex()`, all regex constants
|
||||
326
docs/plans/2026-03-26-feat-cobol-full-language-coverage-plan.md
Normal file
326
docs/plans/2026-03-26-feat-cobol-full-language-coverage-plan.md
Normal file
|
|
@ -0,0 +1,326 @@
|
|||
---
|
||||
title: "feat: Complete COBOL language feature coverage for maximum knowledge graph value"
|
||||
type: feat
|
||||
status: active
|
||||
date: 2026-03-26
|
||||
origin: Feature audit from v3-integration-architect agent (session 8642401e)
|
||||
---
|
||||
|
||||
## Enhancement Summary
|
||||
|
||||
**Deepened on:** 2026-03-26
|
||||
**Research agents used:** COBOL expert (Phase 1+2), graph value analyst, codebase explorer
|
||||
**Sections enhanced:** Phase 1 (5 features), Phase 2 (4 features), graph value ranking
|
||||
|
||||
### Key Improvements from Research
|
||||
1. **CALL USING** is the #1 highest-value edge type (9.2/10) — fixes ~40% of missing caller references
|
||||
2. **EXEC DLI** requires dual-interface support (EXEC DLI + CBLTDLI CALL) for full IMS coverage
|
||||
3. **DECLARATIVES** is lowest-risk Phase 2 item — existing section/paragraph detection already captures structure
|
||||
4. **SET TO TRUE** accounts for 80-90% of all SET statements — prioritize this form
|
||||
5. **INSPECT** needs multi-line accumulator (like SORT) — can span 5+ continuation lines
|
||||
6. **Graph value ranking**: cobol-call-using (9.2) > cobol-error-handler (9.0) > dli-gu (8.2) > cobol-string (6.2)
|
||||
|
||||
### New Edge Cases Discovered
|
||||
- CALL USING supports mixed modes: `USING BY REFERENCE WS-A BY CONTENT WS-B BY VALUE WS-C`
|
||||
- CALL USING `ADDRESS OF` and `OMITTED` must be filtered from parameter lists
|
||||
- EXEC DLI can have multiple SEGMENT levels in hierarchical retrieval (use matchAll)
|
||||
- DECLARATIVES can have multiple USE sections (one per file + catch-all for INPUT/OUTPUT/I-O/EXTEND)
|
||||
- INSPECT TALLYING can have multiple counters in a single statement
|
||||
- STRING/UNSTRING can span multiple lines (need accumulator pattern)
|
||||
|
||||
---
|
||||
|
||||
# Complete COBOL Language Feature Coverage
|
||||
|
||||
## Overview
|
||||
|
||||
Implement the remaining 25 unhandled COBOL language features and fix 10 partial features to achieve ~95% coverage (up from 71.9%). The goal is to build the richest possible knowledge graph from COBOL codebases, enabling a future `modernize` MCP command (out of scope for this plan) that would use the graph to assist with COBOL-to-modern-language migration.
|
||||
|
||||
## Problem Statement
|
||||
|
||||
The COBOL processor currently handles 54 of 89 applicable language features (71.9%). The 25 unhandled features represent real data loss in the knowledge graph:
|
||||
- **Cross-program data flow** is invisible (CALL ... USING parameters not extracted)
|
||||
- **IMS/DB programs** produce empty graphs (EXEC DLI not recognized)
|
||||
- **String transformation logic** is invisible (STRING/UNSTRING/INSPECT not tracked)
|
||||
- **SQL copybook dependencies** are missing (EXEC SQL INCLUDE not mapped)
|
||||
- **Error handling flows** are lost (DECLARATIVES/USE AFTER not captured)
|
||||
|
||||
## Proposed Solution
|
||||
|
||||
Implement features in 4 phases, ordered by graph value density (edges created per LOC of implementation). Each phase is independently shippable and testable.
|
||||
|
||||
## Technical Approach
|
||||
|
||||
### Phase 1: High-Value Data Flow Edges (~150 LOC, ~8 new edge types)
|
||||
|
||||
The highest-ROI features: they create new ACCESSES and IMPORTS edges that directly improve impact analysis.
|
||||
|
||||
**Critical research finding**: Multi-line statement accumulation is the dominant challenge. CALL USING, STRING/UNSTRING, and multi-line data item clauses all span multiple lines in production COBOL. The free-format path processes each line independently — these features need statement accumulators (like SORT/SELECT) or the free-format path needs multi-line awareness. Estimated LOC increased from 110 to 150 to account for accumulator infrastructure.
|
||||
|
||||
#### 1.1 EXEC SQL INCLUDE -> IMPORTS edges
|
||||
- **File:** `cobol-preprocessor.ts` (parseExecSqlBlock)
|
||||
- **What:** Detect `INCLUDE` as the operation, extract member name, emit as a `copies[]` entry
|
||||
- **Graph:** IMPORTS edge from File to included copybook/SQLCA with reason `sql-include`
|
||||
- **Tests:** Unit test for `EXEC SQL INCLUDE SQLCA END-EXEC` and `EXEC SQL INCLUDE CUSTCOPY END-EXEC`
|
||||
|
||||
**Research insights (EXEC SQL INCLUDE):**
|
||||
- DB2 member names can contain underscores: `EXEC SQL INCLUDE CUST_TBL_DCL END-EXEC` — regex must use `[A-Z][A-Z0-9_-]+`
|
||||
- Quoted literal form: `EXEC SQL INCLUDE 'DBRMLIB.MEMBER' END-EXEC` (z/OS PDS qualified name)
|
||||
- SQLCA/SQLDA are DB2 builtins — won't resolve to repo files. Emit unresolved IMPORTS edge (still valuable)
|
||||
- No REPLACING support on EXEC SQL INCLUDE (unlike COPY)
|
||||
- Add `INCLUDE` to `OP_MAP` in `parseExecSqlBlock`; extract member via `RE_SQL_INCLUDE = /^INCLUDE\s+(?:'([^']+)'|"([^"]+)"|([A-Z][A-Z0-9_-]+))/i`
|
||||
|
||||
#### 1.2 CALL ... USING parameter extraction -> ACCESSES edges (Graph value: 9.2/10)
|
||||
- **File:** `cobol-preprocessor.ts` (processLogicalLine CALL section)
|
||||
- **What:** After capturing CALL target, scan for USING clause. Extract parameter names (reuse USING_KEYWORDS filter). Store as `calls[].parameters: string[]`
|
||||
- **Interface:** Add `parameters?: string[]` to calls array type in CobolRegexResults
|
||||
- **File:** `cobol-processor.ts` (CALL edge block)
|
||||
- **Graph:** For each USING parameter, create ACCESSES edge from caller to data item Property node with reason `cobol-call-using`
|
||||
- **Tests:** `CALL 'AUDITLOG' USING CUST-ID WS-AMOUNT` -> 2 ACCESSES edges
|
||||
|
||||
**Research insights (CALL USING forms):**
|
||||
- Mixed modes: `CALL 'PGM' USING BY REFERENCE WS-A BY CONTENT WS-B BY VALUE WS-C`
|
||||
- Pointer passing: `CALL 'PGM' USING ADDRESS OF WS-A`
|
||||
- Placeholder: `CALL 'PGM' USING OMITTED WS-B`
|
||||
- Filter keywords: add `ADDRESS`, `OMITTED`, `LENGTH` to USING_KEYWORDS (already has BY/VALUE/REFERENCE/CONTENT)
|
||||
- **Impact tool enhancement:** CALL-USING edges enable BFS traversal through parameter data flow — single most impactful edge type for COBOL impact analysis
|
||||
|
||||
#### 1.3 STRING/UNSTRING data flow -> ACCESSES edges
|
||||
- **File:** `cobol-preprocessor.ts` (new section in extractProcedure)
|
||||
- **What:** Accumulate multi-line STRING/UNSTRING until period or END-STRING/END-UNSTRING. Extract sources and INTO targets.
|
||||
- **Interface:** Add `strings: Array<{ sources: string[]; target: string; type: 'string' | 'unstring'; line: number; caller: string | null }>` to CobolRegexResults
|
||||
- **Graph:** read-ACCESSES on sources, write-ACCESSES on INTO target with reason `cobol-string-read` / `cobol-string-write`
|
||||
- **Tests:** 2 unit tests + integration test assertions
|
||||
|
||||
**Research insights (STRING/UNSTRING):**
|
||||
- **Needs statement accumulator** — STRING/UNSTRING always span multiple lines in production
|
||||
- Terminate accumulation at: period, END-STRING/END-UNSTRING, or start of next COBOL verb
|
||||
- STRING sources: identifiers before each `DELIMITED BY`. Filter: STRING, DELIMITED, BY, SIZE, ALL, INTO, WITH, POINTER, ON, OVERFLOW, NOT, END-STRING
|
||||
- UNSTRING: source is first identifier after UNSTRING; INTO targets are identifiers after INTO. Filter: DELIMITER, IN, COUNT, TALLYING, OR
|
||||
- WITH POINTER field is both read AND written (starting position updated)
|
||||
- TALLYING IN / COUNT IN fields are write targets
|
||||
- Literal sources (`'text'`) must be filtered — quote-aware tokenization needed
|
||||
- **Edge case**: STRING terminated by next verb, not period — existing fixture has `STRING ... DISPLAY` without period between them
|
||||
|
||||
#### 1.4 OCCURS DEPENDING ON -> ACCESSES edge
|
||||
- **File:** `cobol-preprocessor.ts` (parseDataItemClauses)
|
||||
- **What:** Extend OCCURS regex to capture DEPENDING ON field, KEY fields, and INDEXED BY names
|
||||
- **Interface:** Add `dependingOn?: string`, `occursMax?: number`, `occursKeys?: Array<{direction: string; fields: string[]}>`, `indexedBy?: string[]` to data items
|
||||
- **Graph:** ACCESSES edge from table item to controlling field with reason `cobol-depends-on`
|
||||
- **Tests:** `05 WS-TABLE OCCURS 100 DEPENDING ON WS-COUNT` -> edge
|
||||
|
||||
**Research insights (OCCURS):**
|
||||
- IBM allows `OCCURS 0 TO n DEPENDING ON` (zero minimum) and `OCCURS UNBOUNDED DEPENDING ON` (V6.4)
|
||||
- Subscripted controlling fields: `DEPENDING ON WS-COUNT(WS-IDX)` — strip subscripts before storing
|
||||
- **Pre-existing gap**: Multi-line data item clauses without continuation indicator are NOT captured. `05 WS-TABLE\n OCCURS 100\n DEPENDING ON WS-COUNT.` — the current RE_DATA_ITEM only gets the first line, `rest` is empty. Fixing properly requires a data item accumulator (like SELECT). **Defer full fix to Phase 3; implement same-line capture now.**
|
||||
- KEY IS fields: `ASCENDING KEY IS WS-KEY-1 WS-KEY-2` — capture for SEARCH ALL resolution
|
||||
- INDEXED BY: `INDEXED BY IDX-1 IDX-2` — capture for SET/SEARCH context
|
||||
|
||||
#### 1.5 VALUE clause for standard data items
|
||||
- **File:** `cobol-preprocessor.ts` (parseDataItemClauses)
|
||||
- **What:** Extract VALUE using a pragmatic function that handles quoted strings, numerics, figurative constants, hex/national literals
|
||||
- **Interface:** Already exists as `values?: string[]` on data items (currently only populated for 88-level)
|
||||
- **Graph:** Stored in Property node description (no new edges)
|
||||
- **Tests:** `01 WS-STATUS PIC X VALUE 'A'` -> values: ['A']
|
||||
|
||||
**Research insights (VALUE forms):**
|
||||
- Hex literals: `VALUE X'F1F2F3F4'`, National: `VALUE N'text'`, DBCS: `VALUE G'text'`
|
||||
- Figurative constants: SPACES, ZEROS, ZEROES, LOW-VALUES, HIGH-VALUES, QUOTES, NULL, NULLS
|
||||
- ALL literal: `VALUE ALL '*'`
|
||||
- Numeric with sign/decimal: `VALUE -123.45`, `VALUE +1`
|
||||
- `VALUE IS` optional — both `VALUE 'A'` and `VALUE IS 'A'` valid
|
||||
- **Decimal vs period ambiguity**: `VALUE 100.` — is `.` decimal or terminator? `parseDataItemClauses` already strips trailing period, so this is handled
|
||||
- IBM V6.4: floating-point `VALUE 1.0E5` — extend numeric regex if needed
|
||||
- Implementation: use a pragmatic `extractValue(rest)` function, not a single complex regex
|
||||
|
||||
### Phase 2: EXEC DLI + DECLARATIVES (~90 LOC, ~4 new edge types)
|
||||
|
||||
IMS/DB support and error handling flows.
|
||||
|
||||
#### 2.1 EXEC DLI (IMS/DB) -> ACCESSES edges (Graph value: 8.2/10)
|
||||
- **File:** `cobol-preprocessor.ts` (processLogicalLine — add RE_EXEC_DLI_START check alongside SQL/CICS)
|
||||
- **What:** Accumulate EXEC DLI blocks like EXEC SQL. Parse DLI verbs (GU, GN, GNP, GHU, GHN, GHNP, ISRT, DLET, REPL, CHKP, SCHD, TERM). Extract segment name, PCB number, INTO/FROM areas, WHERE fields, PSB name.
|
||||
- **Interface:** Add `execDliBlocks: Array<{ line: number; verb: string; pcbNumber?: number; segmentName?: string; intoField?: string; fromField?: string; whereField?: string; psbName?: string }>` to CobolRegexResults
|
||||
- **Graph:** CodeElement node + ACCESSES edge to `<ims>:<segmentName>` Record node with reason `dli-{verb}`; ACCESSES edges to INTO/FROM data areas; PSB ACCESSES for SCHD
|
||||
- **Tests:** `EXEC DLI GU USING PCB(1) SEGMENT(CUSTOMER) INTO(WS-CUST) END-EXEC`
|
||||
|
||||
**Research insights (dual IMS interface):**
|
||||
- **EXEC DLI**: Embedded command interface for CICS-DL/I programs only
|
||||
- **CBLTDLI CALL**: Batch interface via `CALL 'CBLTDLI' USING function-code PCB io-area SSA1..SSA15`
|
||||
- CBLTDLI is already captured as a CALL to 'CBLTDLI' — enrich with USING parameter semantics later
|
||||
- Multiple SEGMENT levels in hierarchical retrieval — use `matchAll` on segment regex
|
||||
- DLI verbs: GU (most common), GN, GNP, GHU, GHN, GHNP, ISRT, REPL, DLET, CHKP, SCHD, TERM, ROLL, ROLB
|
||||
- **Edge case**: DLET/REPL have no SEGMENT clause (operate on current position)
|
||||
- **Recommended order**: Implement AFTER DECLARATIVES and SET (lower risk, higher frequency)
|
||||
|
||||
#### 2.2 DECLARATIVES / USE AFTER STANDARD EXCEPTION (Graph value: 9.0/10)
|
||||
- **File:** `cobol-preprocessor.ts` (processLogicalLine — detect DECLARATIVES keyword, track USE AFTER blocks)
|
||||
- **What:** When `DECLARATIVES.` is encountered, switch to declaratives mode. Extract USE statements binding sections to files/modes.
|
||||
- **Interface:** Add `declaratives: Array<{ sectionName: string; useType: 'error' | 'debug' | 'label' | 'reporting'; target: string; line: number }>` to CobolRegexResults
|
||||
- **Graph:** ACCESSES edge from declarative Namespace to file Record with reason `cobol-declarative-error-handler`
|
||||
- **Tests:** Unit test with DECLARATIVES section, integration test for error flow
|
||||
|
||||
**Research insights (DECLARATIVES syntax):**
|
||||
- `USE AFTER STANDARD {EXCEPTION|ERROR} ON {file-name|INPUT|OUTPUT|I-O|EXTEND}`
|
||||
- EXCEPTION and ERROR are synonymous; STANDARD is optional in IBM dialects
|
||||
- Multiple USE sections allowed (one per file + catch-all for I/O modes)
|
||||
- `END DECLARATIVES.` must NOT reset PROCEDURE DIVISION state
|
||||
- `DECLARATIVES` is already in EXCLUDED_PARA_NAMES — no false paragraph risk
|
||||
- Existing section/paragraph detection already captures structural elements — just need USE binding
|
||||
- **Lowest risk Phase 2 item** — implement first
|
||||
|
||||
#### 2.3 SET statement -> ACCESSES edges
|
||||
- **File:** `cobol-preprocessor.ts` (extractProcedure — new RE_SET regex)
|
||||
- **Interface:** Add `sets: Array<{ targets: string[]; form: 'to-true'|'to-value'|'up-by'|'down-by'|'address-of'|'to-null'|'to-entry'; value?: string; entryTarget?: string; entryIsLiteral?: boolean; line: number; caller: string | null }>` to CobolRegexResults
|
||||
- **Graph:** ACCESSES write edge with reason `cobol-set-condition` (TO TRUE), `cobol-set-index` (TO/UP/DOWN), `cobol-set-address` (ADDRESS OF). SET ENTRY with literal -> CALLS edge.
|
||||
- **Tests:** `SET WS-EOF TO TRUE`, `SET IDX-1 TO 5`, `SET IDX-1 UP BY 1`
|
||||
|
||||
**Research insights (SET forms by frequency):**
|
||||
- `SET condition TO TRUE` — 80-90% of all SET usage. Multiple targets: `SET COND-A COND-B TO TRUE`
|
||||
- `SET index TO/UP BY/DOWN BY` — ~8%. Multiple indices: `SET IDX-1 IDX-2 UP BY 1`
|
||||
- `SET pointer TO ADDRESS OF data-item` / `SET ADDRESS OF data-item TO pointer` — ~2%
|
||||
- `SET proc-ptr TO ENTRY "PROGNAME"` — rare but creates CALLS edge (like dynamic CALL)
|
||||
- Filter OF/IN qualifiers: `SET COND-A OF WS-RECORD TO TRUE` (strip OF WS-RECORD)
|
||||
- **Prioritize**: SET TO TRUE alone covers 80-90% — implement this form first
|
||||
|
||||
#### 2.4 INSPECT -> ACCESSES edges
|
||||
- **File:** `cobol-preprocessor.ts` (extractProcedure — new `inspectAccum` accumulator like SORT)
|
||||
- **What:** Accumulate multi-line INSPECT until period. Extract inspected field + tally counters.
|
||||
- **Interface:** Add `inspects: Array<{ inspectedField: string; counters: string[]; form: 'tallying'|'replacing'|'converting'|'tallying-replacing'; line: number; caller: string | null }>` to CobolRegexResults
|
||||
- **Graph:** ACCESSES read on inspected field always; write if REPLACING/CONVERTING. Write edges for tally counters. Reason: `cobol-inspect-read`/`cobol-inspect-write`/`cobol-inspect-tally`
|
||||
- **Tests:** `INSPECT WS-FIELD TALLYING WS-COUNT FOR ALL 'A'` -> read on WS-FIELD, write on WS-COUNT
|
||||
|
||||
**Research insights (INSPECT forms by frequency):**
|
||||
- REPLACING (~60%): `INSPECT WS-STR REPLACING ALL 'A' BY 'B'`
|
||||
- TALLYING (~25%): `INSPECT WS-STR TALLYING WS-CNT FOR ALL 'A'` — multiple counters possible
|
||||
- CONVERTING (~10%): `INSPECT WS-STR CONVERTING 'abc' TO 'ABC'`
|
||||
- Combined (~5%): TALLYING + REPLACING in single statement
|
||||
- **Needs multi-line accumulator** — INSPECT frequently spans 3-5 lines in production
|
||||
- Extract tally counters with `([A-Z][A-Z0-9-]+)\s+FOR\b` matchAll pattern
|
||||
- Filter figurative constants (SPACES, ZEROS) using existing MOVE_SKIP set
|
||||
|
||||
### Phase 3: Completeness Fixes (~60 LOC)
|
||||
|
||||
Fix the 10 partial features and small gaps.
|
||||
|
||||
#### 3.1 CALL ... RETURNING extraction
|
||||
- Extend RE_CALL processing to capture RETURNING target after the USING clause
|
||||
- Store as `calls[].returning?: string`
|
||||
- Graph: ACCESSES write edge with reason `cobol-call-returning`
|
||||
|
||||
#### 3.2 SELECT OPTIONAL flag preservation
|
||||
- Store `isOptional: boolean` in FileDeclaration interface
|
||||
- Include in Record node description
|
||||
|
||||
#### 3.3 ALTERNATE RECORD KEY extraction
|
||||
- Add regex in parseSelectStatement: `/\bALTERNATE\s+RECORD\s+KEY\s+(?:IS\s+)?([A-Z][A-Z0-9-]+)/i`
|
||||
- Store as `alternateKeys?: string[]`
|
||||
|
||||
#### 3.4 COMMON attribute on nested programs
|
||||
- Extend RE_PROGRAM_ID: `/\bPROGRAM-ID\.\s*([A-Z][A-Z0-9-]+)(?:\s+IS\s+COMMON)?/i`
|
||||
- Store `isCommon: boolean` on Module node
|
||||
- Affects cross-program CALL resolution scope
|
||||
|
||||
#### 3.5 IS EXTERNAL / IS GLOBAL as first-class properties
|
||||
- Change from usage string hack to proper boolean fields on data items
|
||||
- Add `isExternal?: boolean`, `isGlobal?: boolean` to data item interface
|
||||
|
||||
#### 3.6 AUTHOR / DATE-WRITTEN mapped to Module node
|
||||
- Already extracted as programMetadata — map to Module node properties
|
||||
- `graph.addNode({ ..., properties: { ..., author, dateWritten } })`
|
||||
|
||||
#### 3.7 REPLACE statement
|
||||
- Track REPLACE / REPLACE OFF state in preprocessor
|
||||
- Apply text substitutions during preprocessing (before regex extraction)
|
||||
- Complex: requires careful scoping rules
|
||||
|
||||
### Phase 4: Niche Features (~30 LOC)
|
||||
|
||||
Low-priority but nice for completeness.
|
||||
|
||||
#### 4.1 INITIALIZE statement -> write ACCESSES
|
||||
- `/\bINITIALIZE\s+([A-Z][A-Z0-9-]+)/i`
|
||||
- ACCESSES write edge with reason `cobol-initialize`
|
||||
|
||||
#### 4.2 Remaining IDENTIFICATION DIVISION paragraphs
|
||||
- DATE-COMPILED, INSTALLATION, SECURITY, REMARKS
|
||||
- Map to Module node description properties
|
||||
|
||||
#### 4.3 EXEC SQL INCLUDE -> IMPORTS edge (expansion)
|
||||
- For EXEC SQL INCLUDE inside EXEC blocks that reference copybooks containing SQL
|
||||
- Create IMPORTS edge similar to COPY
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
### Functional Requirements
|
||||
|
||||
- [ ] Phase 1: All 5 features implemented with unit + integration tests
|
||||
- [ ] Phase 2: All 4 features implemented with unit + integration tests
|
||||
- [ ] Phase 3: All 7 partial features fixed
|
||||
- [ ] Phase 4: At least 2 of 3 niche features implemented
|
||||
- [ ] All existing 145 tests continue to pass
|
||||
- [ ] TypeScript compiles cleanly
|
||||
|
||||
### Non-Functional Requirements
|
||||
|
||||
- [ ] No performance regression: CardDemo benchmark stays under 8s
|
||||
- [ ] No file exceeds 1500 LOC (preprocessor currently 1326)
|
||||
- [ ] ACAS benchmark shows increased node/edge counts (more data extracted)
|
||||
- [ ] CardDemo benchmark shows increased edge counts (CALL USING, STRING, etc.)
|
||||
|
||||
### Quality Gates
|
||||
|
||||
- [ ] Each phase has its own commit
|
||||
- [ ] Integration test assertions updated with exact counts per phase
|
||||
- [ ] Benchmark run after each phase to track graph growth
|
||||
|
||||
## Dependencies & Risks
|
||||
|
||||
### Dependencies
|
||||
- None. All changes are additive to existing COBOL processor code.
|
||||
- No LanguageProvider changes needed.
|
||||
- No graph schema changes needed (all new constructs map to existing node labels + edge types).
|
||||
|
||||
### Risks
|
||||
- **preprocessor.ts size**: Currently 1326 LOC. Phase 1+2 adds ~200 LOC -> 1526 LOC. May need to extract helpers into a separate `cobol-data-flow.ts` module if it exceeds 1500.
|
||||
- **REPLACE statement** (Phase 3.7) is the most complex feature — requires tracking text substitution state across logical lines. Consider deferring to a separate PR if it takes >100 LOC.
|
||||
- **EXEC DLI** (Phase 2.1) is only testable against IMS codebases. Need fixture data or synthetic test cases.
|
||||
|
||||
## Graph Value Ranking by MCP Tool Impact
|
||||
|
||||
Research agent analyzed all 5 MCP tools (query, context, impact, detect_changes, rename) against planned edge types:
|
||||
|
||||
| Edge Type | QUERY | CONTEXT | IMPACT | DETECT | RENAME | **Overall** |
|
||||
|-----------|-------|---------|--------|--------|--------|-------------|
|
||||
| `cobol-call-using` | 4/5 | 5/5 | 5/5 | 4/5 | 4/5 | **9.2/10** |
|
||||
| `cobol-error-handler` | 5/5 | 4/5 | 5/5 | 5/5 | 2/5 | **9.0/10** |
|
||||
| `dli-*` (IMS verbs) | 4/5 | 4/5 | 5/5 | 4/5 | 2/5 | **8.2/10** |
|
||||
| `cobol-string-*` | 4/5 | 3/5 | 3/5 | 3/5 | 2/5 | **6.2/10** |
|
||||
|
||||
**Key finding**: `cobol-call-using` alone would fix ~40% of missing caller references in COBOL graphs.
|
||||
|
||||
## Future Considerations
|
||||
|
||||
This plan provides the graph data foundation for a future `modernize` MCP command (out of scope) that would:
|
||||
- Use CALL USING edges to map data contracts between programs
|
||||
- Use STRING/UNSTRING edges to identify data transformation logic
|
||||
- Use EXEC SQL/DLI edges to map database access patterns
|
||||
- Use DECLARATIVES to understand error handling architecture
|
||||
- Use the complete knowledge graph to generate migration plans
|
||||
|
||||
**MCP tool enhancements needed** (after this plan ships):
|
||||
- Add `cobol-call-using`, `cobol-error-handler`, `dli-*` to IMPACT tool's default `relationTypes` for COBOL repos
|
||||
- Add confidence floors for new edge types in `IMPACT_RELATION_CONFIDENCE`
|
||||
- Register new edge types in `VALID_RELATION_TYPES` set (`local-backend.ts:52`)
|
||||
|
||||
## Sources & References
|
||||
|
||||
### Internal References
|
||||
- Feature audit: session 8642401e (COBOL expert agent, 123 features audited)
|
||||
- Prior plans: `docs/plans/2026-03-25-feat-cobol-100-percent-feature-coverage-plan.md`
|
||||
- Architecture: `docs/code-indexing/cobol/` (7 documentation files)
|
||||
|
||||
### External References
|
||||
- COBOL features reference: mainframestechhelp.com/tutorials/cobol/features.htm
|
||||
- COBOL-85 standard: ISO/IEC 1989:1985
|
||||
- IBM Enterprise COBOL reference
|
||||
|
|
@ -0,0 +1,55 @@
|
|||
---
|
||||
title: "Field Extractors for All Supported Languages"
|
||||
type: feat
|
||||
status: active
|
||||
date: 2026-03-26
|
||||
---
|
||||
|
||||
# Field Extractors for All Supported Languages
|
||||
|
||||
## Overview
|
||||
|
||||
PR #494 adds a `FieldExtractor` infrastructure with only a TypeScript implementation. This plan fills the registry for all 14 supported languages using a table-driven generic extractor, plus unit tests.
|
||||
|
||||
## Approach: Generic Table-Driven Extractor
|
||||
|
||||
Instead of 14 separate 300+ line files, create a `GenericFieldExtractor` configured via a per-language `FieldExtractionConfig`. Each config specifies:
|
||||
- AST node types for type declarations (class_declaration, struct_item, etc.)
|
||||
- AST node types for field declarations within bodies
|
||||
- How to extract field name, type, visibility, static, readonly from the AST
|
||||
- Default visibility and body node type
|
||||
|
||||
## Implementation
|
||||
|
||||
### Phase 1: Generic Field Extractor
|
||||
|
||||
**Create:** `field-extractors/generic.ts`
|
||||
|
||||
A single `createFieldExtractor(config)` factory that returns a `FieldExtractor` for any language.
|
||||
|
||||
### Phase 2: Language Configs
|
||||
|
||||
**Create:** `field-extractors/configs.ts`
|
||||
|
||||
Export configs for all 14 languages. Each config is ~20-40 lines of node type mappings.
|
||||
|
||||
### Phase 3: Register All in Index
|
||||
|
||||
**Modify:** `field-extractors/index.ts`
|
||||
|
||||
Register all 14 extractors. Keep the TypeScript-specific class for backwards compatibility.
|
||||
|
||||
### Phase 4: Unit Tests
|
||||
|
||||
**Create:** `test/unit/field-extraction-all-languages.test.ts`
|
||||
|
||||
For each language, parse a small code snippet and verify the extractor produces correct FieldInfo[].
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
- [x] GenericFieldExtractor created with table-driven config
|
||||
- [ ] Configs for: TS, JS, Python, Java, Kotlin, Go, Rust, C#, C++, C, PHP, Ruby, Swift, Dart
|
||||
- [ ] All extractors registered in index.ts
|
||||
- [ ] Unit tests for all languages
|
||||
- [ ] `npx tsc --noEmit` passes
|
||||
- [ ] All tests pass
|
||||
|
|
@ -42,4 +42,6 @@ export enum SupportedLanguages {
|
|||
Kotlin = 'kotlin',
|
||||
Swift = 'swift',
|
||||
Dart = 'dart',
|
||||
/** Standalone regex processor — no tree-sitter, no LanguageProvider. */
|
||||
Cobol = 'cobol',
|
||||
}
|
||||
|
|
|
|||
|
|
@ -34,7 +34,15 @@ export const createKnowledgeGraph = (): KnowledgeGraph => {
|
|||
};
|
||||
|
||||
/**
|
||||
* Remove all nodes (and their relationships) belonging to a file
|
||||
* Remove a single relationship by id.
|
||||
* Returns true if the relationship existed and was removed, false otherwise.
|
||||
*/
|
||||
const removeRelationship = (relationshipId: string): boolean => {
|
||||
return relationshipMap.delete(relationshipId);
|
||||
};
|
||||
|
||||
/**
|
||||
* Remove all nodes (and their relationships) belonging to a file.
|
||||
*/
|
||||
const removeNodesByFile = (filePath: string): number => {
|
||||
let removed = 0;
|
||||
|
|
@ -75,6 +83,7 @@ export const createKnowledgeGraph = (): KnowledgeGraph => {
|
|||
addRelationship,
|
||||
removeNode,
|
||||
removeNodesByFile,
|
||||
removeRelationship,
|
||||
|
||||
};
|
||||
};
|
||||
|
|
|
|||
|
|
@ -146,4 +146,5 @@ export interface KnowledgeGraph {
|
|||
addRelationship: (relationship: GraphRelationship) => void,
|
||||
removeNode: (nodeId: string) => boolean,
|
||||
removeNodesByFile: (filePath: string) => number,
|
||||
removeRelationship: (relationshipId: string) => boolean,
|
||||
}
|
||||
|
|
|
|||
|
|
@ -11,7 +11,6 @@ import { getLanguageFromFilename } from './utils/language-detection.js';
|
|||
import { isVerboseIngestionEnabled } from './utils/verbose.js';
|
||||
import { yieldToEventLoop } from './utils/event-loop.js';
|
||||
import { FUNCTION_NODE_TYPES, extractFunctionName, findEnclosingClassId } from './utils/ast-helpers.js';
|
||||
import { isBuiltInOrNoise } from './utils/noise-filter.js';
|
||||
import {
|
||||
countCallArguments,
|
||||
inferCallForm,
|
||||
|
|
@ -210,6 +209,26 @@ const findEnclosingFunction = (
|
|||
return generateId(finalLabel, `${filePath}:${funcName}`);
|
||||
}
|
||||
}
|
||||
|
||||
// Language-specific enclosing function resolution (e.g., Dart where
|
||||
// function_body is a sibling of function_signature, not a child).
|
||||
if (provider.enclosingFunctionFinder) {
|
||||
const customResult = provider.enclosingFunctionFinder(current);
|
||||
if (customResult) {
|
||||
// Try SymbolTable first (same pattern as the FUNCTION_NODE_TYPES branch above).
|
||||
const resolved = ctx.resolve(customResult.funcName, filePath);
|
||||
if (resolved?.tier === 'same-file' && resolved.candidates.length > 0) {
|
||||
return resolved.candidates[0].nodeId;
|
||||
}
|
||||
let finalLabel = customResult.label;
|
||||
if (provider.labelOverride) {
|
||||
const override = provider.labelOverride(current.previousSibling!, finalLabel);
|
||||
if (override !== null) finalLabel = override;
|
||||
}
|
||||
return generateId(finalLabel, `${filePath}:${customResult.funcName}`);
|
||||
}
|
||||
}
|
||||
|
||||
current = current.parent;
|
||||
}
|
||||
|
||||
|
|
@ -497,7 +516,7 @@ export const processCalls = async (
|
|||
}
|
||||
}
|
||||
|
||||
if (isBuiltInOrNoise(calledName)) return;
|
||||
if (provider.isBuiltInName(calledName)) return;
|
||||
|
||||
const callNode = captureMap['call'];
|
||||
const callForm = inferCallForm(callNode, nameNode);
|
||||
|
|
|
|||
1308
gitnexus/src/core/ingestion/cobol-processor.ts
Normal file
1308
gitnexus/src/core/ingestion/cobol-processor.ts
Normal file
File diff suppressed because it is too large
Load diff
501
gitnexus/src/core/ingestion/cobol/cobol-copy-expander.ts
Normal file
501
gitnexus/src/core/ingestion/cobol/cobol-copy-expander.ts
Normal file
|
|
@ -0,0 +1,501 @@
|
|||
/**
|
||||
* COBOL COPY statement expansion engine.
|
||||
*
|
||||
* Expands COPY statements by inlining copybook content, applying REPLACING
|
||||
* transformations (LEADING, TRAILING, EXACT), and handling nested copies
|
||||
* with cycle detection.
|
||||
*
|
||||
* This is a preprocessing step that runs BEFORE extractCobolSymbolsWithRegex.
|
||||
* The caller should run preprocessCobolSource first to clean patch markers.
|
||||
*
|
||||
* Supported syntax:
|
||||
* COPY CPSESP.
|
||||
* COPY "WORKGRID.CPY".
|
||||
* COPY CPSESP REPLACING LEADING "ESP-" BY "LK-ESP-"
|
||||
* LEADING "KPSESPL" BY "LK-KPSESPL".
|
||||
* COPY ANAZI REPLACING "ANAZI-KEY" BY "LK-KEY".
|
||||
*/
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Public interfaces
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
export interface CopyReplacing {
|
||||
type: 'LEADING' | 'TRAILING' | 'EXACT';
|
||||
from: string;
|
||||
to: string;
|
||||
isPseudotext?: boolean;
|
||||
}
|
||||
|
||||
export interface CopyResolution {
|
||||
copyTarget: string;
|
||||
resolvedPath: string | null;
|
||||
line: number;
|
||||
replacing: CopyReplacing[];
|
||||
library?: string;
|
||||
}
|
||||
|
||||
export interface CopyExpansionResult {
|
||||
expandedContent: string;
|
||||
copyResolutions: CopyResolution[];
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Constants
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
export const DEFAULT_MAX_DEPTH = 10;
|
||||
|
||||
/** COBOL identifier pattern: starts with letter, contains letters, digits, hyphens. */
|
||||
const RE_COBOL_IDENTIFIER = /\b([A-Z][A-Z0-9-]*)\b/gi;
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Private helpers
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/**
|
||||
* Strip inline comments (Italian-style `|` comments).
|
||||
* Only strips if `|` appears in the code area (col 7+).
|
||||
*/
|
||||
function stripInlineComment(line: string): string {
|
||||
let inQuote: string | null = null;
|
||||
for (let i = 0; i < line.length; i++) {
|
||||
const ch = line[i];
|
||||
if (inQuote) {
|
||||
if (ch === inQuote) inQuote = null;
|
||||
} else if (ch === '"' || ch === "'") {
|
||||
inQuote = ch;
|
||||
} else if (ch === '|') {
|
||||
return line.substring(0, i);
|
||||
}
|
||||
}
|
||||
return line;
|
||||
}
|
||||
|
||||
/**
|
||||
* Check if a line is a COBOL comment (indicator in col 7 is `*` or `/`).
|
||||
*/
|
||||
function isCommentLine(line: string): boolean {
|
||||
return line.length >= 7 && (line[6] === '*' || line[6] === '/');
|
||||
}
|
||||
|
||||
/**
|
||||
* Check if a line is a continuation line (indicator in col 7 is `-`).
|
||||
*/
|
||||
function isContinuationLine(line: string): boolean {
|
||||
return line.length >= 7 && line[6] === '-';
|
||||
}
|
||||
|
||||
/**
|
||||
* Merge continuation lines into their predecessors.
|
||||
* Returns an array of logical lines with their original starting line numbers.
|
||||
*/
|
||||
function mergeLogicalLines(
|
||||
rawLines: string[],
|
||||
): Array<{ text: string; lineNum: number }> {
|
||||
const logical: Array<{ text: string; lineNum: number }> = [];
|
||||
|
||||
for (let i = 0; i < rawLines.length; i++) {
|
||||
const raw = rawLines[i];
|
||||
|
||||
// Skip comment lines
|
||||
if (isCommentLine(raw)) {
|
||||
logical.push({ text: '', lineNum: i + 1 });
|
||||
continue;
|
||||
}
|
||||
|
||||
// Continuation: merge into previous logical line
|
||||
if (isContinuationLine(raw)) {
|
||||
if (logical.length > 0) {
|
||||
const prev = logical[logical.length - 1];
|
||||
const continuation = raw.length > 7 ? raw.substring(7).trimStart() : '';
|
||||
prev.text += continuation;
|
||||
}
|
||||
// Push empty placeholder to preserve line count
|
||||
logical.push({ text: '', lineNum: i + 1 });
|
||||
continue;
|
||||
}
|
||||
|
||||
// Normal line: strip inline comments
|
||||
const cleaned = stripInlineComment(raw);
|
||||
logical.push({ text: cleaned, lineNum: i + 1 });
|
||||
}
|
||||
|
||||
return logical;
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// COPY statement parsing
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
interface ParsedCopyStatement {
|
||||
startLine: number;
|
||||
endLine: number;
|
||||
target: string;
|
||||
replacing: CopyReplacing[];
|
||||
library?: string;
|
||||
}
|
||||
|
||||
/**
|
||||
* Parse REPLACING clause text into structured replacements.
|
||||
*
|
||||
* Input examples:
|
||||
* LEADING "ESP-" BY "LK-ESP-" LEADING "KPSESPL" BY "LK-KPSESPL"
|
||||
* "ANAZI-KEY" BY "LK-KEY"
|
||||
* TRAILING "-IN" BY "-OUT"
|
||||
* ==CUST-== BY ==WS-CUST-==
|
||||
* ==OLD-TEXT== BY ====
|
||||
*/
|
||||
export function parseReplacingClause(text: string): CopyReplacing[] {
|
||||
const replacings: CopyReplacing[] = [];
|
||||
if (!text || text.trim().length === 0) return replacings;
|
||||
|
||||
// Tokenize: ==pseudotext==, "quoted strings", or bare words.
|
||||
// Pseudotext can contain spaces and single = chars but not ==.
|
||||
interface TokenInfo { value: string; isPseudotext: boolean; }
|
||||
const tokens: TokenInfo[] = [];
|
||||
const tokenRe = /==((?:[^=]|=[^=])*)==|"([^"]*)"|(\S+)/g;
|
||||
let tm: RegExpExecArray | null;
|
||||
while ((tm = tokenRe.exec(text)) !== null) {
|
||||
if (tm[1] !== undefined) {
|
||||
// Pseudotext: trim leading/trailing whitespace
|
||||
tokens.push({ value: tm[1].trim(), isPseudotext: true });
|
||||
} else if (tm[2] !== undefined) {
|
||||
tokens.push({ value: tm[2], isPseudotext: false });
|
||||
} else {
|
||||
tokens.push({ value: tm[3], isPseudotext: false });
|
||||
}
|
||||
}
|
||||
|
||||
// Parse token stream: [LEADING|TRAILING]? <from> BY <to>
|
||||
let i = 0;
|
||||
while (i < tokens.length) {
|
||||
let type: CopyReplacing['type'] = 'EXACT';
|
||||
|
||||
// Check for type modifier (only on non-pseudotext tokens)
|
||||
if (!tokens[i].isPseudotext) {
|
||||
const upper = tokens[i].value.toUpperCase();
|
||||
if (upper === 'LEADING') {
|
||||
type = 'LEADING';
|
||||
i++;
|
||||
} else if (upper === 'TRAILING') {
|
||||
type = 'TRAILING';
|
||||
i++;
|
||||
}
|
||||
}
|
||||
|
||||
if (i >= tokens.length) break;
|
||||
const fromToken = tokens[i];
|
||||
i++;
|
||||
|
||||
// Pseudotext always forces EXACT type
|
||||
if (fromToken.isPseudotext) type = 'EXACT';
|
||||
|
||||
// Expect BY keyword
|
||||
if (i >= tokens.length) break;
|
||||
if (tokens[i].value.toUpperCase() !== 'BY') {
|
||||
// Malformed — skip this token and try to resync
|
||||
continue;
|
||||
}
|
||||
i++; // skip BY
|
||||
|
||||
if (i >= tokens.length) break;
|
||||
const toToken = tokens[i];
|
||||
i++;
|
||||
|
||||
replacings.push({ type, from: fromToken.value, to: toToken.value, isPseudotext: fromToken.isPseudotext || undefined });
|
||||
}
|
||||
|
||||
return replacings;
|
||||
}
|
||||
|
||||
/**
|
||||
* Scan logical lines for COPY statements.
|
||||
* COPY statements can span multiple lines and terminate with a period.
|
||||
*/
|
||||
function parseCopyStatements(
|
||||
logicalLines: Array<{ text: string; lineNum: number }>,
|
||||
): ParsedCopyStatement[] {
|
||||
const results: ParsedCopyStatement[] = [];
|
||||
|
||||
let accumulator: string | null = null;
|
||||
let startLine = 0;
|
||||
let endLine = 0;
|
||||
|
||||
for (let i = 0; i < logicalLines.length; i++) {
|
||||
const { text, lineNum } = logicalLines[i];
|
||||
if (text.length === 0) continue;
|
||||
|
||||
// Check for COPY keyword start (not inside a string context)
|
||||
const copyStart = text.match(/\bCOPY\b/i);
|
||||
|
||||
if (accumulator === null) {
|
||||
if (!copyStart) continue;
|
||||
|
||||
// Start accumulating from the COPY keyword onwards
|
||||
const copyIdx = copyStart.index!;
|
||||
accumulator = text.substring(copyIdx);
|
||||
startLine = lineNum;
|
||||
endLine = lineNum;
|
||||
} else {
|
||||
// Continue accumulating
|
||||
accumulator += ' ' + text.trim();
|
||||
endLine = lineNum;
|
||||
}
|
||||
|
||||
// Check if statement terminates (period at end of accumulated text)
|
||||
if (accumulator !== null && /\.\s*$/.test(accumulator)) {
|
||||
const parsed = parseSingleCopyStatement(accumulator, startLine, endLine);
|
||||
if (parsed) {
|
||||
results.push(parsed);
|
||||
}
|
||||
accumulator = null;
|
||||
}
|
||||
}
|
||||
|
||||
// If there's an unterminated COPY (missing period), try to parse what we have
|
||||
if (accumulator !== null) {
|
||||
const parsed = parseSingleCopyStatement(accumulator, startLine, endLine);
|
||||
if (parsed) {
|
||||
results.push(parsed);
|
||||
}
|
||||
}
|
||||
|
||||
return results;
|
||||
}
|
||||
|
||||
/**
|
||||
* Parse a single complete COPY statement string.
|
||||
*
|
||||
* Formats:
|
||||
* COPY target.
|
||||
* COPY "target".
|
||||
* COPY target REPLACING ... .
|
||||
*/
|
||||
function parseSingleCopyStatement(
|
||||
stmt: string,
|
||||
startLine: number,
|
||||
endLine: number,
|
||||
): ParsedCopyStatement | null {
|
||||
// Strip terminating period
|
||||
const text = stmt.replace(/\.\s*$/, '').trim();
|
||||
|
||||
// Extract target: COPY <target> or COPY "<target>" or COPY '<target>'
|
||||
// Optionally followed by IN/OF <library-name> (COBOL-85 standard: IN and OF are synonyms)
|
||||
const targetMatch = text.match(
|
||||
/^COPY\s+(?:"([^"]+)"|'([^']+)'|([A-Z][A-Z0-9-]*))(?:\s+(?:IN|OF)\s+([A-Z][A-Z0-9-]*))?/i,
|
||||
);
|
||||
if (!targetMatch) return null;
|
||||
|
||||
const target = targetMatch[1] ?? targetMatch[2] ?? targetMatch[3];
|
||||
const library = targetMatch[4] || undefined;
|
||||
|
||||
// Extract REPLACING clause if present
|
||||
let replacing: CopyReplacing[] = [];
|
||||
const replacingIdx = text.search(/\bREPLACING\b/i);
|
||||
if (replacingIdx >= 0) {
|
||||
const replacingText = text.substring(replacingIdx + 'REPLACING'.length);
|
||||
replacing = parseReplacingClause(replacingText);
|
||||
}
|
||||
|
||||
return { startLine, endLine, target, replacing, library };
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// REPLACING application
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/**
|
||||
* Apply REPLACING transformations to copybook content.
|
||||
*
|
||||
* LEADING: replace prefix in COBOL identifiers.
|
||||
* TRAILING: replace suffix in COBOL identifiers.
|
||||
* EXACT: replace exact token matches.
|
||||
*/
|
||||
function applyReplacing(content: string, replacings: CopyReplacing[]): string {
|
||||
if (replacings.length === 0) return content;
|
||||
|
||||
// First pass: handle EXACT replacements that contain spaces or non-identifier
|
||||
// characters (pseudotext). These cannot be handled by identifier-level matching.
|
||||
let result = content;
|
||||
for (const r of replacings) {
|
||||
if (r.type === 'EXACT' && (r.isPseudotext || r.from.includes(' ') || !/^[A-Z][A-Z0-9-]*$/i.test(r.from))) {
|
||||
const escaped = r.from.replace(/[.*+?^${}()|[\]\\]/g, '\\$&');
|
||||
const re = new RegExp(escaped, 'gi');
|
||||
result = result.replace(re, r.to);
|
||||
}
|
||||
}
|
||||
|
||||
// Second pass: identifier-level replacements (LEADING, TRAILING, single-word EXACT)
|
||||
const identifierReplacings = replacings.filter(
|
||||
r => !(r.type === 'EXACT' && (r.isPseudotext || r.from.includes(' ') || !/^[A-Z][A-Z0-9-]*$/i.test(r.from))),
|
||||
);
|
||||
if (identifierReplacings.length === 0) return result;
|
||||
|
||||
return result.replace(RE_COBOL_IDENTIFIER, (match) => {
|
||||
for (const r of identifierReplacings) {
|
||||
const upper = match.toUpperCase();
|
||||
const from = r.from.toUpperCase();
|
||||
const to = r.to.toUpperCase();
|
||||
switch (r.type) {
|
||||
case 'LEADING':
|
||||
if (upper.startsWith(from)) {
|
||||
return to + match.substring(from.length);
|
||||
}
|
||||
break;
|
||||
case 'TRAILING':
|
||||
if (upper.endsWith(from)) {
|
||||
return match.substring(0, match.length - from.length) + to;
|
||||
}
|
||||
break;
|
||||
case 'EXACT':
|
||||
if (upper === from) {
|
||||
return to;
|
||||
}
|
||||
break;
|
||||
}
|
||||
}
|
||||
return match;
|
||||
});
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Main expansion engine
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/**
|
||||
* Expand COBOL COPY statements by inlining copybook content.
|
||||
*
|
||||
* @param content - Source COBOL content (after preprocessCobolSource)
|
||||
* @param filePath - Path of the source file (for diagnostics)
|
||||
* @param resolveFile - Maps a COPY target name to a filesystem path, or null if not found
|
||||
* @param readFile - Reads file content by path, or null if unreadable
|
||||
* @param maxDepth - Maximum nesting depth for recursive expansion (default: 10)
|
||||
* @returns Expanded content and resolution metadata
|
||||
*/
|
||||
export function expandCopies(
|
||||
content: string,
|
||||
filePath: string,
|
||||
resolveFile: (name: string) => string | null,
|
||||
readFile: (path: string) => string | null,
|
||||
maxDepth: number = DEFAULT_MAX_DEPTH,
|
||||
): CopyExpansionResult {
|
||||
const allResolutions: CopyResolution[] = [];
|
||||
const warnedCircular = new Set<string>();
|
||||
let totalExpansions = 0;
|
||||
const MAX_TOTAL_EXPANSIONS = 500;
|
||||
|
||||
const expanded = expandRecursive(content, filePath, 0, new Set<string>());
|
||||
|
||||
return {
|
||||
expandedContent: expanded,
|
||||
copyResolutions: allResolutions,
|
||||
};
|
||||
|
||||
/**
|
||||
* Recursively expand COPY statements in content.
|
||||
*
|
||||
* @param src - Source content to expand
|
||||
* @param srcPath - Path of the file being expanded (for cycle detection logging)
|
||||
* @param depth - Current recursion depth
|
||||
* @param visited - Set of already-visited copybook paths (cycle detection)
|
||||
*/
|
||||
function expandRecursive(
|
||||
src: string,
|
||||
srcPath: string,
|
||||
depth: number,
|
||||
visited: Set<string>,
|
||||
): string {
|
||||
const rawLines = src.split(/\r?\n/);
|
||||
const logicalLines = mergeLogicalLines(rawLines);
|
||||
const copyStatements = parseCopyStatements(logicalLines);
|
||||
|
||||
// No COPY statements — return as-is
|
||||
if (copyStatements.length === 0) return src;
|
||||
|
||||
// Process COPY statements in reverse order so line numbers stay valid
|
||||
// as we splice content
|
||||
const outputLines = [...rawLines];
|
||||
|
||||
for (let ci = copyStatements.length - 1; ci >= 0; ci--) {
|
||||
const cs = copyStatements[ci];
|
||||
|
||||
// Resolve the copybook path
|
||||
const resolvedPath = resolveFile(cs.target);
|
||||
|
||||
// Record resolution metadata
|
||||
allResolutions.push({
|
||||
copyTarget: cs.target,
|
||||
resolvedPath,
|
||||
line: cs.startLine,
|
||||
replacing: cs.replacing,
|
||||
library: cs.library,
|
||||
});
|
||||
|
||||
// Cannot resolve — keep original lines
|
||||
if (resolvedPath === null) {
|
||||
continue;
|
||||
}
|
||||
|
||||
// Cycle detection
|
||||
if (visited.has(resolvedPath)) {
|
||||
if (!warnedCircular.has(resolvedPath)) {
|
||||
warnedCircular.add(resolvedPath);
|
||||
console.warn(
|
||||
`[cobol-copy-expander] Circular COPY detected: ${cs.target} (${resolvedPath}) ` +
|
||||
`includes itself. Skipping expansion.`,
|
||||
);
|
||||
}
|
||||
continue;
|
||||
}
|
||||
|
||||
// Max depth exceeded — keep unexpanded
|
||||
if (depth >= maxDepth) {
|
||||
console.warn(
|
||||
`[cobol-copy-expander] Max expansion depth (${maxDepth}) reached for ` +
|
||||
`COPY ${cs.target} in ${srcPath}. Skipping expansion.`,
|
||||
);
|
||||
continue;
|
||||
}
|
||||
|
||||
// Guard against exponential breadth amplification (N copybooks each with N COPYs)
|
||||
if (++totalExpansions > MAX_TOTAL_EXPANSIONS) {
|
||||
if (!warnedCircular.has('__max_total__')) {
|
||||
warnedCircular.add('__max_total__');
|
||||
console.warn(
|
||||
`[cobol-copy-expander] Max total expansions (${MAX_TOTAL_EXPANSIONS}) reached ` +
|
||||
`in ${srcPath}. Skipping further expansions.`,
|
||||
);
|
||||
}
|
||||
continue;
|
||||
}
|
||||
|
||||
// Read the copybook content
|
||||
const copybookContent = readFile(resolvedPath);
|
||||
if (copybookContent === null) {
|
||||
continue;
|
||||
}
|
||||
|
||||
// Apply REPLACING transformations
|
||||
const replaced = applyReplacing(copybookContent, cs.replacing);
|
||||
|
||||
// Recurse into the copybook for nested COPYs
|
||||
const nestedVisited = new Set(visited);
|
||||
nestedVisited.add(resolvedPath);
|
||||
const expandedCopybook = expandRecursive(
|
||||
replaced,
|
||||
resolvedPath,
|
||||
depth + 1,
|
||||
nestedVisited,
|
||||
);
|
||||
|
||||
// Splice: replace the COPY statement lines with expanded content
|
||||
// startLine/endLine are 1-based; convert to 0-based array index
|
||||
const expansionLines = expandedCopybook.split('\n');
|
||||
const removeCount = cs.endLine - cs.startLine + 1;
|
||||
outputLines.splice(cs.startLine - 1, removeCount, ...expansionLines);
|
||||
}
|
||||
|
||||
return outputLines.join('\n');
|
||||
}
|
||||
}
|
||||
1771
gitnexus/src/core/ingestion/cobol/cobol-preprocessor.ts
Normal file
1771
gitnexus/src/core/ingestion/cobol/cobol-preprocessor.ts
Normal file
File diff suppressed because it is too large
Load diff
263
gitnexus/src/core/ingestion/cobol/jcl-parser.ts
Normal file
263
gitnexus/src/core/ingestion/cobol/jcl-parser.ts
Normal file
|
|
@ -0,0 +1,263 @@
|
|||
/**
|
||||
* JCL Parser — Regex single-pass extraction.
|
||||
*
|
||||
* Extracts JCL constructs from mainframe job streams:
|
||||
* - JOB statements (job name, CLASS, MSGCLASS)
|
||||
* - EXEC statements (step -> program or proc)
|
||||
* - DD statements (dataset references, DISP)
|
||||
* - PROC definitions (in-stream and catalogued)
|
||||
* - INCLUDE MEMBER= directives
|
||||
* - SET symbolic parameters
|
||||
* - IF/ELSE/ENDIF conditional execution
|
||||
* - JCLLIB ORDER= search paths
|
||||
*
|
||||
* Pattern follows cobol-preprocessor.ts — regex-only, no tree-sitter.
|
||||
*/
|
||||
|
||||
export interface JclParseResults {
|
||||
jobs: Array<{ name: string; line: number; class?: string; msgclass?: string }>;
|
||||
steps: Array<{ name: string; jobName: string; program?: string; proc?: string; line: number }>;
|
||||
ddStatements: Array<{ ddName: string; stepName: string; dataset?: string; disp?: string; line: number }>;
|
||||
procs: Array<{ name: string; line: number; isInStream: boolean }>;
|
||||
includes: Array<{ member: string; line: number }>;
|
||||
sets: Array<{ variable: string; value: string; line: number }>;
|
||||
jcllib: Array<{ order: string[]; line: number }>;
|
||||
conditionals: Array<{ type: 'IF' | 'ELSE' | 'ENDIF'; condition?: string; line: number }>;
|
||||
}
|
||||
|
||||
// ── JCL statement patterns ─────────────────────────────────────────────
|
||||
|
||||
// JCL continuation: line ends with a non-blank in col 72, next line starts with //
|
||||
// We handle continuations by joining lines before matching.
|
||||
|
||||
/** Match //jobname JOB ... */
|
||||
const JOB_RE = /^\/\/(\w{1,8})\s+JOB\s+(.*)/i;
|
||||
|
||||
/** Match //stepname EXEC PGM=program or //stepname EXEC procname */
|
||||
const EXEC_RE = /^\/\/(\w{1,8})\s+EXEC\s+(.*)/i;
|
||||
|
||||
/** Match //ddname DD ... */
|
||||
const DD_RE = /^\/\/(\w{1,8})\s+DD\s+(.*)/i;
|
||||
|
||||
/** Match // JCLLIB ORDER=(lib1,lib2,...) */
|
||||
const JCLLIB_RE = /^\/\/\s+JCLLIB\s+ORDER=\(([^)]+)\)/i;
|
||||
|
||||
/** Match // IF condition THEN */
|
||||
const IF_RE = /^\/\/\s+IF\s+(.+)\s+THEN/i;
|
||||
|
||||
/** Match // ELSE */
|
||||
const ELSE_RE = /^\/\/\s+ELSE\b/i;
|
||||
|
||||
/** Match // ENDIF */
|
||||
const ENDIF_RE = /^\/\/\s+ENDIF\b/i;
|
||||
|
||||
/** Match // INCLUDE MEMBER=name */
|
||||
const INCLUDE_RE = /^\/\/\s+INCLUDE\s+MEMBER=(\w+)/i;
|
||||
|
||||
/** Match // SET var=value */
|
||||
const SET_RE = /^\/\/\s+SET\s+(\w+)=(.+)/i;
|
||||
|
||||
/** Match // PROC or //name PROC */
|
||||
const PROC_RE = /^\/\/(\w*)\s+PROC\b/i;
|
||||
|
||||
/** Match // PEND */
|
||||
const PEND_RE = /^\/\/\s+PEND\b/i;
|
||||
|
||||
// ── Parameter extractors ───────────────────────────────────────────────
|
||||
|
||||
function extractParam(params: string, key: string): string | undefined {
|
||||
// Match KEY=VALUE or KEY='VALUE' in JCL parameter string
|
||||
const re = new RegExp(`${key}=(?:'([^']*)'|(\\S+?))(?:[,\\s]|$)`, 'i');
|
||||
const m = params.match(re);
|
||||
return m ? (m[1] ?? m[2]) : undefined;
|
||||
}
|
||||
|
||||
function extractPgm(params: string): string | undefined {
|
||||
return extractParam(params, 'PGM');
|
||||
}
|
||||
|
||||
function extractProc(params: string): string | undefined {
|
||||
// If no PGM= keyword, the first positional parameter is the proc name
|
||||
if (/PGM=/i.test(params)) return undefined;
|
||||
const cleaned = params.replace(/,.*/, '').trim();
|
||||
// Proc name is the first token (no = sign)
|
||||
if (cleaned && !cleaned.includes('=')) {
|
||||
return cleaned.replace(/[,\s].*/s, '').toUpperCase();
|
||||
}
|
||||
return undefined;
|
||||
}
|
||||
|
||||
function extractDsn(params: string): string | undefined {
|
||||
return extractParam(params, 'DSN') ?? extractParam(params, 'DSNAME');
|
||||
}
|
||||
|
||||
function extractDisp(params: string): string | undefined {
|
||||
const m = params.match(/DISP=\(?\s*([^),\s]+)/i);
|
||||
return m ? m[1] : undefined;
|
||||
}
|
||||
|
||||
/**
|
||||
* Parse a JCL file and extract all constructs.
|
||||
*
|
||||
* @param content - Raw JCL file content
|
||||
* @param filePath - Path for diagnostics (not used in extraction)
|
||||
* @returns Parsed JCL results
|
||||
*/
|
||||
export function parseJcl(content: string, filePath: string): JclParseResults {
|
||||
const results: JclParseResults = {
|
||||
jobs: [],
|
||||
steps: [],
|
||||
ddStatements: [],
|
||||
procs: [],
|
||||
includes: [],
|
||||
sets: [],
|
||||
jcllib: [],
|
||||
conditionals: [],
|
||||
};
|
||||
|
||||
const rawLines = content.split(/\r?\n/);
|
||||
// Join continuation lines: a line ending with non-blank in col 71 (0-indexed)
|
||||
// followed by a line starting with // is a continuation.
|
||||
const lines: Array<{ text: string; lineNum: number }> = [];
|
||||
let i = 0;
|
||||
while (i < rawLines.length) {
|
||||
let line = rawLines[i];
|
||||
const lineNum = i + 1;
|
||||
|
||||
// JCL continuation: if line is exactly 72+ chars and col 72 is non-blank
|
||||
// and the next line starts with //, join them.
|
||||
while (
|
||||
i + 1 < rawLines.length &&
|
||||
line.length >= 72 &&
|
||||
line[71] !== ' ' &&
|
||||
rawLines[i + 1].startsWith('//')
|
||||
) {
|
||||
i++;
|
||||
// Continuation text starts after // and leading spaces
|
||||
const contText = rawLines[i].substring(2).replace(/^\s+/, ' ');
|
||||
// Remove the continuation marker (col 72+) from current line
|
||||
line = line.substring(0, 71).trimEnd() + contText;
|
||||
}
|
||||
|
||||
lines.push({ text: line, lineNum });
|
||||
i++;
|
||||
}
|
||||
|
||||
let currentJobName = '';
|
||||
let currentStepName = '';
|
||||
let inStreamProcName = '';
|
||||
|
||||
for (const { text, lineNum } of lines) {
|
||||
// Skip JCL comments (starting with //* )
|
||||
if (text.startsWith('//*')) continue;
|
||||
// Skip non-JCL lines (don't start with //)
|
||||
if (!text.startsWith('//')) continue;
|
||||
|
||||
// PROC definition (in-stream)
|
||||
const procMatch = text.match(PROC_RE);
|
||||
if (procMatch) {
|
||||
const procName = procMatch[1] || inStreamProcName;
|
||||
if (procName) {
|
||||
results.procs.push({ name: procName.toUpperCase(), line: lineNum, isInStream: true });
|
||||
}
|
||||
inStreamProcName = procName?.toUpperCase() || '';
|
||||
continue;
|
||||
}
|
||||
|
||||
// PEND (end of in-stream proc)
|
||||
if (PEND_RE.test(text)) {
|
||||
inStreamProcName = '';
|
||||
continue;
|
||||
}
|
||||
|
||||
// JCLLIB ORDER=
|
||||
const jcllibMatch = text.match(JCLLIB_RE);
|
||||
if (jcllibMatch) {
|
||||
const libs = jcllibMatch[1].split(',').map(s => s.trim().replace(/'/g, ''));
|
||||
results.jcllib.push({ order: libs, line: lineNum });
|
||||
continue;
|
||||
}
|
||||
|
||||
// IF/ELSE/ENDIF
|
||||
const ifMatch = text.match(IF_RE);
|
||||
if (ifMatch) {
|
||||
results.conditionals.push({ type: 'IF', condition: ifMatch[1].trim(), line: lineNum });
|
||||
continue;
|
||||
}
|
||||
if (ELSE_RE.test(text)) {
|
||||
results.conditionals.push({ type: 'ELSE', line: lineNum });
|
||||
continue;
|
||||
}
|
||||
if (ENDIF_RE.test(text)) {
|
||||
results.conditionals.push({ type: 'ENDIF', line: lineNum });
|
||||
continue;
|
||||
}
|
||||
|
||||
// INCLUDE MEMBER=
|
||||
const includeMatch = text.match(INCLUDE_RE);
|
||||
if (includeMatch) {
|
||||
results.includes.push({ member: includeMatch[1].toUpperCase(), line: lineNum });
|
||||
continue;
|
||||
}
|
||||
|
||||
// SET var=value
|
||||
const setMatch = text.match(SET_RE);
|
||||
if (setMatch) {
|
||||
results.sets.push({
|
||||
variable: setMatch[1].toUpperCase(),
|
||||
value: setMatch[2].trim().replace(/,\s*$/, ''),
|
||||
line: lineNum,
|
||||
});
|
||||
continue;
|
||||
}
|
||||
|
||||
// JOB statement
|
||||
const jobMatch = text.match(JOB_RE);
|
||||
if (jobMatch) {
|
||||
currentJobName = jobMatch[1].toUpperCase();
|
||||
const params = jobMatch[2];
|
||||
results.jobs.push({
|
||||
name: currentJobName,
|
||||
line: lineNum,
|
||||
class: extractParam(params, 'CLASS'),
|
||||
msgclass: extractParam(params, 'MSGCLASS'),
|
||||
});
|
||||
continue;
|
||||
}
|
||||
|
||||
// EXEC statement
|
||||
const execMatch = text.match(EXEC_RE);
|
||||
if (execMatch) {
|
||||
currentStepName = execMatch[1].toUpperCase();
|
||||
const params = execMatch[2];
|
||||
const pgm = extractPgm(params);
|
||||
const proc = pgm ? undefined : extractProc(params);
|
||||
|
||||
results.steps.push({
|
||||
name: currentStepName,
|
||||
jobName: currentJobName,
|
||||
program: pgm?.toUpperCase(),
|
||||
proc: proc?.toUpperCase(),
|
||||
line: lineNum,
|
||||
});
|
||||
continue;
|
||||
}
|
||||
|
||||
// DD statement
|
||||
const ddMatch = text.match(DD_RE);
|
||||
if (ddMatch) {
|
||||
const ddName = ddMatch[1].toUpperCase();
|
||||
const params = ddMatch[2];
|
||||
results.ddStatements.push({
|
||||
ddName,
|
||||
stepName: currentStepName,
|
||||
dataset: extractDsn(params)?.toUpperCase(),
|
||||
disp: extractDisp(params)?.toUpperCase(),
|
||||
line: lineNum,
|
||||
});
|
||||
continue;
|
||||
}
|
||||
}
|
||||
|
||||
return results;
|
||||
}
|
||||
274
gitnexus/src/core/ingestion/cobol/jcl-processor.ts
Normal file
274
gitnexus/src/core/ingestion/cobol/jcl-processor.ts
Normal file
|
|
@ -0,0 +1,274 @@
|
|||
/**
|
||||
* JCL Processor — Converts JCL parse results into graph nodes and edges.
|
||||
*
|
||||
* Maps JCL entities to existing graph types (no new tables):
|
||||
* - Job -> CodeElement (description: "jcl-job class:A msgclass:X")
|
||||
* - Step -> CodeElement (description: "jcl-step pgm:PROGRAMNAME")
|
||||
* - Dataset -> CodeElement (description: "jcl-dataset disp:SHR")
|
||||
* - PROC -> Module
|
||||
*
|
||||
* Edges:
|
||||
* - Job CONTAINS Step
|
||||
* - Step CALLS Module (when PGM= matches an indexed program)
|
||||
* - Step references Dataset (CALLS edge with reason "jcl-dd")
|
||||
* - Job/Step IMPORTS PROC
|
||||
*
|
||||
* Pattern follows detectCrossProgamContracts() in pipeline.ts.
|
||||
*/
|
||||
|
||||
import { parseJcl, type JclParseResults } from './jcl-parser.js';
|
||||
import type { KnowledgeGraph } from '../../graph/types.js';
|
||||
import { generateId } from '../../../lib/utils.js';
|
||||
|
||||
export interface JclProcessResult {
|
||||
jobCount: number;
|
||||
stepCount: number;
|
||||
datasetCount: number;
|
||||
programLinks: number;
|
||||
}
|
||||
|
||||
/**
|
||||
* Process JCL files and integrate into the knowledge graph.
|
||||
*
|
||||
* @param graph - The in-memory knowledge graph
|
||||
* @param jclPaths - File paths of JCL files
|
||||
* @param jclContents - Map of path -> file content
|
||||
* @returns Summary of what was added
|
||||
*/
|
||||
export function processJclFiles(
|
||||
graph: KnowledgeGraph,
|
||||
jclPaths: string[],
|
||||
jclContents: Map<string, string>,
|
||||
): JclProcessResult {
|
||||
let jobCount = 0;
|
||||
let stepCount = 0;
|
||||
let datasetCount = 0;
|
||||
let programLinks = 0;
|
||||
|
||||
// Collect all Module names for step -> program linking
|
||||
const moduleNames = new Map<string, string>(); // uppercase name -> node id
|
||||
graph.forEachNode(node => {
|
||||
if (node.label === 'Module') {
|
||||
const nodeName = node.properties.name;
|
||||
if (typeof nodeName === 'string') {
|
||||
moduleNames.set(nodeName.toUpperCase(), node.id);
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
for (const filePath of jclPaths) {
|
||||
const content = jclContents.get(filePath);
|
||||
if (!content) continue;
|
||||
|
||||
const parsed = parseJcl(content, filePath);
|
||||
const result = integrateJclResults(graph, parsed, filePath, moduleNames);
|
||||
|
||||
jobCount += result.jobCount;
|
||||
stepCount += result.stepCount;
|
||||
datasetCount += result.datasetCount;
|
||||
programLinks += result.programLinks;
|
||||
}
|
||||
|
||||
return { jobCount, stepCount, datasetCount, programLinks };
|
||||
}
|
||||
|
||||
function integrateJclResults(
|
||||
graph: KnowledgeGraph,
|
||||
parsed: JclParseResults,
|
||||
filePath: string,
|
||||
moduleNames: Map<string, string>,
|
||||
): JclProcessResult {
|
||||
let jobCount = 0;
|
||||
let stepCount = 0;
|
||||
let datasetCount = 0;
|
||||
let programLinks = 0;
|
||||
|
||||
// Track step node IDs for DD -> step linking
|
||||
const stepNodeIds = new Map<string, string>(); // stepName -> nodeId
|
||||
|
||||
// 1. Create Job nodes
|
||||
for (const job of parsed.jobs) {
|
||||
const jobId = generateId('CodeElement', `${filePath}:job:${job.name}`);
|
||||
const classPart = job.class ? ` class:${job.class}` : '';
|
||||
const msgPart = job.msgclass ? ` msgclass:${job.msgclass}` : '';
|
||||
|
||||
graph.addNode({
|
||||
id: jobId,
|
||||
label: 'CodeElement',
|
||||
properties: {
|
||||
name: job.name,
|
||||
filePath,
|
||||
startLine: job.line,
|
||||
endLine: job.line,
|
||||
description: `jcl-job${classPart}${msgPart}`,
|
||||
},
|
||||
});
|
||||
|
||||
// Link File -> Job (CONTAINS)
|
||||
const fileId = generateId('File', filePath);
|
||||
graph.addRelationship({
|
||||
id: `${fileId}_contains_${jobId}`,
|
||||
type: 'CONTAINS',
|
||||
sourceId: fileId,
|
||||
targetId: jobId,
|
||||
confidence: 1.0,
|
||||
reason: 'jcl-job',
|
||||
});
|
||||
|
||||
jobCount++;
|
||||
}
|
||||
|
||||
// 1.5 Pre-register in-stream PROCs so steps can reference them
|
||||
// (fixes ordering bug: steps processed before PROCs were registered)
|
||||
for (const proc of parsed.procs) {
|
||||
const procId = generateId('Module', `${filePath}:proc:${proc.name}`);
|
||||
moduleNames.set(proc.name.toUpperCase(), procId);
|
||||
}
|
||||
|
||||
// 2. Create Step nodes and link to programs
|
||||
for (const step of parsed.steps) {
|
||||
const stepId = generateId('CodeElement', `${filePath}:step:${step.jobName}:${step.name}`);
|
||||
const pgmPart = step.program ? ` pgm:${step.program}` : '';
|
||||
const procPart = step.proc ? ` proc:${step.proc}` : '';
|
||||
|
||||
graph.addNode({
|
||||
id: stepId,
|
||||
label: 'CodeElement',
|
||||
properties: {
|
||||
name: step.name,
|
||||
filePath,
|
||||
startLine: step.line,
|
||||
endLine: step.line,
|
||||
description: `jcl-step${pgmPart}${procPart}`,
|
||||
},
|
||||
});
|
||||
|
||||
stepNodeIds.set(step.name, stepId);
|
||||
|
||||
// Link Job -> Step (CONTAINS)
|
||||
if (step.jobName) {
|
||||
const jobId = generateId('CodeElement', `${filePath}:job:${step.jobName}`);
|
||||
graph.addRelationship({
|
||||
id: `${jobId}_contains_${stepId}`,
|
||||
type: 'CONTAINS',
|
||||
sourceId: jobId,
|
||||
targetId: stepId,
|
||||
confidence: 1.0,
|
||||
reason: 'jcl-step',
|
||||
});
|
||||
}
|
||||
|
||||
// Link Step -> Module (CALLS) when PGM= matches an indexed program
|
||||
if (step.program) {
|
||||
const moduleId = moduleNames.get(step.program.toUpperCase());
|
||||
if (moduleId) {
|
||||
graph.addRelationship({
|
||||
id: `${stepId}_calls_${moduleId}`,
|
||||
type: 'CALLS',
|
||||
sourceId: stepId,
|
||||
targetId: moduleId,
|
||||
confidence: 0.95,
|
||||
reason: 'jcl-exec-pgm',
|
||||
});
|
||||
programLinks++;
|
||||
}
|
||||
}
|
||||
|
||||
// Link Step -> PROC (CALLS) — PROC as Module
|
||||
if (step.proc) {
|
||||
const procModuleId = moduleNames.get(step.proc.toUpperCase());
|
||||
if (procModuleId) {
|
||||
graph.addRelationship({
|
||||
id: `${stepId}_calls_proc_${procModuleId}`,
|
||||
type: 'CALLS',
|
||||
sourceId: stepId,
|
||||
targetId: procModuleId,
|
||||
confidence: 0.9,
|
||||
reason: 'jcl-exec-proc',
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
stepCount++;
|
||||
}
|
||||
|
||||
// 3. Create Dataset nodes from DD statements
|
||||
const seenDatasets = new Set<string>();
|
||||
for (const dd of parsed.ddStatements) {
|
||||
if (!dd.dataset) continue;
|
||||
|
||||
// Create dataset node (deduplicated per file)
|
||||
const datasetKey = `${filePath}:dataset:${dd.dataset}`;
|
||||
const datasetId = generateId('CodeElement', datasetKey);
|
||||
|
||||
if (!seenDatasets.has(dd.dataset)) {
|
||||
const dispPart = dd.disp ? ` disp:${dd.disp}` : '';
|
||||
graph.addNode({
|
||||
id: datasetId,
|
||||
label: 'CodeElement',
|
||||
properties: {
|
||||
name: dd.dataset,
|
||||
filePath,
|
||||
startLine: dd.line,
|
||||
endLine: dd.line,
|
||||
|
||||
description: `jcl-dataset${dispPart}`,
|
||||
},
|
||||
});
|
||||
seenDatasets.add(dd.dataset);
|
||||
datasetCount++;
|
||||
}
|
||||
|
||||
// Link Step -> Dataset (CALLS with reason jcl-dd)
|
||||
const stepId = stepNodeIds.get(dd.stepName);
|
||||
if (stepId) {
|
||||
graph.addRelationship({
|
||||
id: `${stepId}_dd_${dd.ddName}_${datasetId}`,
|
||||
type: 'CALLS',
|
||||
sourceId: stepId,
|
||||
targetId: datasetId,
|
||||
confidence: 0.85,
|
||||
reason: `jcl-dd:${dd.ddName}`,
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
// 4. Create PROC nodes (in-stream procs as Module)
|
||||
for (const proc of parsed.procs) {
|
||||
if (!proc.isInStream) continue;
|
||||
|
||||
const procId = generateId('Module', `${filePath}:proc:${proc.name}`);
|
||||
graph.addNode({
|
||||
id: procId,
|
||||
label: 'Module',
|
||||
properties: {
|
||||
name: proc.name,
|
||||
filePath,
|
||||
startLine: proc.line,
|
||||
endLine: proc.line,
|
||||
description: 'jcl-proc-instream',
|
||||
},
|
||||
});
|
||||
|
||||
// Register for step linking
|
||||
moduleNames.set(proc.name.toUpperCase(), procId);
|
||||
}
|
||||
|
||||
// 5. INCLUDE directives -> IMPORTS edges
|
||||
for (const inc of parsed.includes) {
|
||||
const moduleId = moduleNames.get(inc.member.toUpperCase());
|
||||
if (moduleId) {
|
||||
const fileId = generateId('File', filePath);
|
||||
graph.addRelationship({
|
||||
id: `${fileId}_includes_${moduleId}`,
|
||||
type: 'IMPORTS',
|
||||
sourceId: fileId,
|
||||
targetId: moduleId,
|
||||
confidence: 0.9,
|
||||
reason: 'jcl-include',
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
return { jobCount, stepCount, datasetCount, programLinks };
|
||||
}
|
||||
|
|
@ -226,6 +226,7 @@ export const ENTRY_POINT_PATTERNS = {
|
|||
/^onEvent$/, // BLoC event handler
|
||||
/^mapEventToState$/, // Legacy BLoC pattern
|
||||
],
|
||||
[SupportedLanguages.Cobol]: [], // Standalone regex processor — no tree-sitter entry points
|
||||
} satisfies Record<SupportedLanguages, RegExp[]>;
|
||||
|
||||
/** Pre-computed merged patterns (universal + language-specific) to avoid per-call array allocation. */
|
||||
|
|
@ -325,7 +326,7 @@ export function calculateEntryPointScore(
|
|||
// Check positive patterns
|
||||
const allPatterns = MERGED_ENTRY_POINT_PATTERNS[language];
|
||||
|
||||
if (allPatterns.some(p => p.test(name))) {
|
||||
if (allPatterns?.some(p => p.test(name))) {
|
||||
nameMultiplier = 1.5; // Bonus for matching entry point pattern
|
||||
reasons.push('entry-pattern');
|
||||
}
|
||||
|
|
|
|||
|
|
@ -601,6 +601,7 @@ export const AST_FRAMEWORK_PATTERNS_BY_LANGUAGE = {
|
|||
{ framework: 'flutter', entryPointMultiplier: 2.5, reason: 'flutter-widget', patterns: FRAMEWORK_AST_PATTERNS.flutter },
|
||||
{ framework: 'riverpod', entryPointMultiplier: 2.8, reason: 'riverpod-pattern', patterns: FRAMEWORK_AST_PATTERNS.riverpod },
|
||||
],
|
||||
[SupportedLanguages.Cobol]: [], // Standalone regex processor — no AST framework patterns
|
||||
} satisfies Record<SupportedLanguages, AstFrameworkPatternConfig[]>;
|
||||
|
||||
/** Pre-lowercased patterns for O(1) pattern matching at runtime */
|
||||
|
|
|
|||
|
|
@ -39,6 +39,12 @@ export function resolveDartImport(
|
|||
return null;
|
||||
}
|
||||
|
||||
// Relative imports — use standard resolution
|
||||
return resolveStandard(stripped, filePath, ctx, SupportedLanguages.Dart);
|
||||
// Relative imports — use standard resolution.
|
||||
// Dart relative imports don't require a leading "./" (e.g. `import 'models.dart'`).
|
||||
// The standard resolver only recognises paths starting with "." as relative, so
|
||||
// prepend "./" when the path doesn't already start with "." to ensure correct
|
||||
// same-directory resolution (without this, "models.dart" would be mangled by the
|
||||
// generic dot-to-slash conversion intended for Java-style package imports).
|
||||
const relPath = stripped.startsWith('.') ? stripped : './' + stripped;
|
||||
return resolveStandard(relPath, filePath, ctx, SupportedLanguages.Dart);
|
||||
}
|
||||
|
|
|
|||
|
|
@ -41,7 +41,12 @@ interface LanguageProviderConfig {
|
|||
readonly extensions: readonly string[];
|
||||
|
||||
// ── Parser ────────────────────────────────────────────────────────
|
||||
/** Tree-sitter query strings for definitions, imports, calls, heritage */
|
||||
/** Parse strategy: 'tree-sitter' (default) uses AST parsing via tree-sitter.
|
||||
* 'standalone' means the language has its own regex-based processor and
|
||||
* should be skipped by the tree-sitter pipeline (e.g., COBOL, Markdown). */
|
||||
readonly parseStrategy?: 'tree-sitter' | 'standalone';
|
||||
/** Tree-sitter query strings for definitions, imports, calls, heritage.
|
||||
* Required for tree-sitter languages; empty string for standalone processors. */
|
||||
readonly treeSitterQueries: string;
|
||||
|
||||
// ── Core (required) ───────────────────────────────────────────────
|
||||
|
|
@ -80,6 +85,17 @@ interface LanguageProviderConfig {
|
|||
projectConfig: unknown,
|
||||
) => void;
|
||||
|
||||
// ── Enclosing function resolution ───────────────────────────────
|
||||
/** Resolve the enclosing function name + label from an AST ancestor node
|
||||
* that is NOT a standard FUNCTION_NODE_TYPE. For languages where the
|
||||
* function body is a sibling of the signature (e.g. Dart: function_body ↔
|
||||
* function_signature are siblings under program/class_body), the default
|
||||
* parent walk cannot find the enclosing function. This hook lets the
|
||||
* language provider inspect each ancestor and return the resolved result.
|
||||
* Return null to continue the default walk.
|
||||
* Default: undefined (standard parent walk only). */
|
||||
readonly enclosingFunctionFinder?: (ancestorNode: SyntaxNode) => { funcName: string; label: NodeLabel } | null;
|
||||
|
||||
// ── Labels ────────────────────────────────────────────────────────
|
||||
/** Override the default node label for definition.function captures.
|
||||
* Return null to skip (C/C++ duplicate), a different label to reclassify
|
||||
|
|
@ -115,6 +131,11 @@ interface LanguageProviderConfig {
|
|||
* When true, the worker extracts routes via the language's route extraction logic.
|
||||
* Default: undefined (no route files). */
|
||||
readonly isRouteFile?: (filePath: string) => boolean;
|
||||
|
||||
// ── Noise filtering ────────────────────────────────────────────────
|
||||
/** Built-in/stdlib names that should be filtered from the call graph for this language.
|
||||
* Default: undefined (no language-specific filtering). */
|
||||
readonly builtInNames?: ReadonlySet<string>;
|
||||
}
|
||||
|
||||
/** Runtime type — same as LanguageProviderConfig but with defaults guaranteed present. */
|
||||
|
|
@ -124,6 +145,8 @@ export interface LanguageProvider extends Omit<LanguageProviderConfig,
|
|||
readonly importSemantics: ImportSemantics;
|
||||
readonly heritageDefaultEdge: 'EXTENDS' | 'IMPLEMENTS';
|
||||
readonly mroStrategy: MroStrategy;
|
||||
/** Check if a name is a built-in/stdlib function that should be filtered from the call graph. */
|
||||
readonly isBuiltInName: (name: string) => boolean;
|
||||
}
|
||||
|
||||
const DEFAULTS: Pick<LanguageProvider, 'importSemantics' | 'heritageDefaultEdge' | 'mroStrategy'> = {
|
||||
|
|
@ -134,5 +157,10 @@ const DEFAULTS: Pick<LanguageProvider, 'importSemantics' | 'heritageDefaultEdge'
|
|||
|
||||
/** Define a language provider — required fields must be supplied, optional fields get sensible defaults. */
|
||||
export function defineLanguage(config: LanguageProviderConfig): LanguageProvider {
|
||||
return { ...DEFAULTS, ...config };
|
||||
const builtIns = config.builtInNames;
|
||||
return {
|
||||
...DEFAULTS,
|
||||
...config,
|
||||
isBuiltInName: builtIns ? (name: string) => builtIns.has(name) : () => false,
|
||||
};
|
||||
}
|
||||
|
|
|
|||
|
|
@ -20,6 +20,28 @@ import type { LanguageProvider } from '../language-provider.js';
|
|||
import { createFieldExtractor } from '../field-extractors/generic.js';
|
||||
import { cConfig as cFieldConfig, cppConfig as cppFieldConfig } from '../field-extractors/configs/c-cpp.js';
|
||||
|
||||
const C_BUILT_INS: ReadonlySet<string> = new Set([
|
||||
'printf', 'fprintf', 'sprintf', 'snprintf', 'vprintf', 'vfprintf', 'vsprintf', 'vsnprintf',
|
||||
'scanf', 'fscanf', 'sscanf',
|
||||
'malloc', 'calloc', 'realloc', 'free', 'memcpy', 'memmove', 'memset', 'memcmp',
|
||||
'strlen', 'strcpy', 'strncpy', 'strcat', 'strncat', 'strcmp', 'strncmp', 'strstr', 'strchr', 'strrchr',
|
||||
'atoi', 'atol', 'atof', 'strtol', 'strtoul', 'strtoll', 'strtoull', 'strtod',
|
||||
'sizeof', 'offsetof', 'typeof',
|
||||
'assert', 'abort', 'exit', '_exit',
|
||||
'fopen', 'fclose', 'fread', 'fwrite', 'fseek', 'ftell', 'rewind', 'fflush', 'fgets', 'fputs',
|
||||
'likely', 'unlikely', 'BUG', 'BUG_ON', 'WARN', 'WARN_ON', 'WARN_ONCE',
|
||||
'IS_ERR', 'PTR_ERR', 'ERR_PTR', 'IS_ERR_OR_NULL',
|
||||
'ARRAY_SIZE', 'container_of', 'list_for_each_entry', 'list_for_each_entry_safe',
|
||||
'min', 'max', 'clamp', 'abs', 'swap',
|
||||
'pr_info', 'pr_warn', 'pr_err', 'pr_debug', 'pr_notice', 'pr_crit', 'pr_emerg',
|
||||
'printk', 'dev_info', 'dev_warn', 'dev_err', 'dev_dbg',
|
||||
'GFP_KERNEL', 'GFP_ATOMIC',
|
||||
'spin_lock', 'spin_unlock', 'spin_lock_irqsave', 'spin_unlock_irqrestore',
|
||||
'mutex_lock', 'mutex_unlock', 'mutex_init',
|
||||
'kfree', 'kmalloc', 'kzalloc', 'kcalloc', 'krealloc', 'kvmalloc', 'kvfree',
|
||||
'get', 'put',
|
||||
]);
|
||||
|
||||
/** Label override shared by C and C++: skip function_definition captures inside class/struct
|
||||
* bodies (they're duplicates of definition.method captures). */
|
||||
const cppLabelOverride: NonNullable<LanguageProvider['labelOverride']> = (functionNode, defaultLabel) => {
|
||||
|
|
@ -37,6 +59,7 @@ export const cProvider = defineLanguage({
|
|||
importSemantics: 'wildcard',
|
||||
fieldExtractor: createFieldExtractor(cFieldConfig),
|
||||
labelOverride: cppLabelOverride,
|
||||
builtInNames: C_BUILT_INS,
|
||||
});
|
||||
|
||||
export const cppProvider = defineLanguage({
|
||||
|
|
@ -50,4 +73,5 @@ export const cppProvider = defineLanguage({
|
|||
mroStrategy: 'leftmost-base',
|
||||
fieldExtractor: createFieldExtractor(cppFieldConfig),
|
||||
labelOverride: cppLabelOverride,
|
||||
builtInNames: C_BUILT_INS,
|
||||
});
|
||||
|
|
|
|||
27
gitnexus/src/core/ingestion/languages/cobol.ts
Normal file
27
gitnexus/src/core/ingestion/languages/cobol.ts
Normal file
|
|
@ -0,0 +1,27 @@
|
|||
/**
|
||||
* COBOL Language Provider
|
||||
*
|
||||
* Standalone regex-based processor — no tree-sitter grammar.
|
||||
* COBOL files (.cbl, .cob, .cobol, .cpy, .copybook) are detected and
|
||||
* processed by cobol-processor.ts in pipeline Phase 2.6, not by the
|
||||
* tree-sitter pipeline.
|
||||
*
|
||||
* This provider exists to satisfy the SupportedLanguages exhaustiveness
|
||||
* checks and to declare parseStrategy: 'standalone'.
|
||||
*/
|
||||
import { SupportedLanguages } from '../../../config/supported-languages.js';
|
||||
import { defineLanguage } from '../language-provider.js';
|
||||
|
||||
export const cobolProvider = defineLanguage({
|
||||
id: SupportedLanguages.Cobol,
|
||||
parseStrategy: 'standalone',
|
||||
extensions: [], // COBOL files detected by cobol-processor's isCobolFile/isJclFile
|
||||
treeSitterQueries: '',
|
||||
typeConfig: {
|
||||
declarationNodeTypes: new Set(),
|
||||
extractDeclaration: () => null,
|
||||
extractParameter: () => null,
|
||||
},
|
||||
exportChecker: () => false,
|
||||
importResolver: () => null,
|
||||
});
|
||||
|
|
@ -16,6 +16,27 @@ import { CSHARP_QUERIES } from '../tree-sitter-queries.js';
|
|||
import { createFieldExtractor } from '../field-extractors/generic.js';
|
||||
import { csharpConfig as csharpFieldConfig } from '../field-extractors/configs/csharp.js';
|
||||
|
||||
const BUILT_INS: ReadonlySet<string> = new Set([
|
||||
'Console', 'WriteLine', 'ReadLine', 'Write',
|
||||
'Task', 'Run', 'Wait', 'WhenAll', 'WhenAny', 'FromResult', 'Delay', 'ContinueWith',
|
||||
'ConfigureAwait', 'GetAwaiter', 'GetResult',
|
||||
'ToString', 'GetType', 'Equals', 'GetHashCode', 'ReferenceEquals',
|
||||
'Add', 'Remove', 'Contains', 'Clear', 'Count', 'Any', 'All',
|
||||
'Where', 'Select', 'SelectMany', 'OrderBy', 'OrderByDescending', 'GroupBy',
|
||||
'First', 'FirstOrDefault', 'Single', 'SingleOrDefault', 'Last', 'LastOrDefault',
|
||||
'ToList', 'ToArray', 'ToDictionary', 'AsEnumerable', 'AsQueryable',
|
||||
'Aggregate', 'Sum', 'Average', 'Min', 'Max', 'Distinct', 'Skip', 'Take',
|
||||
'String', 'Format', 'IsNullOrEmpty', 'IsNullOrWhiteSpace', 'Concat', 'Join',
|
||||
'Trim', 'TrimStart', 'TrimEnd', 'Split', 'Replace', 'StartsWith', 'EndsWith',
|
||||
'Convert', 'ToInt32', 'ToDouble', 'ToBoolean', 'ToByte',
|
||||
'Math', 'Abs', 'Ceiling', 'Floor', 'Round', 'Pow', 'Sqrt',
|
||||
'Dispose', 'Close',
|
||||
'TryParse', 'Parse',
|
||||
'AddRange', 'RemoveAt', 'RemoveAll', 'FindAll', 'Exists', 'TrueForAll',
|
||||
'ContainsKey', 'TryGetValue', 'AddOrUpdate',
|
||||
'Throw', 'ThrowIfNull',
|
||||
]);
|
||||
|
||||
export const csharpProvider = defineLanguage({
|
||||
id: SupportedLanguages.CSharp,
|
||||
extensions: ['.cs'],
|
||||
|
|
@ -27,4 +48,5 @@ export const csharpProvider = defineLanguage({
|
|||
interfaceNamePattern: /^I[A-Z]/,
|
||||
mroStrategy: 'implements-split',
|
||||
fieldExtractor: createFieldExtractor(csharpFieldConfig),
|
||||
builtInNames: BUILT_INS,
|
||||
});
|
||||
|
|
|
|||
|
|
@ -5,8 +5,14 @@
|
|||
* - importSemantics: 'wildcard' (Dart imports bring everything public into scope)
|
||||
* - exportChecker: public if no leading underscore
|
||||
* - Dart SDK imports (dart:*) and external packages are skipped
|
||||
* - enclosingFunctionFinder: Dart's tree-sitter grammar places function_body
|
||||
* as a sibling of function_signature/method_signature (not as a child).
|
||||
* The hook resolves the enclosing function by inspecting the previous sibling.
|
||||
*/
|
||||
|
||||
import type { SyntaxNode } from '../utils/ast-helpers.js';
|
||||
import type { NodeLabel } from '../../graph/types.js';
|
||||
import { FUNCTION_NODE_TYPES, extractFunctionName } from '../utils/ast-helpers.js';
|
||||
import { SupportedLanguages } from '../../../config/supported-languages.js';
|
||||
import { defineLanguage } from '../language-provider.js';
|
||||
import { typeConfig as dartConfig } from '../type-extractors/dart.js';
|
||||
|
|
@ -16,6 +22,32 @@ import { DART_QUERIES } from '../tree-sitter-queries.js';
|
|||
import { createFieldExtractor } from '../field-extractors/generic.js';
|
||||
import { dartConfig as dartFieldConfig } from '../field-extractors/configs/dart.js';
|
||||
|
||||
/**
|
||||
* Resolve the enclosing function from a `function_body` node by looking at its
|
||||
* previous sibling. In Dart's tree-sitter grammar, function_signature and
|
||||
* function_body are siblings under program or class_body, unlike most languages
|
||||
* where the function declaration wraps both.
|
||||
*
|
||||
* Delegates name extraction to the shared `extractFunctionName` which already
|
||||
* handles Dart's function_signature and method_signature node types.
|
||||
*/
|
||||
const dartEnclosingFunctionFinder = (node: SyntaxNode): { funcName: string; label: NodeLabel } | null => {
|
||||
if (node.type !== 'function_body') return null;
|
||||
const prev = node.previousSibling;
|
||||
if (!prev || !FUNCTION_NODE_TYPES.has(prev.type)) return null;
|
||||
const { funcName, label } = extractFunctionName(prev);
|
||||
return funcName ? { funcName, label } : null;
|
||||
};
|
||||
|
||||
const BUILT_INS: ReadonlySet<string> = new Set([
|
||||
'setState', 'mounted', 'debugPrint',
|
||||
'runApp', 'showDialog', 'showModalBottomSheet',
|
||||
'Navigator', 'push', 'pushNamed', 'pushReplacement', 'pop', 'maybePop',
|
||||
'ScaffoldMessenger', 'showSnackBar',
|
||||
'deactivate', 'reassemble', 'debugDumpApp', 'debugDumpRenderTree',
|
||||
'then', 'catchError', 'whenComplete', 'listen',
|
||||
]);
|
||||
|
||||
export const dartProvider = defineLanguage({
|
||||
id: SupportedLanguages.Dart,
|
||||
extensions: ['.dart'],
|
||||
|
|
@ -25,4 +57,6 @@ export const dartProvider = defineLanguage({
|
|||
importResolver: resolveDartImport,
|
||||
importSemantics: 'wildcard',
|
||||
fieldExtractor: createFieldExtractor(dartFieldConfig),
|
||||
enclosingFunctionFinder: dartEnclosingFunctionFinder,
|
||||
builtInNames: BUILT_INS,
|
||||
});
|
||||
|
|
|
|||
|
|
@ -23,6 +23,7 @@ import { phpProvider } from './php.js';
|
|||
import { rubyProvider } from './ruby.js';
|
||||
import { swiftProvider } from './swift.js';
|
||||
import { dartProvider } from './dart.js';
|
||||
import { cobolProvider } from './cobol.js';
|
||||
|
||||
export const providers = {
|
||||
[SupportedLanguages.JavaScript]: javascriptProvider,
|
||||
|
|
@ -39,6 +40,7 @@ export const providers = {
|
|||
[SupportedLanguages.Ruby]: rubyProvider,
|
||||
[SupportedLanguages.Swift]: swiftProvider,
|
||||
[SupportedLanguages.Dart]: dartProvider,
|
||||
[SupportedLanguages.Cobol]: cobolProvider,
|
||||
} satisfies Record<SupportedLanguages, LanguageProvider>;
|
||||
|
||||
/** Get provider by language enum (always succeeds for SupportedLanguages). */
|
||||
|
|
|
|||
|
|
@ -19,6 +19,21 @@ import { isKotlinClassMethod } from '../utils/ast-helpers.js';
|
|||
import { createFieldExtractor } from '../field-extractors/generic.js';
|
||||
import { kotlinConfig } from '../field-extractors/configs/jvm.js';
|
||||
|
||||
const BUILT_INS: ReadonlySet<string> = new Set([
|
||||
'println', 'print', 'readLine', 'require', 'requireNotNull', 'check', 'assert', 'lazy', 'error',
|
||||
'listOf', 'mapOf', 'setOf', 'mutableListOf', 'mutableMapOf', 'mutableSetOf',
|
||||
'arrayOf', 'sequenceOf', 'also', 'apply', 'run', 'with', 'takeIf', 'takeUnless',
|
||||
'TODO', 'buildString', 'buildList', 'buildMap', 'buildSet',
|
||||
'repeat', 'synchronized',
|
||||
'launch', 'async', 'runBlocking', 'withContext', 'coroutineScope',
|
||||
'supervisorScope', 'delay',
|
||||
'flow', 'flowOf', 'collect', 'emit', 'onEach', 'catch',
|
||||
'buffer', 'conflate', 'distinctUntilChanged',
|
||||
'flatMapLatest', 'flatMapMerge', 'combine',
|
||||
'stateIn', 'shareIn', 'launchIn',
|
||||
'to', 'until', 'downTo', 'step',
|
||||
]);
|
||||
|
||||
export const kotlinProvider = defineLanguage({
|
||||
id: SupportedLanguages.Kotlin,
|
||||
extensions: ['.kt', '.kts'],
|
||||
|
|
@ -30,6 +45,7 @@ export const kotlinProvider = defineLanguage({
|
|||
importPathPreprocessor: appendKotlinWildcard,
|
||||
mroStrategy: 'implements-split',
|
||||
fieldExtractor: createFieldExtractor(kotlinConfig),
|
||||
builtInNames: BUILT_INS,
|
||||
labelOverride: (functionNode, defaultLabel) => {
|
||||
if (defaultLabel !== 'Function') return defaultLabel;
|
||||
if (isKotlinClassMethod(functionNode)) return 'Method';
|
||||
|
|
|
|||
|
|
@ -18,6 +18,25 @@ import type { NodeLabel } from '../../graph/types.js';
|
|||
import { createFieldExtractor } from '../field-extractors/generic.js';
|
||||
import { phpConfig as phpFieldConfig } from '../field-extractors/configs/php.js';
|
||||
|
||||
const BUILT_INS: ReadonlySet<string> = new Set([
|
||||
'echo', 'isset', 'empty', 'unset', 'list', 'array', 'compact', 'extract',
|
||||
'count', 'strlen', 'strpos', 'strrpos', 'substr', 'strtolower', 'strtoupper', 'trim',
|
||||
'ltrim', 'rtrim', 'str_replace', 'str_contains', 'str_starts_with', 'str_ends_with',
|
||||
'sprintf', 'vsprintf', 'printf', 'number_format',
|
||||
'array_map', 'array_filter', 'array_reduce', 'array_push', 'array_pop', 'array_shift',
|
||||
'array_unshift', 'array_slice', 'array_splice', 'array_merge', 'array_keys', 'array_values',
|
||||
'array_key_exists', 'in_array', 'array_search', 'array_unique', 'usort', 'rsort',
|
||||
'json_encode', 'json_decode', 'serialize', 'unserialize',
|
||||
'intval', 'floatval', 'strval', 'boolval', 'is_null', 'is_string', 'is_int', 'is_array',
|
||||
'is_object', 'is_numeric', 'is_bool', 'is_float',
|
||||
'var_dump', 'print_r', 'var_export',
|
||||
'date', 'time', 'strtotime', 'mktime', 'microtime',
|
||||
'file_exists', 'file_get_contents', 'file_put_contents', 'is_file', 'is_dir',
|
||||
'preg_match', 'preg_match_all', 'preg_replace', 'preg_split',
|
||||
'header', 'session_start', 'session_destroy', 'ob_start', 'ob_end_clean', 'ob_get_clean',
|
||||
'dd', 'dump',
|
||||
]);
|
||||
|
||||
/** Eloquent model properties whose array values are worth indexing. */
|
||||
const ELOQUENT_ARRAY_PROPS = new Set(['fillable', 'casts', 'hidden', 'guarded', 'with', 'appends']);
|
||||
|
||||
|
|
@ -133,4 +152,5 @@ export const phpProvider = defineLanguage({
|
|||
fieldExtractor: createFieldExtractor(phpFieldConfig),
|
||||
descriptionExtractor: phpDescriptionExtractor,
|
||||
isRouteFile: isPhpRouteFile,
|
||||
builtInNames: BUILT_INS,
|
||||
});
|
||||
|
|
|
|||
|
|
@ -20,6 +20,13 @@ import { PYTHON_QUERIES } from '../tree-sitter-queries.js';
|
|||
import { createFieldExtractor } from '../field-extractors/generic.js';
|
||||
import { pythonConfig as pythonFieldConfig } from '../field-extractors/configs/python.js';
|
||||
|
||||
const BUILT_INS: ReadonlySet<string> = new Set([
|
||||
'print', 'len', 'range', 'str', 'int', 'float', 'list', 'dict', 'set', 'tuple',
|
||||
'append', 'extend', 'update',
|
||||
'type', 'isinstance', 'issubclass', 'getattr', 'setattr', 'hasattr',
|
||||
'enumerate', 'zip', 'sorted', 'reversed', 'min', 'max', 'sum', 'abs',
|
||||
]);
|
||||
|
||||
export const pythonProvider = defineLanguage({
|
||||
id: SupportedLanguages.Python,
|
||||
extensions: ['.py'],
|
||||
|
|
@ -31,4 +38,5 @@ export const pythonProvider = defineLanguage({
|
|||
importSemantics: 'namespace',
|
||||
mroStrategy: 'c3',
|
||||
fieldExtractor: createFieldExtractor(pythonFieldConfig),
|
||||
builtInNames: BUILT_INS,
|
||||
});
|
||||
|
|
|
|||
|
|
@ -17,6 +17,22 @@ import { RUBY_QUERIES } from '../tree-sitter-queries.js';
|
|||
import { createFieldExtractor } from '../field-extractors/generic.js';
|
||||
import { rubyConfig as rubyFieldConfig } from '../field-extractors/configs/ruby.js';
|
||||
|
||||
const BUILT_INS: ReadonlySet<string> = new Set([
|
||||
'puts', 'p', 'pp', 'raise', 'fail',
|
||||
'require', 'require_relative', 'load', 'autoload',
|
||||
'include', 'extend', 'prepend',
|
||||
'attr_accessor', 'attr_reader', 'attr_writer',
|
||||
'public', 'private', 'protected', 'module_function',
|
||||
'lambda', 'proc', 'block_given?',
|
||||
'nil?', 'is_a?', 'kind_of?', 'instance_of?', 'respond_to?',
|
||||
'freeze', 'frozen?', 'dup', 'tap', 'yield_self',
|
||||
'each', 'select', 'reject', 'detect', 'collect',
|
||||
'inject', 'flat_map', 'each_with_object', 'each_with_index',
|
||||
'any?', 'all?', 'none?', 'count', 'first', 'last',
|
||||
'sort_by', 'min_by', 'max_by',
|
||||
'group_by', 'partition', 'compact', 'flatten', 'uniq',
|
||||
]);
|
||||
|
||||
export const rubyProvider = defineLanguage({
|
||||
id: SupportedLanguages.Ruby,
|
||||
extensions: ['.rb', '.rake', '.gemspec'],
|
||||
|
|
@ -27,4 +43,5 @@ export const rubyProvider = defineLanguage({
|
|||
callRouter: routeRubyCall,
|
||||
importSemantics: 'wildcard',
|
||||
fieldExtractor: createFieldExtractor(rubyFieldConfig),
|
||||
builtInNames: BUILT_INS,
|
||||
});
|
||||
|
|
|
|||
|
|
@ -20,6 +20,19 @@ import { RUST_QUERIES } from '../tree-sitter-queries.js';
|
|||
import { createFieldExtractor } from '../field-extractors/generic.js';
|
||||
import { rustConfig as rustFieldConfig } from '../field-extractors/configs/rust.js';
|
||||
|
||||
const BUILT_INS: ReadonlySet<string> = new Set([
|
||||
'unwrap', 'expect', 'unwrap_or', 'unwrap_or_else', 'unwrap_or_default',
|
||||
'ok', 'err', 'is_ok', 'is_err', 'map', 'map_err', 'and_then', 'or_else',
|
||||
'clone', 'to_string', 'to_owned', 'into', 'from', 'as_ref', 'as_mut',
|
||||
'iter', 'into_iter', 'collect', 'filter', 'fold', 'for_each',
|
||||
'len', 'is_empty', 'push', 'pop', 'insert', 'remove', 'contains',
|
||||
'format', 'write', 'writeln', 'panic', 'unreachable', 'todo', 'unimplemented',
|
||||
'vec', 'println', 'eprintln', 'dbg',
|
||||
'lock', 'read', 'try_lock',
|
||||
'spawn', 'join', 'sleep',
|
||||
'Some', 'None', 'Ok', 'Err',
|
||||
]);
|
||||
|
||||
export const rustProvider = defineLanguage({
|
||||
id: SupportedLanguages.Rust,
|
||||
extensions: ['.rs'],
|
||||
|
|
@ -30,4 +43,5 @@ export const rustProvider = defineLanguage({
|
|||
namedBindingExtractor: extractRustNamedBindings,
|
||||
mroStrategy: 'qualified-syntax',
|
||||
fieldExtractor: createFieldExtractor(rustFieldConfig),
|
||||
builtInNames: BUILT_INS,
|
||||
});
|
||||
|
|
|
|||
|
|
@ -103,6 +103,34 @@ function wireSwiftImplicitImports(
|
|||
}
|
||||
}
|
||||
|
||||
const BUILT_INS: ReadonlySet<string> = new Set([
|
||||
'print', 'debugPrint', 'dump', 'fatalError', 'precondition', 'preconditionFailure',
|
||||
'assert', 'assertionFailure', 'NSLog',
|
||||
'abs', 'min', 'max', 'zip', 'stride', 'sequence', 'repeatElement',
|
||||
'swap', 'withUnsafePointer', 'withUnsafeMutablePointer', 'withUnsafeBytes',
|
||||
'autoreleasepool', 'unsafeBitCast', 'unsafeDowncast', 'numericCast',
|
||||
'type', 'MemoryLayout',
|
||||
'map', 'flatMap', 'compactMap', 'filter', 'reduce', 'forEach', 'contains',
|
||||
'first', 'last', 'prefix', 'suffix', 'dropFirst', 'dropLast',
|
||||
'sorted', 'reversed', 'enumerated', 'joined', 'split',
|
||||
'append', 'insert', 'remove', 'removeAll', 'removeFirst', 'removeLast',
|
||||
'isEmpty', 'count', 'index', 'startIndex', 'endIndex',
|
||||
'addSubview', 'removeFromSuperview', 'layoutSubviews', 'setNeedsLayout',
|
||||
'layoutIfNeeded', 'setNeedsDisplay', 'invalidateIntrinsicContentSize',
|
||||
'addTarget', 'removeTarget', 'addGestureRecognizer',
|
||||
'addConstraint', 'addConstraints', 'removeConstraint', 'removeConstraints',
|
||||
'NSLocalizedString', 'Bundle',
|
||||
'reloadData', 'reloadSections', 'reloadRows', 'performBatchUpdates',
|
||||
'register', 'dequeueReusableCell', 'dequeueReusableSupplementaryView',
|
||||
'beginUpdates', 'endUpdates', 'insertRows', 'deleteRows', 'insertSections', 'deleteSections',
|
||||
'present', 'dismiss', 'pushViewController', 'popViewController', 'popToRootViewController',
|
||||
'performSegue', 'prepare',
|
||||
'DispatchQueue', 'async', 'sync', 'asyncAfter',
|
||||
'Task', 'withCheckedContinuation', 'withCheckedThrowingContinuation',
|
||||
'sink', 'store', 'assign', 'receive', 'subscribe',
|
||||
'addObserver', 'removeObserver', 'post', 'NotificationCenter',
|
||||
]);
|
||||
|
||||
export const swiftProvider = defineLanguage({
|
||||
id: SupportedLanguages.Swift,
|
||||
extensions: ['.swift'],
|
||||
|
|
@ -114,4 +142,5 @@ export const swiftProvider = defineLanguage({
|
|||
heritageDefaultEdge: 'IMPLEMENTS',
|
||||
fieldExtractor: createFieldExtractor(swiftFieldConfig),
|
||||
implicitImportWirer: wireSwiftImplicitImports,
|
||||
builtInNames: BUILT_INS,
|
||||
});
|
||||
|
|
|
|||
|
|
@ -18,6 +18,27 @@ import { typescriptFieldExtractor } from '../field-extractors/typescript.js';
|
|||
import { createFieldExtractor } from '../field-extractors/generic.js';
|
||||
import { javascriptConfig } from '../field-extractors/configs/typescript-javascript.js';
|
||||
|
||||
const BUILT_INS: ReadonlySet<string> = new Set([
|
||||
'console', 'log', 'warn', 'error', 'info', 'debug',
|
||||
'setTimeout', 'setInterval', 'clearTimeout', 'clearInterval',
|
||||
'parseInt', 'parseFloat', 'isNaN', 'isFinite',
|
||||
'encodeURI', 'decodeURI', 'encodeURIComponent', 'decodeURIComponent',
|
||||
'JSON', 'parse', 'stringify',
|
||||
'Object', 'Array', 'String', 'Number', 'Boolean', 'Symbol', 'BigInt',
|
||||
'Map', 'Set', 'WeakMap', 'WeakSet',
|
||||
'Promise', 'resolve', 'reject', 'then', 'catch', 'finally',
|
||||
'Math', 'Date', 'RegExp', 'Error',
|
||||
'require', 'import', 'export', 'fetch', 'Response', 'Request',
|
||||
'useState', 'useEffect', 'useCallback', 'useMemo', 'useRef', 'useContext',
|
||||
'useReducer', 'useLayoutEffect', 'useImperativeHandle', 'useDebugValue',
|
||||
'createElement', 'createContext', 'createRef', 'forwardRef', 'memo', 'lazy',
|
||||
'map', 'filter', 'reduce', 'forEach', 'find', 'findIndex', 'some', 'every',
|
||||
'includes', 'indexOf', 'slice', 'splice', 'concat', 'join', 'split',
|
||||
'push', 'pop', 'shift', 'unshift', 'sort', 'reverse',
|
||||
'keys', 'values', 'entries', 'assign', 'freeze', 'seal',
|
||||
'hasOwnProperty', 'toString', 'valueOf',
|
||||
]);
|
||||
|
||||
export const typescriptProvider = defineLanguage({
|
||||
id: SupportedLanguages.TypeScript,
|
||||
extensions: ['.ts', '.tsx'],
|
||||
|
|
@ -27,6 +48,7 @@ export const typescriptProvider = defineLanguage({
|
|||
importResolver: resolveTypescriptImport,
|
||||
namedBindingExtractor: extractTsNamedBindings,
|
||||
fieldExtractor: typescriptFieldExtractor,
|
||||
builtInNames: BUILT_INS,
|
||||
});
|
||||
|
||||
export const javascriptProvider = defineLanguage({
|
||||
|
|
@ -38,4 +60,5 @@ export const javascriptProvider = defineLanguage({
|
|||
importResolver: resolveJavascriptImport,
|
||||
namedBindingExtractor: extractTsNamedBindings,
|
||||
fieldExtractor: createFieldExtractor(javascriptConfig),
|
||||
builtInNames: BUILT_INS,
|
||||
});
|
||||
|
|
|
|||
|
|
@ -1,6 +1,7 @@
|
|||
import { createKnowledgeGraph } from '../graph/graph.js';
|
||||
import { processStructure } from './structure-processor.js';
|
||||
import { processMarkdown } from './markdown-processor.js';
|
||||
import { processCobol, isCobolFile, isJclFile } from './cobol-processor.js';
|
||||
import { processParsing } from './parsing-processor.js';
|
||||
import {
|
||||
processImports,
|
||||
|
|
@ -464,6 +465,14 @@ async function runScanAndStructure(
|
|||
stats: { filesProcessed: totalFiles, totalFiles, nodesCreated: graph.nodeCount },
|
||||
});
|
||||
|
||||
// ── Custom (non-tree-sitter) processors ─────────────────────────────
|
||||
// Each custom processor follows the pattern in markdown-processor.ts:
|
||||
// 1. Export a process function: (graph, files, allPathSet) => result
|
||||
// 2. Export a file detection function: (path) => boolean
|
||||
// 3. Filter files by extension, write nodes/edges directly to graph
|
||||
// To add a new language: create a new processor file, import it here,
|
||||
// and add a filter-read-call-log block following the pattern below.
|
||||
|
||||
// ── Phase 2.5: Markdown processing (headings + cross-links) ────────
|
||||
const mdScanned = scannedFiles.filter(f => f.path.endsWith('.md') || f.path.endsWith('.mdx'));
|
||||
if (mdScanned.length > 0) {
|
||||
|
|
@ -478,6 +487,26 @@ async function runScanAndStructure(
|
|||
}
|
||||
}
|
||||
|
||||
// ── Phase 2.6: COBOL processing (regex extraction, no tree-sitter) ──
|
||||
const cobolScanned = scannedFiles.filter(f => isCobolFile(f.path) || isJclFile(f.path));
|
||||
if (cobolScanned.length > 0) {
|
||||
const cobolContents = await readFileContents(repoPath, cobolScanned.map(f => f.path));
|
||||
const cobolFiles = cobolScanned
|
||||
.filter(f => cobolContents.has(f.path))
|
||||
.map(f => ({ path: f.path, content: cobolContents.get(f.path)! }));
|
||||
const allPathSet = new Set(allPaths);
|
||||
const cobolResult = processCobol(graph, cobolFiles, allPathSet);
|
||||
if (isDev) {
|
||||
console.log(` COBOL: ${cobolResult.programs} programs, ${cobolResult.paragraphs} paragraphs, ${cobolResult.sections} sections from ${cobolFiles.length} files`);
|
||||
if (cobolResult.execSqlBlocks > 0 || cobolResult.execCicsBlocks > 0 || cobolResult.entryPoints > 0) {
|
||||
console.log(` COBOL enriched: ${cobolResult.execSqlBlocks} SQL blocks, ${cobolResult.execCicsBlocks} CICS blocks, ${cobolResult.entryPoints} entry points, ${cobolResult.moves} moves, ${cobolResult.fileDeclarations} file declarations`);
|
||||
}
|
||||
if (cobolResult.jclJobs > 0) {
|
||||
console.log(` JCL: ${cobolResult.jclJobs} jobs, ${cobolResult.jclSteps} steps`);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
return { scannedFiles, allPaths, totalFiles };
|
||||
}
|
||||
|
||||
|
|
|
|||
|
|
@ -808,7 +808,7 @@ export const RUBY_QUERIES = `
|
|||
; NOTE: This may over-capture variable reads as calls (e.g. 'result' at
|
||||
; statement level). Ruby's grammar makes bare identifiers ambiguous — they
|
||||
; could be local variables or zero-arity method calls. Post-processing via
|
||||
; isBuiltInOrNoise and symbol resolution filtering suppresses most false
|
||||
; provider.isBuiltInName and symbol resolution filtering suppresses most false
|
||||
; positives, but a variable name that coincidentally matches a method name
|
||||
; elsewhere may produce a false CALLS edge.
|
||||
(body_statement
|
||||
|
|
@ -1054,6 +1054,20 @@ export const DART_QUERIES = `
|
|||
(factory_constructor_signature
|
||||
(identifier) @name . (formal_parameter_list))) @definition.constructor
|
||||
|
||||
; ── Field declarations (String name = '', Address address = Address()) ──────
|
||||
(declaration
|
||||
(type_identifier)
|
||||
(initialized_identifier_list
|
||||
(initialized_identifier
|
||||
(identifier) @name))) @definition.property
|
||||
|
||||
; ── Nullable field declarations (String? name) ──────────────────────────────
|
||||
(declaration
|
||||
(nullable_type)
|
||||
(initialized_identifier_list
|
||||
(initialized_identifier
|
||||
(identifier) @name))) @definition.property
|
||||
|
||||
; ── Getters ──────────────────────────────────────────────────────────────────
|
||||
(method_signature
|
||||
(getter_signature
|
||||
|
|
@ -1097,6 +1111,22 @@ export const DART_QUERIES = `
|
|||
(library_export
|
||||
(configurable_uri) @import.source)) @import
|
||||
|
||||
; ── Write access: obj.field = value ──────────────────────────────────────────
|
||||
(assignment_expression
|
||||
left: (assignable_expression
|
||||
(identifier) @assignment.receiver
|
||||
(unconditional_assignable_selector
|
||||
(identifier) @assignment.property))
|
||||
right: (_)) @assignment
|
||||
|
||||
; ── Write access: this.field = value ─────────────────────────────────────────
|
||||
(assignment_expression
|
||||
left: (assignable_expression
|
||||
(this) @assignment.receiver
|
||||
(unconditional_assignable_selector
|
||||
(identifier) @assignment.property))
|
||||
right: (_)) @assignment
|
||||
|
||||
; ── Heritage: extends ────────────────────────────────────────────────────────
|
||||
(class_definition
|
||||
name: (identifier) @heritage.class
|
||||
|
|
@ -1134,4 +1164,5 @@ export const LANGUAGE_QUERIES: Record<SupportedLanguages, string> = {
|
|||
[SupportedLanguages.Ruby]: RUBY_QUERIES,
|
||||
[SupportedLanguages.Swift]: SWIFT_QUERIES,
|
||||
[SupportedLanguages.Dart]: DART_QUERIES,
|
||||
[SupportedLanguages.Cobol]: '', // Standalone regex processor — no tree-sitter queries
|
||||
};
|
||||
|
|
|
|||
|
|
@ -1,6 +1,5 @@
|
|||
import { type SyntaxNode, FUNCTION_NODE_TYPES, extractFunctionName, CLASS_CONTAINER_TYPES } from './utils/ast-helpers.js';
|
||||
import { CALL_EXPRESSION_TYPES } from './utils/call-analysis.js';
|
||||
import { isBuiltInOrNoise } from './utils/noise-filter.js';
|
||||
import { SupportedLanguages } from '../../config/supported-languages.js';
|
||||
import { TYPED_PARAMETER_TYPES } from './type-extractors/shared.js';
|
||||
import { getProvider } from './languages/index.js';
|
||||
|
|
@ -732,7 +731,7 @@ export const buildTypeEnv = (
|
|||
lookupReturnType(callee: string): string | undefined {
|
||||
// SymbolTable is authoritative when it has an unambiguous match
|
||||
if (symbolTable) {
|
||||
if (isBuiltInOrNoise(callee)) return undefined;
|
||||
if (provider.isBuiltInName(callee)) return undefined;
|
||||
const callables = symbolTable.lookupFuzzyCallable(callee);
|
||||
if (callables.length === 1) {
|
||||
const rawReturn = callables[0].returnType;
|
||||
|
|
@ -746,7 +745,7 @@ export const buildTypeEnv = (
|
|||
},
|
||||
lookupRawReturnType(callee: string): string | undefined {
|
||||
if (symbolTable) {
|
||||
if (isBuiltInOrNoise(callee)) return undefined;
|
||||
if (provider.isBuiltInName(callee)) return undefined;
|
||||
const callables = symbolTable.lookupFuzzyCallable(callee);
|
||||
if (callables.length === 1) return callables[0].returnType;
|
||||
// Ambiguous (2+) → return undefined (conservative, no cross-file fallback)
|
||||
|
|
|
|||
|
|
@ -111,6 +111,25 @@ function hasDartTypeAnnotation(node: SyntaxNode): boolean {
|
|||
// ── Tier 0: Explicit Type Annotations ───────────────────────────────────
|
||||
|
||||
const extractDartDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
|
||||
// initialized_identifier: comma-separated variable (String a, b, c) — type is on parent
|
||||
if (node.type === 'initialized_identifier') {
|
||||
const parent = node.parent;
|
||||
if (!parent) return;
|
||||
let typeNode = findChild(parent, 'type_identifier');
|
||||
if (!typeNode) {
|
||||
const nullable = findChild(parent, 'nullable_type');
|
||||
if (nullable) typeNode = findChild(nullable, 'type_identifier');
|
||||
}
|
||||
if (!typeNode) return;
|
||||
const typeName = extractSimpleTypeName(typeNode);
|
||||
if (!typeName || typeName === 'dynamic') return;
|
||||
const nameNode = findChild(node, 'identifier');
|
||||
if (!nameNode) return;
|
||||
const varName = extractVarName(nameNode);
|
||||
if (varName) env.set(varName, typeName);
|
||||
return;
|
||||
}
|
||||
|
||||
let typeNode = findChild(node, 'type_identifier');
|
||||
if (!typeNode) {
|
||||
const nullable = findChild(node, 'nullable_type');
|
||||
|
|
|
|||
|
|
@ -82,6 +82,9 @@ export const FUNCTION_NODE_TYPES = new Set([
|
|||
// Ruby
|
||||
'method', // def foo
|
||||
'singleton_method', // def self.foo
|
||||
// Dart
|
||||
'function_signature',
|
||||
'method_signature',
|
||||
]);
|
||||
|
||||
/**
|
||||
|
|
@ -430,6 +433,34 @@ export const extractFunctionName = (node: SyntaxNode): { funcName: string | null
|
|||
}
|
||||
funcName = nameNode?.text;
|
||||
label = 'Method';
|
||||
} else if (node.type === 'function_signature') {
|
||||
// Dart: top-level function signatures
|
||||
let nameNode = node.childForFieldName?.('name');
|
||||
if (!nameNode) {
|
||||
for (let i = 0; i < node.childCount; i++) {
|
||||
const c = node.child(i);
|
||||
if (c?.type === 'identifier') { nameNode = c; break; }
|
||||
}
|
||||
}
|
||||
funcName = nameNode?.text ?? null;
|
||||
} else if (node.type === 'method_signature') {
|
||||
// Dart: method_signature wraps function_signature
|
||||
let funcSig: SyntaxNode | null = null;
|
||||
for (let i = 0; i < node.childCount; i++) {
|
||||
const c = node.child(i);
|
||||
if (c?.type === 'function_signature') { funcSig = c; break; }
|
||||
}
|
||||
if (funcSig) {
|
||||
let nameNode = funcSig.childForFieldName?.('name');
|
||||
if (!nameNode) {
|
||||
for (let i = 0; i < funcSig.childCount; i++) {
|
||||
const c = funcSig.child(i);
|
||||
if (c?.type === 'identifier') { nameNode = c; break; }
|
||||
}
|
||||
}
|
||||
funcName = nameNode?.text ?? null;
|
||||
}
|
||||
label = 'Method';
|
||||
}
|
||||
|
||||
return { funcName, label };
|
||||
|
|
@ -471,6 +502,7 @@ export const extractMethodSignature = (node: SyntaxNode | null | undefined): Met
|
|||
const paramListTypes = new Set([
|
||||
'formal_parameters', 'parameters', 'parameter_list',
|
||||
'function_parameters', 'method_parameters', 'function_value_parameters',
|
||||
'formal_parameter_list', // Dart
|
||||
]);
|
||||
|
||||
// Node types that indicate variadic/rest parameters
|
||||
|
|
|
|||
|
|
@ -64,6 +64,7 @@ const MEMBER_ACCESS_NODE_TYPES = new Set([
|
|||
'selector_expression', // Go: obj.Method()
|
||||
'navigation_suffix', // Kotlin/Swift: obj.method() — nameNode sits inside navigation_suffix
|
||||
'member_binding_expression', // C#: user?.Method() — null-conditional access
|
||||
'unconditional_assignable_selector', // Dart: obj.method() — nameNode inside selector > unconditional_assignable_selector
|
||||
]);
|
||||
|
||||
/**
|
||||
|
|
@ -208,6 +209,16 @@ export const extractReceiverName = (
|
|||
}
|
||||
}
|
||||
|
||||
// Dart: unconditional_assignable_selector is inside a `selector`, which is a sibling
|
||||
// of the receiver in the expression_statement. For `user.save()`, the previous named
|
||||
// sibling of the `selector` is `identifier [user]`.
|
||||
if (!receiver && parent.type === 'unconditional_assignable_selector') {
|
||||
const selectorNode = parent.parent; // selector [.save]
|
||||
if (selectorNode) {
|
||||
receiver = selectorNode.previousNamedSibling;
|
||||
}
|
||||
}
|
||||
|
||||
// C# null-conditional: user?.Save() → conditional_access_expression wraps member_binding_expression
|
||||
if (!receiver && parent.type === 'member_binding_expression') {
|
||||
const condAccess = parent.parent;
|
||||
|
|
@ -291,6 +302,14 @@ export const extractReceiverNode = (
|
|||
}
|
||||
}
|
||||
|
||||
// Dart: unconditional_assignable_selector — receiver is previous sibling of the selector
|
||||
if (!receiver && parent.type === 'unconditional_assignable_selector') {
|
||||
const selectorNode = parent.parent;
|
||||
if (selectorNode) {
|
||||
receiver = selectorNode.previousNamedSibling;
|
||||
}
|
||||
}
|
||||
|
||||
if (!receiver && parent.type === 'member_binding_expression') {
|
||||
const condAccess = parent.parent;
|
||||
if (condAccess?.type === 'conditional_access_expression') {
|
||||
|
|
@ -527,6 +546,27 @@ export function extractMixedChain(
|
|||
} else {
|
||||
return { chain, baseReceiverName: innerObject.text || undefined };
|
||||
}
|
||||
} else if (current.type === 'selector') {
|
||||
// ── Dart: flat selector siblings (user.address.save() uses selector nodes) ──
|
||||
// Extract field name from unconditional_assignable_selector child
|
||||
const uas = current.namedChildren?.find(
|
||||
(c: SyntaxNode) => c.type === 'unconditional_assignable_selector',
|
||||
);
|
||||
const propertyName = uas?.namedChildren?.find(
|
||||
(c: SyntaxNode) => c.type === 'identifier',
|
||||
)?.text;
|
||||
if (!propertyName) break;
|
||||
chain.unshift({ kind: 'field', name: propertyName });
|
||||
|
||||
// Walk to previous sibling for the next step in the chain
|
||||
const prev = current.previousNamedSibling;
|
||||
if (!prev) break;
|
||||
if (prev.type === 'selector') {
|
||||
current = prev;
|
||||
} else {
|
||||
// Base receiver (identifier or other terminal)
|
||||
return { chain, baseReceiverName: prev.text || undefined };
|
||||
}
|
||||
} else {
|
||||
// Simple identifier — this is the base receiver
|
||||
return chain.length > 0
|
||||
|
|
|
|||
|
|
@ -1,175 +0,0 @@
|
|||
/**
|
||||
* Built-in name filtering — identifies standard library functions and common noise
|
||||
* that should not be tracked as call targets in the knowledge graph.
|
||||
*
|
||||
* Covers: JS/TS, Python, Kotlin, C/C++, C#, PHP, Swift, Rust, Ruby, Dart/Flutter standard libraries.
|
||||
*/
|
||||
|
||||
export const BUILT_IN_NAMES = new Set([
|
||||
// JavaScript/TypeScript
|
||||
'console', 'log', 'warn', 'error', 'info', 'debug',
|
||||
'setTimeout', 'setInterval', 'clearTimeout', 'clearInterval',
|
||||
'parseInt', 'parseFloat', 'isNaN', 'isFinite',
|
||||
'encodeURI', 'decodeURI', 'encodeURIComponent', 'decodeURIComponent',
|
||||
'JSON', 'parse', 'stringify',
|
||||
'Object', 'Array', 'String', 'Number', 'Boolean', 'Symbol', 'BigInt',
|
||||
'Map', 'Set', 'WeakMap', 'WeakSet',
|
||||
'Promise', 'resolve', 'reject', 'then', 'catch', 'finally',
|
||||
'Math', 'Date', 'RegExp', 'Error',
|
||||
'require', 'import', 'export', 'fetch', 'Response', 'Request',
|
||||
'useState', 'useEffect', 'useCallback', 'useMemo', 'useRef', 'useContext',
|
||||
'useReducer', 'useLayoutEffect', 'useImperativeHandle', 'useDebugValue',
|
||||
'createElement', 'createContext', 'createRef', 'forwardRef', 'memo', 'lazy',
|
||||
'map', 'filter', 'reduce', 'forEach', 'find', 'findIndex', 'some', 'every',
|
||||
'includes', 'indexOf', 'slice', 'splice', 'concat', 'join', 'split',
|
||||
'push', 'pop', 'shift', 'unshift', 'sort', 'reverse',
|
||||
'keys', 'values', 'entries', 'assign', 'freeze', 'seal',
|
||||
'hasOwnProperty', 'toString', 'valueOf',
|
||||
// Python
|
||||
'print', 'len', 'range', 'str', 'int', 'float', 'list', 'dict', 'set', 'tuple',
|
||||
'append', 'extend', 'update',
|
||||
// NOTE: 'open', 'read', 'write', 'close' removed — these are real C POSIX syscalls
|
||||
'type', 'isinstance', 'issubclass', 'getattr', 'setattr', 'hasattr',
|
||||
'enumerate', 'zip', 'sorted', 'reversed', 'min', 'max', 'sum', 'abs',
|
||||
// Kotlin stdlib
|
||||
'println', 'print', 'readLine', 'require', 'requireNotNull', 'check', 'assert', 'lazy', 'error',
|
||||
'listOf', 'mapOf', 'setOf', 'mutableListOf', 'mutableMapOf', 'mutableSetOf',
|
||||
'arrayOf', 'sequenceOf', 'also', 'apply', 'run', 'with', 'takeIf', 'takeUnless',
|
||||
'TODO', 'buildString', 'buildList', 'buildMap', 'buildSet',
|
||||
'repeat', 'synchronized',
|
||||
// Kotlin coroutine builders & scope functions
|
||||
'launch', 'async', 'runBlocking', 'withContext', 'coroutineScope',
|
||||
'supervisorScope', 'delay',
|
||||
// Kotlin Flow operators
|
||||
'flow', 'flowOf', 'collect', 'emit', 'onEach', 'catch',
|
||||
'buffer', 'conflate', 'distinctUntilChanged',
|
||||
'flatMapLatest', 'flatMapMerge', 'combine',
|
||||
'stateIn', 'shareIn', 'launchIn',
|
||||
// Kotlin infix stdlib functions
|
||||
'to', 'until', 'downTo', 'step',
|
||||
// C/C++ standard library
|
||||
'printf', 'fprintf', 'sprintf', 'snprintf', 'vprintf', 'vfprintf', 'vsprintf', 'vsnprintf',
|
||||
'scanf', 'fscanf', 'sscanf',
|
||||
'malloc', 'calloc', 'realloc', 'free', 'memcpy', 'memmove', 'memset', 'memcmp',
|
||||
'strlen', 'strcpy', 'strncpy', 'strcat', 'strncat', 'strcmp', 'strncmp', 'strstr', 'strchr', 'strrchr',
|
||||
'atoi', 'atol', 'atof', 'strtol', 'strtoul', 'strtoll', 'strtoull', 'strtod',
|
||||
'sizeof', 'offsetof', 'typeof',
|
||||
'assert', 'abort', 'exit', '_exit',
|
||||
'fopen', 'fclose', 'fread', 'fwrite', 'fseek', 'ftell', 'rewind', 'fflush', 'fgets', 'fputs',
|
||||
// Linux kernel common macros/helpers (not real call targets)
|
||||
'likely', 'unlikely', 'BUG', 'BUG_ON', 'WARN', 'WARN_ON', 'WARN_ONCE',
|
||||
'IS_ERR', 'PTR_ERR', 'ERR_PTR', 'IS_ERR_OR_NULL',
|
||||
'ARRAY_SIZE', 'container_of', 'list_for_each_entry', 'list_for_each_entry_safe',
|
||||
'min', 'max', 'clamp', 'abs', 'swap',
|
||||
'pr_info', 'pr_warn', 'pr_err', 'pr_debug', 'pr_notice', 'pr_crit', 'pr_emerg',
|
||||
'printk', 'dev_info', 'dev_warn', 'dev_err', 'dev_dbg',
|
||||
'GFP_KERNEL', 'GFP_ATOMIC',
|
||||
'spin_lock', 'spin_unlock', 'spin_lock_irqsave', 'spin_unlock_irqrestore',
|
||||
'mutex_lock', 'mutex_unlock', 'mutex_init',
|
||||
'kfree', 'kmalloc', 'kzalloc', 'kcalloc', 'krealloc', 'kvmalloc', 'kvfree',
|
||||
'get', 'put',
|
||||
// C# / .NET built-ins
|
||||
'Console', 'WriteLine', 'ReadLine', 'Write',
|
||||
'Task', 'Run', 'Wait', 'WhenAll', 'WhenAny', 'FromResult', 'Delay', 'ContinueWith',
|
||||
'ConfigureAwait', 'GetAwaiter', 'GetResult',
|
||||
'ToString', 'GetType', 'Equals', 'GetHashCode', 'ReferenceEquals',
|
||||
'Add', 'Remove', 'Contains', 'Clear', 'Count', 'Any', 'All',
|
||||
'Where', 'Select', 'SelectMany', 'OrderBy', 'OrderByDescending', 'GroupBy',
|
||||
'First', 'FirstOrDefault', 'Single', 'SingleOrDefault', 'Last', 'LastOrDefault',
|
||||
'ToList', 'ToArray', 'ToDictionary', 'AsEnumerable', 'AsQueryable',
|
||||
'Aggregate', 'Sum', 'Average', 'Min', 'Max', 'Distinct', 'Skip', 'Take',
|
||||
'String', 'Format', 'IsNullOrEmpty', 'IsNullOrWhiteSpace', 'Concat', 'Join',
|
||||
'Trim', 'TrimStart', 'TrimEnd', 'Split', 'Replace', 'StartsWith', 'EndsWith',
|
||||
'Convert', 'ToInt32', 'ToDouble', 'ToBoolean', 'ToByte',
|
||||
'Math', 'Abs', 'Ceiling', 'Floor', 'Round', 'Pow', 'Sqrt',
|
||||
'Dispose', 'Close',
|
||||
'TryParse', 'Parse',
|
||||
'AddRange', 'RemoveAt', 'RemoveAll', 'FindAll', 'Exists', 'TrueForAll',
|
||||
'ContainsKey', 'TryGetValue', 'AddOrUpdate',
|
||||
'Throw', 'ThrowIfNull',
|
||||
// PHP built-ins
|
||||
'echo', 'isset', 'empty', 'unset', 'list', 'array', 'compact', 'extract',
|
||||
'count', 'strlen', 'strpos', 'strrpos', 'substr', 'strtolower', 'strtoupper', 'trim',
|
||||
'ltrim', 'rtrim', 'str_replace', 'str_contains', 'str_starts_with', 'str_ends_with',
|
||||
'sprintf', 'vsprintf', 'printf', 'number_format',
|
||||
'array_map', 'array_filter', 'array_reduce', 'array_push', 'array_pop', 'array_shift',
|
||||
'array_unshift', 'array_slice', 'array_splice', 'array_merge', 'array_keys', 'array_values',
|
||||
'array_key_exists', 'in_array', 'array_search', 'array_unique', 'usort', 'rsort',
|
||||
'json_encode', 'json_decode', 'serialize', 'unserialize',
|
||||
'intval', 'floatval', 'strval', 'boolval', 'is_null', 'is_string', 'is_int', 'is_array',
|
||||
'is_object', 'is_numeric', 'is_bool', 'is_float',
|
||||
'var_dump', 'print_r', 'var_export',
|
||||
'date', 'time', 'strtotime', 'mktime', 'microtime',
|
||||
'file_exists', 'file_get_contents', 'file_put_contents', 'is_file', 'is_dir',
|
||||
'preg_match', 'preg_match_all', 'preg_replace', 'preg_split',
|
||||
'header', 'session_start', 'session_destroy', 'ob_start', 'ob_end_clean', 'ob_get_clean',
|
||||
'dd', 'dump',
|
||||
// Swift/iOS built-ins and standard library
|
||||
'print', 'debugPrint', 'dump', 'fatalError', 'precondition', 'preconditionFailure',
|
||||
'assert', 'assertionFailure', 'NSLog',
|
||||
'abs', 'min', 'max', 'zip', 'stride', 'sequence', 'repeatElement',
|
||||
'swap', 'withUnsafePointer', 'withUnsafeMutablePointer', 'withUnsafeBytes',
|
||||
'autoreleasepool', 'unsafeBitCast', 'unsafeDowncast', 'numericCast',
|
||||
'type', 'MemoryLayout',
|
||||
// Swift collection/string methods (common noise)
|
||||
'map', 'flatMap', 'compactMap', 'filter', 'reduce', 'forEach', 'contains',
|
||||
'first', 'last', 'prefix', 'suffix', 'dropFirst', 'dropLast',
|
||||
'sorted', 'reversed', 'enumerated', 'joined', 'split',
|
||||
'append', 'insert', 'remove', 'removeAll', 'removeFirst', 'removeLast',
|
||||
'isEmpty', 'count', 'index', 'startIndex', 'endIndex',
|
||||
// UIKit/Foundation common methods (noise in call graph)
|
||||
'addSubview', 'removeFromSuperview', 'layoutSubviews', 'setNeedsLayout',
|
||||
'layoutIfNeeded', 'setNeedsDisplay', 'invalidateIntrinsicContentSize',
|
||||
'addTarget', 'removeTarget', 'addGestureRecognizer',
|
||||
'addConstraint', 'addConstraints', 'removeConstraint', 'removeConstraints',
|
||||
'NSLocalizedString', 'Bundle',
|
||||
'reloadData', 'reloadSections', 'reloadRows', 'performBatchUpdates',
|
||||
'register', 'dequeueReusableCell', 'dequeueReusableSupplementaryView',
|
||||
'beginUpdates', 'endUpdates', 'insertRows', 'deleteRows', 'insertSections', 'deleteSections',
|
||||
'present', 'dismiss', 'pushViewController', 'popViewController', 'popToRootViewController',
|
||||
'performSegue', 'prepare',
|
||||
// GCD / async
|
||||
'DispatchQueue', 'async', 'sync', 'asyncAfter',
|
||||
'Task', 'withCheckedContinuation', 'withCheckedThrowingContinuation',
|
||||
// Combine
|
||||
'sink', 'store', 'assign', 'receive', 'subscribe',
|
||||
// Notification / KVO
|
||||
'addObserver', 'removeObserver', 'post', 'NotificationCenter',
|
||||
// Rust standard library (common noise in call graphs)
|
||||
'unwrap', 'expect', 'unwrap_or', 'unwrap_or_else', 'unwrap_or_default',
|
||||
'ok', 'err', 'is_ok', 'is_err', 'map', 'map_err', 'and_then', 'or_else',
|
||||
'clone', 'to_string', 'to_owned', 'into', 'from', 'as_ref', 'as_mut',
|
||||
'iter', 'into_iter', 'collect', 'map', 'filter', 'fold', 'for_each',
|
||||
'len', 'is_empty', 'push', 'pop', 'insert', 'remove', 'contains',
|
||||
'format', 'write', 'writeln', 'panic', 'unreachable', 'todo', 'unimplemented',
|
||||
'vec', 'println', 'eprintln', 'dbg',
|
||||
'lock', 'read', 'write', 'try_lock',
|
||||
'spawn', 'join', 'sleep',
|
||||
'Some', 'None', 'Ok', 'Err',
|
||||
// Ruby built-ins and Kernel methods
|
||||
'puts', 'p', 'pp', 'raise', 'fail',
|
||||
'require', 'require_relative', 'load', 'autoload',
|
||||
'include', 'extend', 'prepend',
|
||||
'attr_accessor', 'attr_reader', 'attr_writer',
|
||||
'public', 'private', 'protected', 'module_function',
|
||||
'lambda', 'proc', 'block_given?',
|
||||
'nil?', 'is_a?', 'kind_of?', 'instance_of?', 'respond_to?',
|
||||
'freeze', 'frozen?', 'dup', 'tap', 'yield_self',
|
||||
// Dart / Flutter
|
||||
'setState', 'mounted', 'debugPrint',
|
||||
'runApp', 'showDialog', 'showModalBottomSheet',
|
||||
'Navigator', 'push', 'pushNamed', 'pushReplacement', 'pop', 'maybePop',
|
||||
'ScaffoldMessenger', 'showSnackBar',
|
||||
'deactivate', 'reassemble', 'debugDumpApp', 'debugDumpRenderTree',
|
||||
// Dart async
|
||||
'then', 'catchError', 'whenComplete', 'listen',
|
||||
// Ruby enumerables
|
||||
'each', 'select', 'reject', 'detect', 'collect',
|
||||
'inject', 'flat_map', 'each_with_object', 'each_with_index',
|
||||
'any?', 'all?', 'none?', 'count', 'first', 'last',
|
||||
'sort_by', 'min_by', 'max_by',
|
||||
'group_by', 'partition', 'compact', 'flatten', 'uniq',
|
||||
]);
|
||||
|
||||
/** Check if a name is a built-in function or common noise that should be filtered out */
|
||||
export const isBuiltInOrNoise = (name: string): boolean => BUILT_IN_NAMES.has(name);
|
||||
|
|
@ -29,7 +29,6 @@ try { Dart = _require('tree-sitter-dart'); } catch {}
|
|||
let Kotlin: any = null;
|
||||
try { Kotlin = _require('tree-sitter-kotlin'); } catch {}
|
||||
import { getLanguageFromFilename } from '../utils/language-detection.js';
|
||||
import { isBuiltInOrNoise } from '../utils/noise-filter.js';
|
||||
import {
|
||||
FUNCTION_NODE_TYPES,
|
||||
extractFunctionName,
|
||||
|
|
@ -388,6 +387,23 @@ const findEnclosingFunctionId = (node: any, filePath: string, provider: Language
|
|||
return result;
|
||||
}
|
||||
}
|
||||
|
||||
// Language-specific enclosing function resolution (e.g., Dart where
|
||||
// function_body is a sibling of function_signature, not a child).
|
||||
if (provider.enclosingFunctionFinder) {
|
||||
const customResult = provider.enclosingFunctionFinder(current);
|
||||
if (customResult) {
|
||||
let finalLabel: NodeLabel = customResult.label;
|
||||
if (provider.labelOverride) {
|
||||
const override = provider.labelOverride(current.previousSibling, finalLabel);
|
||||
if (override !== null) finalLabel = override;
|
||||
}
|
||||
const result = generateId(finalLabel, `${filePath}:${customResult.funcName}`);
|
||||
functionIdCache.set(node, result);
|
||||
return result;
|
||||
}
|
||||
}
|
||||
|
||||
current = current.parent;
|
||||
}
|
||||
functionIdCache.set(node, null);
|
||||
|
|
@ -1268,7 +1284,7 @@ const processFileGroup = (
|
|||
// kind === 'call' — fall through to normal call processing below
|
||||
}
|
||||
|
||||
if (!isBuiltInOrNoise(calledName)) {
|
||||
if (!provider.isBuiltInName(calledName)) {
|
||||
const callNode = captureMap['call'];
|
||||
const sourceId = findEnclosingFunctionId(callNode, file.path, provider)
|
||||
|| generateId('File', file.path);
|
||||
|
|
|
|||
|
|
@ -551,9 +551,10 @@ export const closeLbug = async (repoId?: string): Promise<void> => {
|
|||
export const isLbugReady = (repoId: string): boolean => pool.has(repoId);
|
||||
|
||||
/** Regex to detect write operations in user-supplied Cypher queries */
|
||||
export const CYPHER_WRITE_RE = /\b(CREATE|DELETE|SET|MERGE|REMOVE|DROP|ALTER|COPY|DETACH)\b/i;
|
||||
export const CYPHER_WRITE_RE = /(?<!:)\b(CREATE|DELETE|SET|MERGE|REMOVE|DROP|ALTER|COPY|DETACH|FOREACH)\b/i;
|
||||
|
||||
/** Check if a Cypher query contains write operations */
|
||||
export function isWriteQuery(query: string): boolean {
|
||||
return CYPHER_WRITE_RE.test(query);
|
||||
}
|
||||
|
||||
|
|
|
|||
|
|
@ -8,7 +8,8 @@
|
|||
|
||||
import fs from 'fs/promises';
|
||||
import path from 'path';
|
||||
import { initLbug, executeQuery, executeParameterized, closeLbug, isLbugReady } from '../core/lbug-adapter.js';
|
||||
import { initLbug, executeQuery, executeParameterized, closeLbug, isLbugReady, isWriteQuery } from '../core/lbug-adapter.js';
|
||||
export { isWriteQuery };
|
||||
// Embedding imports are lazy (dynamic import) to avoid loading onnxruntime-node
|
||||
// at MCP server startup — crashes on unsupported Node ABI versions (#89)
|
||||
// git utilities available if needed
|
||||
|
|
@ -89,13 +90,6 @@ export const IMPACT_RELATION_CONFIDENCE: Readonly<Record<string, number>> = {
|
|||
const confidenceForRelType = (relType: string | undefined): number =>
|
||||
IMPACT_RELATION_CONFIDENCE[relType ?? ''] ?? 0.5;
|
||||
|
||||
/** Regex to detect write operations in user-supplied Cypher queries */
|
||||
export const CYPHER_WRITE_RE = /\b(CREATE|DELETE|SET|MERGE|REMOVE|DROP|ALTER|COPY|DETACH)\b/i;
|
||||
|
||||
/** Check if a Cypher query contains write operations */
|
||||
export function isWriteQuery(query: string): boolean {
|
||||
return CYPHER_WRITE_RE.test(query);
|
||||
}
|
||||
|
||||
/** Structured error logging for query failures — replaces empty catch blocks */
|
||||
function logQueryError(context: string, err: unknown): void {
|
||||
|
|
@ -777,7 +771,7 @@ export class LocalBackend {
|
|||
}
|
||||
|
||||
// Block write operations (defense-in-depth — DB is already read-only)
|
||||
if (CYPHER_WRITE_RE.test(params.query)) {
|
||||
if (isWriteQuery(params.query)) {
|
||||
return { error: 'Write operations (CREATE, DELETE, SET, MERGE, REMOVE, DROP, ALTER, COPY, DETACH) are not allowed. The knowledge graph is read-only.' };
|
||||
}
|
||||
|
||||
|
|
@ -1722,69 +1716,219 @@ export class LocalBackend {
|
|||
let affectedModules: any[] = [];
|
||||
|
||||
if (impacted.length > 0) {
|
||||
// Cap IN-clause to 100 IDs to prevent oversized queries that crash
|
||||
// the native DB engine on arm64 macOS (#292)
|
||||
const cappedImpacted = impacted.slice(0, 100);
|
||||
const allIds = cappedImpacted.map(i => `'${String(i.id ?? '').replace(/'/g, "''")}'`).join(', ');
|
||||
const d1Items = (grouped[1] || []).slice(0, 100);
|
||||
const d1Ids = d1Items.map((i: any) => `'${String(i.id ?? '').replace(/'/g, "''")}'`).join(', ');
|
||||
const CHUNK_SIZE = 100;
|
||||
// Max number of chunks to process to avoid unbounded DB round-trips.
|
||||
// Configurable via env IMPACT_MAX_CHUNKS, default 10 => max items = 1000
|
||||
const MAX_CHUNKS = parseInt(process.env.IMPACT_MAX_CHUNKS || '10', 10);
|
||||
|
||||
// Enrichment queries: sequential on arm64 macOS to avoid SIGSEGV from
|
||||
// concurrent native DB access (#285, #290, #292); parallel elsewhere
|
||||
// to preserve performance on unaffected platforms.
|
||||
const isArm64Mac = process.platform === 'darwin' && process.arch === 'arm64';
|
||||
// ── Process enrichment: batched chunking (bounded by MAX_CHUNKS) ─
|
||||
// Uses merged Cypher query (WITH + OPTIONAL MATCH) to fetch
|
||||
// process + entry point info in 1 round-trip per chunk. Converted to
|
||||
// parameterized queries to avoid manual string escaping and long query strings.
|
||||
const entryPointMap = new Map<string, {
|
||||
name: string; type: string; filePath: string;
|
||||
affected_process_count: number;
|
||||
total_hits: number;
|
||||
earliest_broken_step: number;
|
||||
}>();
|
||||
|
||||
const processQuery = executeQuery(repo.id, `
|
||||
MATCH (s)-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process)
|
||||
WHERE s.id IN [${allIds}]
|
||||
RETURN p.heuristicLabel AS name, COUNT(DISTINCT s.id) AS hits, MIN(r.step) AS minStep, p.stepCount AS stepCount
|
||||
ORDER BY hits DESC
|
||||
LIMIT 20
|
||||
`).catch(() => []);
|
||||
const moduleQuery = () => executeQuery(repo.id, `
|
||||
MATCH (s)-[:CodeRelation {type: 'MEMBER_OF'}]->(c:Community)
|
||||
WHERE s.id IN [${allIds}]
|
||||
RETURN c.heuristicLabel AS name, COUNT(DISTINCT s.id) AS hits
|
||||
ORDER BY hits DESC
|
||||
LIMIT 20
|
||||
`).catch(() => []);
|
||||
const directModuleQuery = () => d1Ids
|
||||
? executeQuery(repo.id, `
|
||||
MATCH (s)-[:CodeRelation {type: 'MEMBER_OF'}]->(c:Community)
|
||||
WHERE s.id IN [${d1Ids}]
|
||||
RETURN DISTINCT c.heuristicLabel AS name
|
||||
`).catch(() => [])
|
||||
: Promise.resolve([]);
|
||||
// Map process id -> entryPointId to allow fixing missing minStep values later
|
||||
const processToEntryPoint = new Map<string, string>();
|
||||
// Collect process ids where MIN(r.step) returned null so we can retry in batch
|
||||
const processesMissingMinStep = new Set<string>();
|
||||
|
||||
let processRows: any[], moduleRows: any[], directModuleRows: any[];
|
||||
if (isArm64Mac) {
|
||||
// Sequential: avoid concurrent native DB access
|
||||
processRows = await processQuery;
|
||||
moduleRows = await moduleQuery();
|
||||
directModuleRows = await directModuleQuery();
|
||||
} else {
|
||||
// Parallel: safe on non-arm64 platforms
|
||||
processRows = await processQuery;
|
||||
[moduleRows, directModuleRows] = await Promise.all([moduleQuery(), directModuleQuery()]);
|
||||
let chunksProcessed = 0;
|
||||
for (let i = 0; i < impacted.length && chunksProcessed < MAX_CHUNKS; i += CHUNK_SIZE, chunksProcessed++) {
|
||||
const chunk = impacted.slice(i, i + CHUNK_SIZE);
|
||||
const ids = chunk.map(item => String(item.id ?? ''));
|
||||
|
||||
try {
|
||||
// Use parameterized list to avoid building long query strings
|
||||
const rows = await executeParameterized(repo.id, `
|
||||
MATCH (s)-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process)
|
||||
WHERE s.id IN $ids
|
||||
WITH p, COUNT(DISTINCT s.id) AS hits, MIN(r.step) AS minStep
|
||||
OPTIONAL MATCH (ep {id: p.entryPointId})
|
||||
RETURN p.id AS pId, p.heuristicLabel AS name, p.processType AS processType,
|
||||
p.entryPointId AS entryPointId, hits, minStep, p.stepCount AS stepCount,
|
||||
ep.name AS epName, labels(ep)[0] AS epType, ep.filePath AS epFilePath
|
||||
`, { ids }).catch(() => []);
|
||||
|
||||
for (const row of rows) {
|
||||
const pId = row.pId ?? row[0];
|
||||
const epId = row.entryPointId ?? row[3] ?? row.pId ?? row[0];
|
||||
// Track mapping from process -> entryPoint so we can backfill missing minStep
|
||||
if (pId) processToEntryPoint.set(String(pId), String(epId));
|
||||
|
||||
// Normalize epName: prefer epName, fall back to other columns, and
|
||||
// ensure we don't keep an empty string (labels(...) can return "").
|
||||
const epNameRaw = row.epName ?? row[7] ?? row.name ?? row[1] ?? 'unknown';
|
||||
const epName = (typeof epNameRaw === 'string' && epNameRaw.trim().length > 0) ? epNameRaw.trim() : 'unknown';
|
||||
|
||||
// Normalize epType: labels(ep)[0] can return an empty string in
|
||||
// some DBs (LadybugDB). Using nullish coalescing (??) preserves
|
||||
// empty strings, which results in empty `type` values being
|
||||
// propagated. Treat empty-string labels as missing and fall back
|
||||
// to the next candidate or a sensible default.
|
||||
const epTypeRaw = row.epType ?? row[8] ?? '';
|
||||
const epType = (typeof epTypeRaw === 'string' && epTypeRaw.trim().length > 0)
|
||||
? epTypeRaw.trim()
|
||||
: 'Function';
|
||||
|
||||
const epFilePath = row.epFilePath ?? row[9] ?? '';
|
||||
const hits = row.hits ?? row[4] ?? 0;
|
||||
const minStep = row.minStep ?? row[5];
|
||||
// If the DB returned null for minStep, note the process id so we
|
||||
// can run a follow-up query using a different aggregation strategy.
|
||||
if (minStep === null || minStep === undefined) {
|
||||
if (pId) processesMissingMinStep.add(String(pId));
|
||||
}
|
||||
if (!entryPointMap.has(epId)) {
|
||||
entryPointMap.set(epId, {
|
||||
name: epName,
|
||||
type: epType,
|
||||
filePath: epFilePath,
|
||||
affected_process_count: 0,
|
||||
total_hits: 0,
|
||||
earliest_broken_step: Infinity,
|
||||
});
|
||||
}
|
||||
const ep = entryPointMap.get(epId)!;
|
||||
ep.affected_process_count += 1;
|
||||
ep.total_hits += hits;
|
||||
ep.earliest_broken_step = Math.min(ep.earliest_broken_step, minStep ?? Infinity);
|
||||
}
|
||||
} catch (e) {
|
||||
logQueryError('impact:process-chunk', e);
|
||||
}
|
||||
}
|
||||
|
||||
// If some processes returned null minStep, try a batched follow-up query
|
||||
// using the full impacted id set. This handles older indexes or DBs
|
||||
// where MIN(r.step) can come back null even when step properties exist.
|
||||
if (processesMissingMinStep.size > 0) {
|
||||
try {
|
||||
const pIds = Array.from(processesMissingMinStep);
|
||||
const allImpactedIds = impacted.map(it => String(it.id ?? ''));
|
||||
const missingRows = await executeParameterized(repo.id, `
|
||||
MATCH (s)-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process)
|
||||
WHERE p.id IN $pIds AND s.id IN $ids
|
||||
RETURN p.id AS pid, MIN(r.step) AS minStep
|
||||
`, { pIds, ids: allImpactedIds }).catch(() => []);
|
||||
|
||||
for (const mr of missingRows) {
|
||||
const pid = mr.pid ?? mr[0];
|
||||
const minStep = mr.minStep ?? mr[1];
|
||||
const epId = processToEntryPoint.get(String(pid));
|
||||
if (!epId) continue;
|
||||
const ep = entryPointMap.get(epId);
|
||||
if (!ep) continue;
|
||||
if (typeof minStep === 'number') {
|
||||
ep.earliest_broken_step = Math.min(ep.earliest_broken_step, minStep);
|
||||
}
|
||||
}
|
||||
} catch (e) {
|
||||
logQueryError('impact:process-chunk-backfill', e);
|
||||
}
|
||||
}
|
||||
|
||||
affectedProcesses = processRows.map((r: any) => ({
|
||||
name: r.name || r[0],
|
||||
hits: r.hits || r[1],
|
||||
broken_at_step: r.minStep ?? r[2],
|
||||
step_count: r.stepCount ?? r[3],
|
||||
}));
|
||||
// If we capped chunks, mark traversal incomplete so caller knows results are partial
|
||||
if (chunksProcessed * CHUNK_SIZE < impacted.length) {
|
||||
traversalComplete = false;
|
||||
}
|
||||
|
||||
const directModuleSet = new Set(directModuleRows.map((r: any) => r.name || r[0]));
|
||||
affectedProcesses = Array.from(entryPointMap.values())
|
||||
.map(ep => ({
|
||||
...ep,
|
||||
earliest_broken_step: ep.earliest_broken_step === Infinity ? null : ep.earliest_broken_step,
|
||||
}))
|
||||
.sort((a, b) => b.total_hits - a.total_hits);
|
||||
|
||||
// ── Module enrichment: use same cap as process enrichment and parameterized queries
|
||||
const maxItems = Math.min(impacted.length, MAX_CHUNKS * CHUNK_SIZE);
|
||||
const cappedImpacted = impacted.slice(0, maxItems);
|
||||
const allIdsArr = cappedImpacted.map((i: any) => String(i.id ?? ''));
|
||||
const d1Items = (grouped[1] || []).slice(0, maxItems);
|
||||
const d1IdsArr = d1Items.map((i: any) => String(i.id ?? ''));
|
||||
|
||||
// Chunked module enrichment: run the MEMBER_OF queries in chunks
|
||||
// to avoid large single queries or concurrent Kuzu calls that can
|
||||
// crash (SIGSEGV) on arm64 macOS; behavior preserves existing maxItems cap and returns equivalent aggregated results.
|
||||
const moduleHitsMap = new Map<string, number>();
|
||||
const directModuleSet = new Set<string>();
|
||||
|
||||
// Helper to run a single module chunk and accumulate hits by name
|
||||
const runModuleChunk = async (idsChunk: string[]) => {
|
||||
if (!idsChunk || idsChunk.length === 0) return;
|
||||
try {
|
||||
const rows = await executeParameterized(repo.id, `
|
||||
MATCH (s)-[:CodeRelation {type: 'MEMBER_OF'}]->(c:Community)
|
||||
WHERE s.id IN $ids
|
||||
RETURN c.heuristicLabel AS name, COUNT(DISTINCT s.id) AS hits
|
||||
ORDER BY hits DESC
|
||||
LIMIT 20
|
||||
`, { ids: idsChunk }).catch(() => []);
|
||||
|
||||
for (const r of rows) {
|
||||
const name = r.name ?? r[0] ?? null;
|
||||
const hits = (r.hits ?? r[1]) || 0;
|
||||
if (!name) continue;
|
||||
moduleHitsMap.set(name, (moduleHitsMap.get(name) || 0) + hits);
|
||||
}
|
||||
} catch (e) {
|
||||
logQueryError('impact:module-chunk', e);
|
||||
}
|
||||
};
|
||||
|
||||
// Run module query chunks sequentially (safe on arm64 macOS)
|
||||
for (let i = 0; i < allIdsArr.length; i += CHUNK_SIZE) {
|
||||
const chunkIds = allIdsArr.slice(i, i + CHUNK_SIZE);
|
||||
await runModuleChunk(chunkIds);
|
||||
}
|
||||
|
||||
// Run direct module query similarly (distinct heuristic labels for depth-1 items)
|
||||
const runDirectModuleChunk = async (idsChunk: string[]) => {
|
||||
if (!idsChunk || idsChunk.length === 0) return;
|
||||
try {
|
||||
const rows = await executeParameterized(repo.id, `
|
||||
MATCH (s)-[:CodeRelation {type: 'MEMBER_OF'}]->(c:Community)
|
||||
WHERE s.id IN $ids
|
||||
RETURN DISTINCT c.heuristicLabel AS name
|
||||
`, { ids: idsChunk }).catch(() => []);
|
||||
for (const r of rows) {
|
||||
const name = r.name ?? r[0] ?? null;
|
||||
if (name) directModuleSet.add(name);
|
||||
}
|
||||
} catch (e) {
|
||||
logQueryError('impact:direct-module-chunk', e);
|
||||
}
|
||||
};
|
||||
|
||||
for (let i = 0; i < d1IdsArr.length; i += CHUNK_SIZE) {
|
||||
const chunkIds = d1IdsArr.slice(i, i + CHUNK_SIZE);
|
||||
await runDirectModuleChunk(chunkIds);
|
||||
}
|
||||
|
||||
// Build final moduleRows array from aggregated hits map, sorted & limited
|
||||
const moduleRows = Array.from(moduleHitsMap.entries())
|
||||
.map(([name, hits]) => ({ name, hits }))
|
||||
.sort((a, b) => b.hits - a.hits)
|
||||
.slice(0, 20);
|
||||
|
||||
const directModuleRows = Array.from(directModuleSet).map(name => ({ name }));
|
||||
|
||||
// Build affectedModules in the same shape as original implementation
|
||||
const directModuleNameSet = new Set(directModuleRows.map((r: any) => r.name || r[0]));
|
||||
affectedModules = moduleRows.map((r: any) => {
|
||||
const name = r.name || r[0];
|
||||
const name = r.name ?? r[0];
|
||||
const hits = r.hits ?? r[1] ?? 0;
|
||||
return {
|
||||
name,
|
||||
hits: r.hits || r[1],
|
||||
impact: directModuleSet.has(name) ? 'direct' : 'indirect',
|
||||
hits,
|
||||
impact: directModuleNameSet.has(name) ? 'direct' : 'indirect',
|
||||
};
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
// Risk scoring
|
||||
const processCount = affectedProcesses.length;
|
||||
|
|
|
|||
25
gitnexus/test/fixtures/lang-resolution/cobol-app/AUDITLOG.cbl
vendored
Normal file
25
gitnexus/test/fixtures/lang-resolution/cobol-app/AUDITLOG.cbl
vendored
Normal file
|
|
@ -0,0 +1,25 @@
|
|||
IDENTIFICATION DIVISION.
|
||||
PROGRAM-ID. AUDITLOG.
|
||||
|
||||
DATA DIVISION.
|
||||
WORKING-STORAGE SECTION.
|
||||
01 WS-LOG-MESSAGE PIC X(80).
|
||||
01 WS-TIMESTAMP PIC X(26).
|
||||
|
||||
LINKAGE SECTION.
|
||||
01 LS-CUST-ID PIC 9(8).
|
||||
01 LS-AMOUNT PIC 9(7)V99.
|
||||
|
||||
PROCEDURE DIVISION USING LS-CUST-ID LS-AMOUNT.
|
||||
MAIN-PARAGRAPH.
|
||||
PERFORM WRITE-LOG
|
||||
GOBACK.
|
||||
|
||||
WRITE-LOG.
|
||||
STRING 'Customer ' LS-CUST-ID ' amount ' LS-AMOUNT
|
||||
DELIMITED BY SIZE INTO WS-LOG-MESSAGE
|
||||
DISPLAY WS-LOG-MESSAGE.
|
||||
|
||||
ENTRY "AUDITLOG-BATCH" USING LS-CUST-ID.
|
||||
DISPLAY 'Batch audit for ' LS-CUST-ID
|
||||
GOBACK.
|
||||
3
gitnexus/test/fixtures/lang-resolution/cobol-app/COPYLIB.cpy
vendored
Normal file
3
gitnexus/test/fixtures/lang-resolution/cobol-app/COPYLIB.cpy
vendored
Normal file
|
|
@ -0,0 +1,3 @@
|
|||
01 PREFIX-RECORD.
|
||||
05 PREFIX-CODE PIC X(10).
|
||||
05 PREFIX-NAME PIC X(30).
|
||||
6
gitnexus/test/fixtures/lang-resolution/cobol-app/CUSTDAT.cpy
vendored
Normal file
6
gitnexus/test/fixtures/lang-resolution/cobol-app/CUSTDAT.cpy
vendored
Normal file
|
|
@ -0,0 +1,6 @@
|
|||
01 WS-CUSTOMER-DATA.
|
||||
05 WS-CUST-CODE PIC X(10).
|
||||
05 WS-CUST-TYPE PIC X(3).
|
||||
88 PREMIUM-CUSTOMER VALUE 'PRM'.
|
||||
88 REGULAR-CUSTOMER VALUE 'REG'.
|
||||
05 WS-CUST-ADDR PIC X(50).
|
||||
74
gitnexus/test/fixtures/lang-resolution/cobol-app/CUSTUPDT.cbl
vendored
Normal file
74
gitnexus/test/fixtures/lang-resolution/cobol-app/CUSTUPDT.cbl
vendored
Normal file
|
|
@ -0,0 +1,74 @@
|
|||
IDENTIFICATION DIVISION.
|
||||
PROGRAM-ID. CUSTUPDT.
|
||||
AUTHOR. TEST.
|
||||
|
||||
ENVIRONMENT DIVISION.
|
||||
INPUT-OUTPUT SECTION.
|
||||
FILE-CONTROL.
|
||||
SELECT CUSTOMER-FILE ASSIGN TO 'CUSTFILE'
|
||||
ORGANIZATION IS INDEXED
|
||||
ACCESS IS DYNAMIC
|
||||
RECORD KEY IS CUST-ID
|
||||
FILE STATUS IS WS-FILE-STATUS.
|
||||
|
||||
DATA DIVISION.
|
||||
FILE SECTION.
|
||||
FD CUSTOMER-FILE.
|
||||
01 CUSTOMER-RECORD.
|
||||
05 CUST-ID PIC 9(8).
|
||||
05 CUST-NAME PIC X(30).
|
||||
05 CUST-BALANCE PIC 9(7)V99.
|
||||
|
||||
WORKING-STORAGE SECTION.
|
||||
01 WS-FILE-STATUS PIC XX.
|
||||
01 WS-CUSTOMER-NAME PIC X(30).
|
||||
01 WS-AMOUNT PIC 9(7)V99.
|
||||
01 WS-EOF PIC 9 VALUE 0.
|
||||
88 END-OF-FILE VALUE 1.
|
||||
01 WS-AMT PIC 9(5)V99.
|
||||
01 WS-PROG-NAME PIC X(8).
|
||||
01 FIELD-A PIC 9(5)V99.
|
||||
01 FIELD-B PIC 9(5)V99.
|
||||
COPY COPYLIB REPLACING ==PREFIX-== BY ==WS-==.
|
||||
|
||||
LINKAGE SECTION.
|
||||
01 LS-PARAM PIC X(20).
|
||||
|
||||
PROCEDURE DIVISION.
|
||||
INIT-SECTION SECTION.
|
||||
MAIN-PARAGRAPH.
|
||||
PERFORM INIT-PARAGRAPH
|
||||
PERFORM PROCESS-PARAGRAPH
|
||||
PERFORM CLEANUP-PARAGRAPH
|
||||
STOP RUN.
|
||||
|
||||
INIT-PARAGRAPH.
|
||||
OPEN I-O CUSTOMER-FILE
|
||||
MOVE SPACES TO WS-CUSTOMER-NAME.
|
||||
|
||||
PROCESSING-SECTION SECTION.
|
||||
PROCESS-PARAGRAPH.
|
||||
PERFORM READ-CUSTOMER THRU WRITE-CUSTOMER
|
||||
CALL "AUDITLOG" USING CUST-ID WS-AMOUNT
|
||||
CALL WS-PROG-NAME.
|
||||
|
||||
READ-CUSTOMER.
|
||||
READ CUSTOMER-FILE
|
||||
NOT AT END
|
||||
MOVE CUST-NAME TO WS-CUSTOMER-NAME
|
||||
END-READ.
|
||||
|
||||
UPDATE-BALANCE.
|
||||
ADD WS-AMOUNT TO CUST-BALANCE
|
||||
MOVE WS-AMOUNT TO CUST-BALANCE
|
||||
MOVE WS-AMT TO FIELD-A FIELD-B.
|
||||
|
||||
WRITE-CUSTOMER.
|
||||
REWRITE CUSTOMER-RECORD.
|
||||
|
||||
CLEANUP-PARAGRAPH.
|
||||
CLOSE CUSTOMER-FILE.
|
||||
|
||||
ENTRY 'ALTENTRY' USING LS-PARAM.
|
||||
DISPLAY 'ALTERNATE ENTRY POINT'
|
||||
GOBACK.
|
||||
33
gitnexus/test/fixtures/lang-resolution/cobol-app/NESTED.cbl
vendored
Normal file
33
gitnexus/test/fixtures/lang-resolution/cobol-app/NESTED.cbl
vendored
Normal file
|
|
@ -0,0 +1,33 @@
|
|||
IDENTIFICATION DIVISION.
|
||||
PROGRAM-ID. OUTER-PROG.
|
||||
|
||||
DATA DIVISION.
|
||||
WORKING-STORAGE SECTION.
|
||||
01 WS-OUTER-FLAG PIC 9 VALUE 0.
|
||||
|
||||
PROCEDURE DIVISION.
|
||||
OUTER-MAIN.
|
||||
PERFORM OUTER-PROCESS
|
||||
CALL "INNER-PROG"
|
||||
STOP RUN.
|
||||
|
||||
OUTER-PROCESS.
|
||||
DISPLAY 'OUTER PROCESSING'.
|
||||
|
||||
IDENTIFICATION DIVISION.
|
||||
PROGRAM-ID. INNER-PROG.
|
||||
|
||||
DATA DIVISION.
|
||||
WORKING-STORAGE SECTION.
|
||||
01 WS-INNER-CODE PIC X(5).
|
||||
|
||||
PROCEDURE DIVISION.
|
||||
INNER-MAIN.
|
||||
PERFORM INNER-PROCESS
|
||||
GOBACK.
|
||||
|
||||
INNER-PROCESS.
|
||||
DISPLAY 'INNER PROCESSING'.
|
||||
|
||||
END PROGRAM INNER-PROG.
|
||||
END PROGRAM OUTER-PROG.
|
||||
94
gitnexus/test/fixtures/lang-resolution/cobol-app/RPTGEN.cbl
vendored
Normal file
94
gitnexus/test/fixtures/lang-resolution/cobol-app/RPTGEN.cbl
vendored
Normal file
|
|
@ -0,0 +1,94 @@
|
|||
IDENTIFICATION DIVISION.
|
||||
PROGRAM-ID. RPTGEN.
|
||||
|
||||
DATA DIVISION.
|
||||
WORKING-STORAGE SECTION.
|
||||
COPY CUSTDAT.
|
||||
01 WS-REPORT-LINE PIC X(132).
|
||||
01 WS-SQL-CODE PIC S9(9) COMP.
|
||||
01 WS-COUNT PIC 9(4).
|
||||
01 WS-MAP-NAME PIC X(8).
|
||||
01 WS-SORT-FILE PIC X(8).
|
||||
01 WS-QUEUE-NAME PIC X(16).
|
||||
01 WS-NEXT-PGM PIC X(8).
|
||||
|
||||
PROCEDURE DIVISION.
|
||||
MAIN-PARAGRAPH.
|
||||
PERFORM FETCH-DATA
|
||||
PERFORM FORMAT-REPORT
|
||||
PERFORM SEND-SCREEN
|
||||
CALL "CUSTUPDT"
|
||||
GO TO EXIT-PARAGRAPH.
|
||||
|
||||
FETCH-DATA.
|
||||
EXEC SQL
|
||||
SELECT CUST_NAME, CUST_BALANCE
|
||||
FROM CUSTOMER
|
||||
WHERE CUST_ID = :WS-CUST-CODE
|
||||
END-EXEC.
|
||||
|
||||
FORMAT-REPORT.
|
||||
PERFORM WS-COUNT TIMES
|
||||
MOVE WS-CUST-CODE TO WS-REPORT-LINE
|
||||
END-PERFORM
|
||||
PERFORM MAIN-PARAGRAPH THRU FORMAT-REPORT
|
||||
IF WS-COUNT > 0 PERFORM FETCH-DATA
|
||||
ELSE PERFORM SEND-SCREEN
|
||||
END-IF
|
||||
SORT WS-SORT-FILE USING CUSTOMER-DATA
|
||||
GIVING WS-REPORT-LINE.
|
||||
SORT WS-SORT-FILE ON ASCENDING KEY WS-COUNT
|
||||
INPUT PROCEDURE IS BUILD-SORT-INPUT
|
||||
OUTPUT PROCEDURE IS WRITE-SORTED.
|
||||
MOVE CORR WS-CUSTOMER-DATA TO WS-REPORT-LINE
|
||||
SEARCH WS-CUSTOMER-DATA
|
||||
GO TO FETCH-DATA FORMAT-REPORT SEND-SCREEN
|
||||
DEPENDING ON WS-COUNT.
|
||||
|
||||
SEND-SCREEN.
|
||||
EXEC CICS
|
||||
SEND MAP(WS-MAP-NAME) MAPSET('CUSTSET')
|
||||
FROM(WS-REPORT-LINE)
|
||||
END-EXEC.
|
||||
|
||||
EXEC CICS
|
||||
LINK PROGRAM('AUDITLOG')
|
||||
END-EXEC.
|
||||
|
||||
EXEC CICS
|
||||
XCTL PROGRAM('CUSTUPDT')
|
||||
END-EXEC.
|
||||
|
||||
EXEC CICS
|
||||
READ FILE('CUSTFILE')
|
||||
INTO(WS-CUSTOMER-DATA)
|
||||
END-EXEC.
|
||||
|
||||
EXEC CICS
|
||||
WRITEQ TS QUEUE('RPTQUEUE')
|
||||
FROM(WS-REPORT-LINE)
|
||||
END-EXEC.
|
||||
|
||||
EXEC CICS
|
||||
HANDLE ABEND LABEL(ABEND-HANDLER)
|
||||
END-EXEC.
|
||||
|
||||
EXEC CICS
|
||||
RETURN TRANSID('RPTG')
|
||||
END-EXEC.
|
||||
|
||||
EXEC CICS
|
||||
XCTL PROGRAM(WS-NEXT-PGM)
|
||||
END-EXEC.
|
||||
|
||||
BUILD-SORT-INPUT.
|
||||
DISPLAY 'BUILDING SORT INPUT'.
|
||||
|
||||
WRITE-SORTED.
|
||||
DISPLAY 'WRITING SORTED OUTPUT'.
|
||||
|
||||
ABEND-HANDLER.
|
||||
DISPLAY 'ABEND OCCURRED'.
|
||||
|
||||
EXIT-PARAGRAPH.
|
||||
STOP RUN.
|
||||
5
gitnexus/test/fixtures/lang-resolution/cobol-app/RUNJOBS.jcl
vendored
Normal file
5
gitnexus/test/fixtures/lang-resolution/cobol-app/RUNJOBS.jcl
vendored
Normal file
|
|
@ -0,0 +1,5 @@
|
|||
//CUSTJOB JOB (ACCT),'CUSTOMER UPDATE',CLASS=A,MSGCLASS=X
|
||||
//STEP1 EXEC PGM=CUSTUPDT
|
||||
//CUSTFILE DD DSN=PROD.CUSTOMER.MASTER,DISP=SHR
|
||||
//STEP2 EXEC PGM=RPTGEN
|
||||
//SYSOUT DD SYSOUT=*
|
||||
6
gitnexus/test/fixtures/lang-resolution/dart-call-result-binding/app.dart
vendored
Normal file
6
gitnexus/test/fixtures/lang-resolution/dart-call-result-binding/app.dart
vendored
Normal file
|
|
@ -0,0 +1,6 @@
|
|||
import 'models.dart';
|
||||
|
||||
void processUser() {
|
||||
var user = getUser('alice');
|
||||
user.save();
|
||||
}
|
||||
11
gitnexus/test/fixtures/lang-resolution/dart-call-result-binding/models.dart
vendored
Normal file
11
gitnexus/test/fixtures/lang-resolution/dart-call-result-binding/models.dart
vendored
Normal file
|
|
@ -0,0 +1,11 @@
|
|||
class User {
|
||||
String name = '';
|
||||
|
||||
bool save() {
|
||||
return true;
|
||||
}
|
||||
}
|
||||
|
||||
User getUser(String name) {
|
||||
return User();
|
||||
}
|
||||
5
gitnexus/test/fixtures/lang-resolution/dart-field-types/app.dart
vendored
Normal file
5
gitnexus/test/fixtures/lang-resolution/dart-field-types/app.dart
vendored
Normal file
|
|
@ -0,0 +1,5 @@
|
|||
import 'models.dart';
|
||||
|
||||
void processUser(User user) {
|
||||
user.address.save();
|
||||
}
|
||||
16
gitnexus/test/fixtures/lang-resolution/dart-field-types/models.dart
vendored
Normal file
16
gitnexus/test/fixtures/lang-resolution/dart-field-types/models.dart
vendored
Normal file
|
|
@ -0,0 +1,16 @@
|
|||
class Address {
|
||||
String city = '';
|
||||
|
||||
void save() {
|
||||
// persist address
|
||||
}
|
||||
}
|
||||
|
||||
class User {
|
||||
String name = '';
|
||||
Address address = Address();
|
||||
|
||||
String greet() {
|
||||
return name;
|
||||
}
|
||||
}
|
||||
608
gitnexus/test/integration/resolvers/cobol.test.ts
Normal file
608
gitnexus/test/integration/resolvers/cobol.test.ts
Normal file
|
|
@ -0,0 +1,608 @@
|
|||
/**
|
||||
* COBOL: Exhaustive strict integration test.
|
||||
*
|
||||
* Every single node and edge produced by the COBOL/JCL pipeline is asserted
|
||||
* with exact counts AND exact sorted edge-pair lists. No fuzzy assertions.
|
||||
*
|
||||
* Ground truth captured from the cobol-app fixture:
|
||||
* CUSTUPDT.cbl, AUDITLOG.cbl, RPTGEN.cbl, NESTED.cbl,
|
||||
* CUSTDAT.cpy, COPYLIB.cpy, RUNJOBS.jcl
|
||||
*/
|
||||
import { describe, it, expect, beforeAll } from 'vitest';
|
||||
import path from 'path';
|
||||
import {
|
||||
FIXTURES, getRelationships, getNodesByLabel, edgeSet,
|
||||
runPipelineFromRepo, type PipelineResult,
|
||||
} from './helpers.js';
|
||||
|
||||
describe('COBOL full system extraction', () => {
|
||||
let result: PipelineResult;
|
||||
|
||||
beforeAll(async () => {
|
||||
result = await runPipelineFromRepo(
|
||||
path.join(FIXTURES, 'cobol-app'),
|
||||
() => {},
|
||||
{ skipGraphPhases: true },
|
||||
);
|
||||
}, 60000);
|
||||
|
||||
// =====================================================================
|
||||
// NODE COMPLETENESS — exact count + exact sorted name list per label
|
||||
// =====================================================================
|
||||
|
||||
describe('node completeness', () => {
|
||||
|
||||
it('produces exactly 5 Module nodes', () => {
|
||||
const nodes = getNodesByLabel(result, 'Module');
|
||||
expect(nodes.length).toBe(5);
|
||||
expect(nodes).toEqual(['AUDITLOG', 'CUSTUPDT', 'INNER-PROG', 'OUTER-PROG', 'RPTGEN']);
|
||||
});
|
||||
|
||||
it('produces exactly 21 Function nodes', () => {
|
||||
const nodes = getNodesByLabel(result, 'Function');
|
||||
expect(nodes.length).toBe(21);
|
||||
expect(nodes).toEqual([
|
||||
'ABEND-HANDLER', 'BUILD-SORT-INPUT', 'CLEANUP-PARAGRAPH',
|
||||
'EXIT-PARAGRAPH', 'FETCH-DATA', 'FORMAT-REPORT', 'INIT-PARAGRAPH',
|
||||
'INNER-MAIN', 'INNER-PROCESS',
|
||||
'MAIN-PARAGRAPH', 'MAIN-PARAGRAPH', 'MAIN-PARAGRAPH',
|
||||
'OUTER-MAIN', 'OUTER-PROCESS',
|
||||
'PROCESS-PARAGRAPH', 'READ-CUSTOMER', 'SEND-SCREEN',
|
||||
'UPDATE-BALANCE', 'WRITE-CUSTOMER', 'WRITE-LOG', 'WRITE-SORTED',
|
||||
]);
|
||||
});
|
||||
|
||||
it('produces exactly 2 Namespace nodes', () => {
|
||||
expect(getNodesByLabel(result, 'Namespace')).toEqual(['INIT-SECTION', 'PROCESSING-SECTION']);
|
||||
});
|
||||
|
||||
it('produces exactly 36 Property nodes', () => {
|
||||
const nodes = getNodesByLabel(result, 'Property');
|
||||
expect(nodes.length).toBe(36);
|
||||
expect(nodes).toEqual([
|
||||
'CUST-BALANCE', 'CUST-ID', 'CUST-NAME', 'CUSTOMER-RECORD',
|
||||
'END-OF-FILE', 'FIELD-A', 'FIELD-B',
|
||||
'LS-AMOUNT', 'LS-CUST-ID', 'LS-PARAM',
|
||||
'PREMIUM-CUSTOMER', 'REGULAR-CUSTOMER',
|
||||
'WS-AMOUNT', 'WS-AMT', 'WS-CODE', 'WS-COUNT',
|
||||
'WS-CUST-ADDR', 'WS-CUST-CODE', 'WS-CUST-TYPE',
|
||||
'WS-CUSTOMER-DATA', 'WS-CUSTOMER-NAME', 'WS-EOF',
|
||||
'WS-FILE-STATUS', 'WS-INNER-CODE', 'WS-LOG-MESSAGE',
|
||||
'WS-MAP-NAME', 'WS-NAME', 'WS-NEXT-PGM', 'WS-OUTER-FLAG',
|
||||
'WS-PROG-NAME', 'WS-QUEUE-NAME', 'WS-RECORD',
|
||||
'WS-REPORT-LINE', 'WS-SORT-FILE', 'WS-SQL-CODE', 'WS-TIMESTAMP',
|
||||
]);
|
||||
});
|
||||
|
||||
it('produces exactly 1 Record node', () => {
|
||||
expect(getNodesByLabel(result, 'Record')).toEqual(['CUSTOMER-FILE']);
|
||||
});
|
||||
|
||||
it('produces exactly 15 CodeElement nodes', () => {
|
||||
const nodes = getNodesByLabel(result, 'CodeElement');
|
||||
expect(nodes.length).toBe(15);
|
||||
expect(nodes).toEqual([
|
||||
'CALL WS-PROG-NAME', 'CICS XCTL WS-NEXT-PGM', 'CUSTJOB',
|
||||
'EXEC CICS HANDLE ABEND', 'EXEC CICS LINK', 'EXEC CICS READ',
|
||||
'EXEC CICS RETURN', 'EXEC CICS SEND MAP', 'EXEC CICS WRITEQ TS',
|
||||
'EXEC CICS XCTL', 'EXEC CICS XCTL', 'EXEC SQL SELECT',
|
||||
'PROD.CUSTOMER.MASTER', 'STEP1', 'STEP2',
|
||||
]);
|
||||
});
|
||||
|
||||
it('produces exactly 2 Constructor nodes', () => {
|
||||
expect(getNodesByLabel(result, 'Constructor')).toEqual(['ALTENTRY', 'AUDITLOG-BATCH']);
|
||||
});
|
||||
});
|
||||
|
||||
// =====================================================================
|
||||
// CALLS EDGES — exact count + exact sorted pairs per reason
|
||||
// =====================================================================
|
||||
|
||||
describe('CALLS edge completeness', () => {
|
||||
|
||||
it('produces exactly 15 CALLS edges with reason cobol-perform', () => {
|
||||
const edges = getRelationships(result, 'CALLS').filter(e => e.rel.reason === 'cobol-perform');
|
||||
expect(edges.length).toBe(15);
|
||||
expect(edgeSet(edges)).toEqual([
|
||||
'FORMAT-REPORT \u2192 BUILD-SORT-INPUT',
|
||||
'FORMAT-REPORT \u2192 FETCH-DATA',
|
||||
'FORMAT-REPORT \u2192 MAIN-PARAGRAPH',
|
||||
'FORMAT-REPORT \u2192 SEND-SCREEN',
|
||||
'FORMAT-REPORT \u2192 WRITE-SORTED',
|
||||
'INNER-MAIN \u2192 INNER-PROCESS',
|
||||
'MAIN-PARAGRAPH \u2192 CLEANUP-PARAGRAPH',
|
||||
'MAIN-PARAGRAPH \u2192 FETCH-DATA',
|
||||
'MAIN-PARAGRAPH \u2192 FORMAT-REPORT',
|
||||
'MAIN-PARAGRAPH \u2192 INIT-PARAGRAPH',
|
||||
'MAIN-PARAGRAPH \u2192 PROCESS-PARAGRAPH',
|
||||
'MAIN-PARAGRAPH \u2192 SEND-SCREEN',
|
||||
'MAIN-PARAGRAPH \u2192 WRITE-LOG',
|
||||
'OUTER-MAIN \u2192 OUTER-PROCESS',
|
||||
'PROCESS-PARAGRAPH \u2192 READ-CUSTOMER',
|
||||
]);
|
||||
});
|
||||
|
||||
it('produces exactly 2 CALLS edges with reason cobol-perform-thru', () => {
|
||||
const edges = getRelationships(result, 'CALLS').filter(e => e.rel.reason === 'cobol-perform-thru');
|
||||
expect(edges.length).toBe(2);
|
||||
expect(edgeSet(edges)).toEqual([
|
||||
'FORMAT-REPORT \u2192 FORMAT-REPORT',
|
||||
'PROCESS-PARAGRAPH \u2192 WRITE-CUSTOMER',
|
||||
]);
|
||||
});
|
||||
|
||||
it('produces exactly 3 CALLS edges with reason cobol-call', () => {
|
||||
const edges = getRelationships(result, 'CALLS').filter(e => e.rel.reason === 'cobol-call');
|
||||
expect(edges.length).toBe(3);
|
||||
expect(edgeSet(edges)).toEqual([
|
||||
'CUSTUPDT \u2192 AUDITLOG',
|
||||
'OUTER-PROG \u2192 INNER-PROG',
|
||||
'RPTGEN \u2192 CUSTUPDT',
|
||||
]);
|
||||
});
|
||||
|
||||
it('produces exactly 4 CALLS edges with reason cobol-goto', () => {
|
||||
const edges = getRelationships(result, 'CALLS').filter(e => e.rel.reason === 'cobol-goto');
|
||||
expect(edges.length).toBe(4);
|
||||
expect(edgeSet(edges)).toEqual([
|
||||
'FORMAT-REPORT \u2192 FETCH-DATA',
|
||||
'FORMAT-REPORT \u2192 FORMAT-REPORT',
|
||||
'FORMAT-REPORT \u2192 SEND-SCREEN',
|
||||
'MAIN-PARAGRAPH \u2192 EXIT-PARAGRAPH',
|
||||
]);
|
||||
});
|
||||
|
||||
it('produces exactly 1 CALLS edge with reason cics-link', () => {
|
||||
const edges = getRelationships(result, 'CALLS').filter(e => e.rel.reason === 'cics-link');
|
||||
expect(edges.length).toBe(1);
|
||||
expect(edgeSet(edges)).toEqual(['RPTGEN \u2192 AUDITLOG']);
|
||||
});
|
||||
|
||||
it('produces exactly 1 CALLS edge with reason cics-xctl', () => {
|
||||
const edges = getRelationships(result, 'CALLS').filter(e => e.rel.reason === 'cics-xctl');
|
||||
expect(edges.length).toBe(1);
|
||||
expect(edgeSet(edges)).toEqual(['RPTGEN \u2192 CUSTUPDT']);
|
||||
});
|
||||
|
||||
it('produces exactly 1 CALLS edge with reason cics-handle-abend', () => {
|
||||
const edges = getRelationships(result, 'CALLS').filter(e => e.rel.reason === 'cics-handle-abend');
|
||||
expect(edges.length).toBe(1);
|
||||
expect(edgeSet(edges)).toEqual(['RPTGEN \u2192 ABEND-HANDLER']);
|
||||
});
|
||||
|
||||
it('produces exactly 1 CALLS edge with reason cics-return-transid', () => {
|
||||
const edges = getRelationships(result, 'CALLS').filter(e => e.rel.reason === 'cics-return-transid');
|
||||
expect(edges.length).toBe(1);
|
||||
});
|
||||
|
||||
it('produces exactly 2 CALLS edges with reason jcl-exec-pgm', () => {
|
||||
const edges = getRelationships(result, 'CALLS').filter(e => e.rel.reason === 'jcl-exec-pgm');
|
||||
expect(edges.length).toBe(2);
|
||||
expect(edgeSet(edges)).toEqual(['STEP1 \u2192 CUSTUPDT', 'STEP2 \u2192 RPTGEN']);
|
||||
});
|
||||
|
||||
it('produces exactly 1 CALLS edge with reason jcl-dd:CUSTFILE', () => {
|
||||
const edges = getRelationships(result, 'CALLS').filter(e => e.rel.reason === 'jcl-dd:CUSTFILE');
|
||||
expect(edges.length).toBe(1);
|
||||
expect(edgeSet(edges)).toEqual(['STEP1 \u2192 PROD.CUSTOMER.MASTER']);
|
||||
});
|
||||
|
||||
it('produces zero unresolved CALLS edges', () => {
|
||||
expect(getRelationships(result, 'CALLS').filter(e => e.rel.reason.endsWith('-unresolved')).length).toBe(0);
|
||||
});
|
||||
});
|
||||
|
||||
// =====================================================================
|
||||
// CONTAINS EDGES — exact count + exact sorted pairs per reason
|
||||
// =====================================================================
|
||||
|
||||
describe('CONTAINS edge completeness', () => {
|
||||
|
||||
it('produces exactly 4 CONTAINS edges with reason cobol-program-id', () => {
|
||||
const edges = getRelationships(result, 'CONTAINS').filter(e => e.rel.reason === 'cobol-program-id');
|
||||
expect(edges.length).toBe(4);
|
||||
expect(edgeSet(edges)).toEqual([
|
||||
'AUDITLOG.cbl \u2192 AUDITLOG',
|
||||
'CUSTUPDT.cbl \u2192 CUSTUPDT',
|
||||
'NESTED.cbl \u2192 OUTER-PROG',
|
||||
'RPTGEN.cbl \u2192 RPTGEN',
|
||||
]);
|
||||
});
|
||||
|
||||
it('produces exactly 1 CONTAINS edge with reason cobol-nested-program', () => {
|
||||
const edges = getRelationships(result, 'CONTAINS').filter(e => e.rel.reason === 'cobol-nested-program');
|
||||
expect(edges.length).toBe(1);
|
||||
expect(edgeSet(edges)).toEqual(['OUTER-PROG \u2192 INNER-PROG']);
|
||||
});
|
||||
|
||||
it('produces exactly 2 CONTAINS edges with reason cobol-section', () => {
|
||||
const edges = getRelationships(result, 'CONTAINS').filter(e => e.rel.reason === 'cobol-section');
|
||||
expect(edges.length).toBe(2);
|
||||
expect(edgeSet(edges)).toEqual([
|
||||
'CUSTUPDT \u2192 INIT-SECTION',
|
||||
'CUSTUPDT \u2192 PROCESSING-SECTION',
|
||||
]);
|
||||
});
|
||||
|
||||
it('produces exactly 21 CONTAINS edges with reason cobol-paragraph', () => {
|
||||
const edges = getRelationships(result, 'CONTAINS').filter(e => e.rel.reason === 'cobol-paragraph');
|
||||
expect(edges.length).toBe(21);
|
||||
expect(edgeSet(edges)).toEqual([
|
||||
'AUDITLOG \u2192 MAIN-PARAGRAPH',
|
||||
'AUDITLOG \u2192 WRITE-LOG',
|
||||
'INIT-SECTION \u2192 INIT-PARAGRAPH',
|
||||
'INIT-SECTION \u2192 MAIN-PARAGRAPH',
|
||||
'INNER-PROG \u2192 INNER-MAIN',
|
||||
'INNER-PROG \u2192 INNER-PROCESS',
|
||||
'OUTER-PROG \u2192 OUTER-MAIN',
|
||||
'OUTER-PROG \u2192 OUTER-PROCESS',
|
||||
'PROCESSING-SECTION \u2192 CLEANUP-PARAGRAPH',
|
||||
'PROCESSING-SECTION \u2192 PROCESS-PARAGRAPH',
|
||||
'PROCESSING-SECTION \u2192 READ-CUSTOMER',
|
||||
'PROCESSING-SECTION \u2192 UPDATE-BALANCE',
|
||||
'PROCESSING-SECTION \u2192 WRITE-CUSTOMER',
|
||||
'RPTGEN \u2192 ABEND-HANDLER',
|
||||
'RPTGEN \u2192 BUILD-SORT-INPUT',
|
||||
'RPTGEN \u2192 EXIT-PARAGRAPH',
|
||||
'RPTGEN \u2192 FETCH-DATA',
|
||||
'RPTGEN \u2192 FORMAT-REPORT',
|
||||
'RPTGEN \u2192 MAIN-PARAGRAPH',
|
||||
'RPTGEN \u2192 SEND-SCREEN',
|
||||
'RPTGEN \u2192 WRITE-SORTED',
|
||||
]);
|
||||
});
|
||||
|
||||
it('produces exactly 36 CONTAINS edges with reason cobol-data-item', () => {
|
||||
const edges = getRelationships(result, 'CONTAINS').filter(e => e.rel.reason === 'cobol-data-item');
|
||||
expect(edges.length).toBe(36);
|
||||
expect(edgeSet(edges)).toEqual([
|
||||
'AUDITLOG \u2192 LS-AMOUNT',
|
||||
'AUDITLOG \u2192 LS-CUST-ID',
|
||||
'AUDITLOG \u2192 WS-LOG-MESSAGE',
|
||||
'AUDITLOG \u2192 WS-TIMESTAMP',
|
||||
'CUSTUPDT \u2192 CUST-BALANCE',
|
||||
'CUSTUPDT \u2192 CUST-ID',
|
||||
'CUSTUPDT \u2192 CUST-NAME',
|
||||
'CUSTUPDT \u2192 CUSTOMER-RECORD',
|
||||
'CUSTUPDT \u2192 END-OF-FILE',
|
||||
'CUSTUPDT \u2192 FIELD-A',
|
||||
'CUSTUPDT \u2192 FIELD-B',
|
||||
'CUSTUPDT \u2192 LS-PARAM',
|
||||
'CUSTUPDT \u2192 WS-AMOUNT',
|
||||
'CUSTUPDT \u2192 WS-AMT',
|
||||
'CUSTUPDT \u2192 WS-CODE',
|
||||
'CUSTUPDT \u2192 WS-CUSTOMER-NAME',
|
||||
'CUSTUPDT \u2192 WS-EOF',
|
||||
'CUSTUPDT \u2192 WS-FILE-STATUS',
|
||||
'CUSTUPDT \u2192 WS-NAME',
|
||||
'CUSTUPDT \u2192 WS-PROG-NAME',
|
||||
'CUSTUPDT \u2192 WS-RECORD',
|
||||
'INNER-PROG \u2192 WS-INNER-CODE',
|
||||
'OUTER-PROG \u2192 WS-OUTER-FLAG',
|
||||
'RPTGEN \u2192 PREMIUM-CUSTOMER',
|
||||
'RPTGEN \u2192 REGULAR-CUSTOMER',
|
||||
'RPTGEN \u2192 WS-COUNT',
|
||||
'RPTGEN \u2192 WS-CUST-ADDR',
|
||||
'RPTGEN \u2192 WS-CUST-CODE',
|
||||
'RPTGEN \u2192 WS-CUST-TYPE',
|
||||
'RPTGEN \u2192 WS-CUSTOMER-DATA',
|
||||
'RPTGEN \u2192 WS-MAP-NAME',
|
||||
'RPTGEN \u2192 WS-NEXT-PGM',
|
||||
'RPTGEN \u2192 WS-QUEUE-NAME',
|
||||
'RPTGEN \u2192 WS-REPORT-LINE',
|
||||
'RPTGEN \u2192 WS-SORT-FILE',
|
||||
'RPTGEN \u2192 WS-SQL-CODE',
|
||||
]);
|
||||
});
|
||||
|
||||
it('produces exactly 8 CONTAINS edges with reason cobol-exec-cics', () => {
|
||||
const edges = getRelationships(result, 'CONTAINS').filter(e => e.rel.reason === 'cobol-exec-cics');
|
||||
expect(edges.length).toBe(8);
|
||||
expect(edgeSet(edges)).toEqual([
|
||||
'RPTGEN \u2192 EXEC CICS HANDLE ABEND',
|
||||
'RPTGEN \u2192 EXEC CICS LINK',
|
||||
'RPTGEN \u2192 EXEC CICS READ',
|
||||
'RPTGEN \u2192 EXEC CICS RETURN',
|
||||
'RPTGEN \u2192 EXEC CICS SEND MAP',
|
||||
'RPTGEN \u2192 EXEC CICS WRITEQ TS',
|
||||
'RPTGEN \u2192 EXEC CICS XCTL',
|
||||
'RPTGEN \u2192 EXEC CICS XCTL',
|
||||
]);
|
||||
});
|
||||
|
||||
it('produces exactly 1 CONTAINS edge with reason cobol-exec-sql', () => {
|
||||
expect(edgeSet(getRelationships(result, 'CONTAINS').filter(e => e.rel.reason === 'cobol-exec-sql')))
|
||||
.toEqual(['RPTGEN \u2192 EXEC SQL SELECT']);
|
||||
});
|
||||
|
||||
it('produces exactly 1 CONTAINS edge with reason cics-dynamic-program', () => {
|
||||
expect(edgeSet(getRelationships(result, 'CONTAINS').filter(e => e.rel.reason === 'cics-dynamic-program')))
|
||||
.toEqual(['RPTGEN \u2192 CICS XCTL WS-NEXT-PGM']);
|
||||
});
|
||||
|
||||
it('produces exactly 1 CONTAINS edge with reason cobol-dynamic-call', () => {
|
||||
expect(edgeSet(getRelationships(result, 'CONTAINS').filter(e => e.rel.reason === 'cobol-dynamic-call')))
|
||||
.toEqual(['CUSTUPDT \u2192 CALL WS-PROG-NAME']);
|
||||
});
|
||||
|
||||
it('produces exactly 2 CONTAINS edges with reason cobol-entry-point', () => {
|
||||
expect(edgeSet(getRelationships(result, 'CONTAINS').filter(e => e.rel.reason === 'cobol-entry-point')))
|
||||
.toEqual(['AUDITLOG \u2192 AUDITLOG-BATCH', 'CUSTUPDT \u2192 ALTENTRY']);
|
||||
});
|
||||
|
||||
it('produces exactly 1 CONTAINS edge with reason cobol-file-declaration', () => {
|
||||
expect(edgeSet(getRelationships(result, 'CONTAINS').filter(e => e.rel.reason === 'cobol-file-declaration')))
|
||||
.toEqual(['CUSTUPDT \u2192 CUSTOMER-FILE']);
|
||||
});
|
||||
|
||||
it('produces exactly 1 CONTAINS edge with reason jcl-job', () => {
|
||||
expect(edgeSet(getRelationships(result, 'CONTAINS').filter(e => e.rel.reason === 'jcl-job')))
|
||||
.toEqual(['RUNJOBS.jcl \u2192 CUSTJOB']);
|
||||
});
|
||||
|
||||
it('produces exactly 2 CONTAINS edges with reason jcl-step', () => {
|
||||
expect(edgeSet(getRelationships(result, 'CONTAINS').filter(e => e.rel.reason === 'jcl-step')))
|
||||
.toEqual(['CUSTJOB \u2192 STEP1', 'CUSTJOB \u2192 STEP2']);
|
||||
});
|
||||
});
|
||||
|
||||
// =====================================================================
|
||||
// ACCESSES EDGES — exact count + exact sorted pairs per reason
|
||||
// =====================================================================
|
||||
|
||||
describe('ACCESSES edge completeness', () => {
|
||||
|
||||
it('produces exactly 4 ACCESSES edges with reason cobol-move-read', () => {
|
||||
const edges = getRelationships(result, 'ACCESSES').filter(e => e.rel.reason === 'cobol-move-read');
|
||||
expect(edges.length).toBe(4);
|
||||
expect(edgeSet(edges)).toEqual([
|
||||
'FORMAT-REPORT \u2192 WS-CUST-CODE',
|
||||
'READ-CUSTOMER \u2192 CUST-NAME',
|
||||
'UPDATE-BALANCE \u2192 WS-AMOUNT',
|
||||
'UPDATE-BALANCE \u2192 WS-AMT',
|
||||
]);
|
||||
});
|
||||
|
||||
it('produces exactly 5 ACCESSES edges with reason cobol-move-write', () => {
|
||||
const edges = getRelationships(result, 'ACCESSES').filter(e => e.rel.reason === 'cobol-move-write');
|
||||
expect(edges.length).toBe(5);
|
||||
expect(edgeSet(edges)).toEqual([
|
||||
'FORMAT-REPORT \u2192 WS-REPORT-LINE',
|
||||
'READ-CUSTOMER \u2192 WS-CUSTOMER-NAME',
|
||||
'UPDATE-BALANCE \u2192 CUST-BALANCE',
|
||||
'UPDATE-BALANCE \u2192 FIELD-A',
|
||||
'UPDATE-BALANCE \u2192 FIELD-B',
|
||||
]);
|
||||
});
|
||||
|
||||
it('produces exactly 1 ACCESSES edge with reason cics-file-read', () => {
|
||||
expect(getRelationships(result, 'ACCESSES').filter(e => e.rel.reason === 'cics-file-read').length).toBe(1);
|
||||
});
|
||||
|
||||
it('produces exactly 1 ACCESSES edge with reason cics-map', () => {
|
||||
expect(getRelationships(result, 'ACCESSES').filter(e => e.rel.reason === 'cics-map').length).toBe(1);
|
||||
});
|
||||
|
||||
it('produces exactly 1 ACCESSES edge with reason cics-queue-write', () => {
|
||||
expect(getRelationships(result, 'ACCESSES').filter(e => e.rel.reason === 'cics-queue-write').length).toBe(1);
|
||||
});
|
||||
|
||||
it('produces exactly 1 ACCESSES edge with reason cics-receive-into', () => {
|
||||
const edges = getRelationships(result, 'ACCESSES').filter(e => e.rel.reason === 'cics-receive-into');
|
||||
expect(edges.length).toBe(1);
|
||||
expect(edges[0].target).toBe('WS-CUSTOMER-DATA');
|
||||
});
|
||||
|
||||
it('produces exactly 2 ACCESSES edges with reason cics-send-from', () => {
|
||||
const edges = getRelationships(result, 'ACCESSES').filter(e => e.rel.reason === 'cics-send-from');
|
||||
expect(edges.length).toBe(2);
|
||||
expect(edgeSet(edges)).toEqual([
|
||||
'EXEC CICS SEND MAP \u2192 WS-REPORT-LINE',
|
||||
'EXEC CICS WRITEQ TS \u2192 WS-REPORT-LINE',
|
||||
]);
|
||||
});
|
||||
|
||||
it('produces exactly 1 ACCESSES edge with reason cobol-search', () => {
|
||||
const edges = getRelationships(result, 'ACCESSES').filter(e => e.rel.reason === 'cobol-search');
|
||||
expect(edges.length).toBe(1);
|
||||
expect(edgeSet(edges)).toEqual(['RPTGEN \u2192 WS-CUSTOMER-DATA']);
|
||||
});
|
||||
|
||||
it('produces exactly 1 ACCESSES edge with reason sort-using', () => {
|
||||
expect(getRelationships(result, 'ACCESSES').filter(e => e.rel.reason === 'sort-using').length).toBe(1);
|
||||
});
|
||||
|
||||
it('produces exactly 1 ACCESSES edge with reason sort-giving (multi-line SORT)', () => {
|
||||
expect(getRelationships(result, 'ACCESSES').filter(e => e.rel.reason === 'sort-giving').length).toBe(1);
|
||||
});
|
||||
|
||||
it('produces exactly 2 ACCESSES edges with reason cobol-procedure-using', () => {
|
||||
const edges = getRelationships(result, 'ACCESSES').filter(e => e.rel.reason === 'cobol-procedure-using');
|
||||
expect(edges.length).toBe(2);
|
||||
expect(edgeSet(edges)).toEqual([
|
||||
'AUDITLOG \u2192 LS-AMOUNT',
|
||||
'AUDITLOG \u2192 LS-CUST-ID',
|
||||
]);
|
||||
});
|
||||
|
||||
it('produces exactly 1 ACCESSES edge with reason sql-select', () => {
|
||||
expect(getRelationships(result, 'ACCESSES').filter(e => e.rel.reason === 'sql-select').length).toBe(1);
|
||||
});
|
||||
});
|
||||
|
||||
// =====================================================================
|
||||
// IMPORTS EDGES — exact pairs
|
||||
// =====================================================================
|
||||
|
||||
describe('IMPORTS edge completeness', () => {
|
||||
|
||||
it('produces exactly 2 IMPORTS edges with reason cobol-copy', () => {
|
||||
const edges = getRelationships(result, 'IMPORTS').filter(e => e.rel.reason === 'cobol-copy');
|
||||
expect(edges.length).toBe(2);
|
||||
});
|
||||
});
|
||||
|
||||
// =====================================================================
|
||||
// FEATURE-SPECIFIC ASSERTIONS — validates all review findings resolved
|
||||
// =====================================================================
|
||||
|
||||
describe('multi-PERFORM on same line (Finding #III)', () => {
|
||||
|
||||
it('captures both PERFORMs in IF/ELSE on a single logical line', () => {
|
||||
// IF WS-COUNT > 0 PERFORM FETCH-DATA ELSE PERFORM SEND-SCREEN
|
||||
const edges = getRelationships(result, 'CALLS').filter(
|
||||
e => e.rel.reason === 'cobol-perform' && e.source === 'FORMAT-REPORT',
|
||||
);
|
||||
const targets = edges.map(e => e.target).sort();
|
||||
expect(targets).toContain('FETCH-DATA');
|
||||
expect(targets).toContain('SEND-SCREEN');
|
||||
});
|
||||
});
|
||||
|
||||
describe('INPUT/OUTPUT PROCEDURE IS in SORT (Finding #iii)', () => {
|
||||
|
||||
it('creates CALLS edges for INPUT PROCEDURE and OUTPUT PROCEDURE targets', () => {
|
||||
const edges = getRelationships(result, 'CALLS').filter(
|
||||
e => e.rel.reason === 'cobol-perform' && e.source === 'FORMAT-REPORT',
|
||||
);
|
||||
const targets = edges.map(e => e.target).sort();
|
||||
expect(targets).toContain('BUILD-SORT-INPUT');
|
||||
expect(targets).toContain('WRITE-SORTED');
|
||||
});
|
||||
|
||||
it('creates paragraph nodes for INPUT/OUTPUT PROCEDURE targets', () => {
|
||||
const nodes = getNodesByLabel(result, 'Function');
|
||||
expect(nodes).toContain('BUILD-SORT-INPUT');
|
||||
expect(nodes).toContain('WRITE-SORTED');
|
||||
});
|
||||
});
|
||||
|
||||
describe('GO TO DEPENDING ON multi-target (Finding #iv)', () => {
|
||||
|
||||
it('captures all three targets from GO TO ... DEPENDING ON', () => {
|
||||
// GO TO FETCH-DATA FORMAT-REPORT SEND-SCREEN DEPENDING ON WS-COUNT
|
||||
const edges = getRelationships(result, 'CALLS').filter(
|
||||
e => e.rel.reason === 'cobol-goto' && e.source === 'FORMAT-REPORT',
|
||||
);
|
||||
expect(edges.length).toBe(3);
|
||||
expect(edgeSet(edges)).toEqual([
|
||||
'FORMAT-REPORT \u2192 FETCH-DATA',
|
||||
'FORMAT-REPORT \u2192 FORMAT-REPORT',
|
||||
'FORMAT-REPORT \u2192 SEND-SCREEN',
|
||||
]);
|
||||
});
|
||||
});
|
||||
|
||||
describe('MOVE CORR abbreviation (Finding #IV)', () => {
|
||||
|
||||
it('produces ACCESSES edges for MOVE CORR with corresponding reason', () => {
|
||||
const readEdges = getRelationships(result, 'ACCESSES').filter(
|
||||
e => e.rel.reason === 'cobol-move-corresponding-read',
|
||||
);
|
||||
expect(readEdges.length).toBe(1);
|
||||
expect(edgeSet(readEdges)).toEqual(['FORMAT-REPORT \u2192 WS-CUSTOMER-DATA']);
|
||||
|
||||
const writeEdges = getRelationships(result, 'ACCESSES').filter(
|
||||
e => e.rel.reason === 'cobol-move-corresponding-write',
|
||||
);
|
||||
expect(writeEdges.length).toBe(1);
|
||||
expect(edgeSet(writeEdges)).toEqual(['FORMAT-REPORT \u2192 WS-REPORT-LINE']);
|
||||
});
|
||||
});
|
||||
|
||||
describe('nested program CONTAINS attribution (Finding #I, #II)', () => {
|
||||
|
||||
it('attributes INNER-PROG paragraphs to INNER-PROG, not OUTER-PROG', () => {
|
||||
const edges = getRelationships(result, 'CONTAINS').filter(
|
||||
e => e.rel.reason === 'cobol-paragraph' && e.target === 'INNER-MAIN',
|
||||
);
|
||||
expect(edges.length).toBe(1);
|
||||
expect(edges[0].source).toBe('INNER-PROG');
|
||||
});
|
||||
|
||||
it('attributes INNER-PROG data items to INNER-PROG, not OUTER-PROG', () => {
|
||||
const edges = getRelationships(result, 'CONTAINS').filter(
|
||||
e => e.rel.reason === 'cobol-data-item' && e.target === 'WS-INNER-CODE',
|
||||
);
|
||||
expect(edges.length).toBe(1);
|
||||
expect(edges[0].source).toBe('INNER-PROG');
|
||||
});
|
||||
|
||||
it('attributes OUTER-PROG data items to OUTER-PROG', () => {
|
||||
const edges = getRelationships(result, 'CONTAINS').filter(
|
||||
e => e.rel.reason === 'cobol-data-item' && e.target === 'WS-OUTER-FLAG',
|
||||
);
|
||||
expect(edges.length).toBe(1);
|
||||
expect(edges[0].source).toBe('OUTER-PROG');
|
||||
});
|
||||
});
|
||||
|
||||
describe('per-program PROCEDURE DIVISION USING (Finding #III partial)', () => {
|
||||
|
||||
it('creates ACCESSES edges from AUDITLOG, not from wrong program', () => {
|
||||
const edges = getRelationships(result, 'ACCESSES').filter(
|
||||
e => e.rel.reason === 'cobol-procedure-using',
|
||||
);
|
||||
expect(edges.length).toBe(2);
|
||||
// Both edges should source from AUDITLOG (the program that declares USING)
|
||||
for (const e of edges) {
|
||||
expect(e.source).toBe('AUDITLOG');
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
describe('PERFORM THRU edge correctness', () => {
|
||||
|
||||
it('captures FORMAT-REPORT PERFORM THRU from MAIN-PARAGRAPH', () => {
|
||||
const edges = getRelationships(result, 'CALLS').filter(
|
||||
e => e.rel.reason === 'cobol-perform-thru',
|
||||
);
|
||||
expect(edgeSet(edges)).toContain('FORMAT-REPORT \u2192 FORMAT-REPORT');
|
||||
});
|
||||
});
|
||||
|
||||
describe('nested program CALLS attribution', () => {
|
||||
|
||||
it('attributes INNER-PROG PERFORM edges to INNER-PROG paragraphs', () => {
|
||||
const edges = getRelationships(result, 'CALLS').filter(
|
||||
e => e.rel.reason === 'cobol-perform' && e.source === 'INNER-MAIN',
|
||||
);
|
||||
expect(edges.length).toBe(1);
|
||||
expect(edges[0].target).toBe('INNER-PROCESS');
|
||||
});
|
||||
});
|
||||
|
||||
// =====================================================================
|
||||
// GRAND TOTALS — catch any unexpected edge leakage
|
||||
// =====================================================================
|
||||
|
||||
describe('grand totals', () => {
|
||||
|
||||
it('produces exactly 31 total CALLS edges', () => {
|
||||
// 15 perform + 2 perform-thru + 3 call + 4 goto + 1 link + 1 xctl
|
||||
// + 1 handle-abend + 1 return-transid + 2 jcl-exec-pgm + 1 jcl-dd
|
||||
expect(getRelationships(result, 'CALLS').length).toBe(31);
|
||||
});
|
||||
|
||||
it('produces exactly 81 total CONTAINS edges', () => {
|
||||
// 4 program-id + 1 nested-program + 2 section + 21 paragraph
|
||||
// + 36 data-item + 8 exec-cics + 1 exec-sql + 1 dynamic-call
|
||||
// + 1 cics-dynamic-program + 2 entry-point + 1 file-declaration
|
||||
// + 1 jcl-job + 2 jcl-step
|
||||
expect(getRelationships(result, 'CONTAINS').length).toBe(81);
|
||||
});
|
||||
|
||||
it('produces exactly 2 total IMPORTS edges', () => {
|
||||
expect(getRelationships(result, 'IMPORTS').length).toBe(2);
|
||||
});
|
||||
|
||||
it('produces exactly 25 total ACCESSES edges', () => {
|
||||
// 4 move-read + 5 move-write + 1 move-corresponding-read + 1 move-corresponding-write
|
||||
// + 1 file-read + 1 map + 1 queue-write
|
||||
// + 1 receive-into + 2 send-from + 1 search + 1 sort-using + 1 sort-giving
|
||||
// + 2 procedure-using + 1 sql-select + 2 call-using
|
||||
expect(getRelationships(result, 'ACCESSES').length).toBe(25);
|
||||
});
|
||||
});
|
||||
});
|
||||
147
gitnexus/test/integration/resolvers/dart.test.ts
Normal file
147
gitnexus/test/integration/resolvers/dart.test.ts
Normal file
|
|
@ -0,0 +1,147 @@
|
|||
/**
|
||||
* Dart: field-type resolution and call-result binding.
|
||||
* Verifies that class fields are captured as Property nodes with HAS_PROPERTY
|
||||
* edges, and that calls (including chained and call-result-bound) are resolved.
|
||||
*
|
||||
* Remaining known Dart gaps (field-chain ACCESSES) are documented as
|
||||
* it.todo() tests to be filled when the pipeline is extended.
|
||||
*/
|
||||
import { describe, it, expect, beforeAll } from 'vitest';
|
||||
import path from 'path';
|
||||
import {
|
||||
FIXTURES, getRelationships, getNodesByLabel, edgeSet,
|
||||
runPipelineFromRepo, type PipelineResult,
|
||||
} from './helpers.js';
|
||||
import { isLanguageAvailable } from '../../../src/core/tree-sitter/parser-loader.js';
|
||||
import { SupportedLanguages } from '../../../src/config/supported-languages.js';
|
||||
|
||||
const dartAvailable = isLanguageAvailable(SupportedLanguages.Dart);
|
||||
|
||||
// ── Phase 8: Field-type resolution ──────────────────────────────────────
|
||||
|
||||
describe.skipIf(!dartAvailable)('Dart field-type resolution', () => {
|
||||
let result: PipelineResult;
|
||||
|
||||
beforeAll(async () => {
|
||||
result = await runPipelineFromRepo(
|
||||
path.join(FIXTURES, 'dart-field-types'),
|
||||
() => {},
|
||||
);
|
||||
}, 60000);
|
||||
|
||||
it('detects classes and their properties', () => {
|
||||
expect(getNodesByLabel(result, 'Class')).toEqual(
|
||||
expect.arrayContaining(['Address', 'User']),
|
||||
);
|
||||
const properties = getNodesByLabel(result, 'Property');
|
||||
expect(properties).toContain('address');
|
||||
expect(properties).toContain('city');
|
||||
expect(properties).toContain('name');
|
||||
});
|
||||
|
||||
it('emits HAS_PROPERTY edges from class to field', () => {
|
||||
const propEdges = getRelationships(result, 'HAS_PROPERTY');
|
||||
expect(edgeSet(propEdges)).toEqual(
|
||||
expect.arrayContaining([
|
||||
'User → address',
|
||||
'User → name',
|
||||
'Address → city',
|
||||
]),
|
||||
);
|
||||
});
|
||||
|
||||
it('resolves save() call from field-chain user.address.save()', () => {
|
||||
const calls = getRelationships(result, 'CALLS');
|
||||
// Dart attributes calls to the enclosing Function
|
||||
const saveCalls = calls.filter(
|
||||
(c) => c.target === 'save' && c.sourceFilePath.includes('app.dart'),
|
||||
);
|
||||
expect(saveCalls.length).toBe(1);
|
||||
expect(saveCalls[0]!.targetFilePath).toContain('models.dart');
|
||||
});
|
||||
|
||||
it('attributes save() call source to processUser, not File', () => {
|
||||
const calls = getRelationships(result, 'CALLS');
|
||||
const saveCalls = calls.filter(
|
||||
(c) => c.target === 'save' && c.sourceFilePath.includes('app.dart'),
|
||||
);
|
||||
expect(saveCalls.length).toBe(1);
|
||||
expect(saveCalls[0]!.source).toBe('processUser');
|
||||
expect(saveCalls[0]!.sourceLabel).toBe('Function');
|
||||
});
|
||||
|
||||
it('creates IMPORTS edge between app.dart and models.dart', () => {
|
||||
const imports = getRelationships(result, 'IMPORTS');
|
||||
const appImports = imports.filter(
|
||||
(e) => e.sourceFilePath.includes('app.dart') && e.targetFilePath.includes('models.dart'),
|
||||
);
|
||||
expect(appImports.length).toBe(1);
|
||||
});
|
||||
|
||||
// Dart field-chain ACCESSES edges require the call-processor's chain-resolution
|
||||
// tier (Step 1c) to fire. This needs the type-env's scoped parameter binding
|
||||
// (processUser's `user: User`) to propagate to processCallsFromExtracted so
|
||||
// walkMixedChain can resolve User → address → Address and emit ACCESSES.
|
||||
// The chain extraction (extractMixedChain) and member detection
|
||||
// (MEMBER_ACCESS_NODE_TYPES) are wired, but the base receiver type lookup
|
||||
// from the type-env currently returns undefined for Dart function parameters
|
||||
// in the call-processor context. Tracked for follow-up.
|
||||
it.skip('emits ACCESSES edges for field reads in chains', () => {
|
||||
const accesses = getRelationships(result, 'ACCESSES');
|
||||
const addressReads = accesses.filter(
|
||||
(e) => e.target === 'address' && e.rel.reason === 'read',
|
||||
);
|
||||
expect(addressReads.length).toBe(1);
|
||||
expect(addressReads[0]!.source).toBe('processUser');
|
||||
expect(addressReads[0]!.targetLabel).toBe('Property');
|
||||
});
|
||||
});
|
||||
|
||||
// ── Phase 9: Call-result binding ────────────────────────────────────────
|
||||
|
||||
describe.skipIf(!dartAvailable)('Dart call-result binding', () => {
|
||||
let result: PipelineResult;
|
||||
|
||||
beforeAll(async () => {
|
||||
result = await runPipelineFromRepo(
|
||||
path.join(FIXTURES, 'dart-call-result-binding'),
|
||||
() => {},
|
||||
);
|
||||
}, 60000);
|
||||
|
||||
it('detects classes, methods, and functions', () => {
|
||||
expect(getNodesByLabel(result, 'Class')).toContain('User');
|
||||
expect(getNodesByLabel(result, 'Function')).toEqual(
|
||||
expect.arrayContaining(['getUser', 'processUser']),
|
||||
);
|
||||
});
|
||||
|
||||
it('resolves save() call via call-result binding', () => {
|
||||
const calls = getRelationships(result, 'CALLS');
|
||||
// Dart attributes calls to the enclosing Function
|
||||
const saveCalls = calls.filter(
|
||||
(c) => c.target === 'save' && c.sourceFilePath.includes('app.dart'),
|
||||
);
|
||||
expect(saveCalls.length).toBe(1);
|
||||
expect(saveCalls[0]!.targetFilePath).toContain('models.dart');
|
||||
});
|
||||
|
||||
it('resolves getUser() call', () => {
|
||||
const calls = getRelationships(result, 'CALLS');
|
||||
const getUserCalls = calls.filter(
|
||||
(c) => c.target === 'getUser' && c.sourceFilePath.includes('app.dart'),
|
||||
);
|
||||
expect(getUserCalls.length).toBe(1);
|
||||
});
|
||||
|
||||
it('attributes calls to processUser, not File', () => {
|
||||
const calls = getRelationships(result, 'CALLS');
|
||||
const appCalls = calls.filter(
|
||||
(c) => c.sourceFilePath.includes('app.dart'),
|
||||
);
|
||||
for (const call of appCalls) {
|
||||
expect(call.source).toBe('processUser');
|
||||
expect(call.sourceLabel).toBe('Function');
|
||||
}
|
||||
});
|
||||
});
|
||||
|
|
@ -60,11 +60,12 @@ describe('Java heritage resolution', () => {
|
|||
expect(extends_.some(e => e.target === 'Validatable')).toBe(false);
|
||||
});
|
||||
|
||||
it('emits exactly 2 CALLS edges', () => {
|
||||
it('emits exactly 3 CALLS edges', () => {
|
||||
const calls = getRelationships(result, 'CALLS');
|
||||
expect(calls.length).toBe(2);
|
||||
expect(calls.length).toBe(3);
|
||||
expect(edgeSet(calls)).toEqual([
|
||||
'processUser → save',
|
||||
'processUser → serialize',
|
||||
'processUser → validate',
|
||||
]);
|
||||
});
|
||||
|
|
|
|||
|
|
@ -10,13 +10,17 @@
|
|||
import { describe, it, expect, vi, beforeEach } from 'vitest';
|
||||
|
||||
// We need to mock the LadybugDB adapter and repo-manager BEFORE importing LocalBackend
|
||||
vi.mock('../../src/mcp/core/lbug-adapter.js', () => ({
|
||||
initLbug: vi.fn().mockResolvedValue(undefined),
|
||||
executeQuery: vi.fn().mockResolvedValue([]),
|
||||
executeParameterized: vi.fn().mockResolvedValue([]),
|
||||
closeLbug: vi.fn().mockResolvedValue(undefined),
|
||||
isLbugReady: vi.fn().mockReturnValue(true),
|
||||
}));
|
||||
vi.mock('../../src/mcp/core/lbug-adapter.js', async (importOriginal) => {
|
||||
const actual = await importOriginal();
|
||||
return {
|
||||
...actual,
|
||||
initLbug: vi.fn().mockResolvedValue(undefined),
|
||||
executeQuery: vi.fn().mockResolvedValue([]),
|
||||
executeParameterized: vi.fn().mockResolvedValue([]),
|
||||
closeLbug: vi.fn().mockResolvedValue(undefined),
|
||||
isLbugReady: vi.fn().mockReturnValue(true),
|
||||
};
|
||||
});
|
||||
|
||||
vi.mock('../../src/storage/repo-manager.js', () => ({
|
||||
listRegisteredRepos: vi.fn().mockResolvedValue([]),
|
||||
|
|
|
|||
69
gitnexus/test/unit/cobol-copy-expander.test.ts
Normal file
69
gitnexus/test/unit/cobol-copy-expander.test.ts
Normal file
|
|
@ -0,0 +1,69 @@
|
|||
/**
|
||||
* Unit Tests: COBOL Copy Expander — pseudotext REPLACING support
|
||||
*/
|
||||
import { describe, it, expect } from 'vitest';
|
||||
import { parseReplacingClause } from '../../src/core/ingestion/cobol/cobol-copy-expander.js';
|
||||
|
||||
describe('parseReplacingClause', () => {
|
||||
// Existing quoted-string behavior preserved
|
||||
it('parses quoted EXACT replacement', () => {
|
||||
const result = parseReplacingClause(' "OLD-NAME" BY "NEW-NAME" ');
|
||||
expect(result).toEqual([{ type: 'EXACT', from: 'OLD-NAME', to: 'NEW-NAME' }]);
|
||||
});
|
||||
|
||||
it('parses LEADING replacement', () => {
|
||||
const result = parseReplacingClause(' LEADING "ESP-" BY "LK-ESP-" ');
|
||||
expect(result).toEqual([{ type: 'LEADING', from: 'ESP-', to: 'LK-ESP-' }]);
|
||||
});
|
||||
|
||||
it('parses TRAILING replacement', () => {
|
||||
const result = parseReplacingClause(' TRAILING "-IN" BY "-OUT" ');
|
||||
expect(result).toEqual([{ type: 'TRAILING', from: '-IN', to: '-OUT' }]);
|
||||
});
|
||||
|
||||
// Pseudotext ==...== support (isPseudotext flag propagated)
|
||||
it('parses basic pseudotext: ==OLD== BY ==NEW==', () => {
|
||||
const result = parseReplacingClause(' ==WS-OLD== BY ==WS-NEW== ');
|
||||
expect(result).toEqual([{ type: 'EXACT', from: 'WS-OLD', to: 'WS-NEW', isPseudotext: true }]);
|
||||
});
|
||||
|
||||
it('parses empty pseudotext (deletion): ==TEXT== BY ====', () => {
|
||||
const result = parseReplacingClause(' ==REMOVE-ME== BY ==== ');
|
||||
expect(result).toEqual([{ type: 'EXACT', from: 'REMOVE-ME', to: '', isPseudotext: true }]);
|
||||
});
|
||||
|
||||
it('parses pseudotext with spaces: ==SOME TEXT== BY ==OTHER TEXT==', () => {
|
||||
const result = parseReplacingClause(' ==WORKING STORAGE== BY ==LOCAL STORAGE== ');
|
||||
expect(result).toEqual([{ type: 'EXACT', from: 'WORKING STORAGE', to: 'LOCAL STORAGE', isPseudotext: true }]);
|
||||
});
|
||||
|
||||
it('parses pseudotext with single = inside: ==A=B== BY ==C=D==', () => {
|
||||
const result = parseReplacingClause(' ==A=B== BY ==C=D== ');
|
||||
expect(result).toEqual([{ type: 'EXACT', from: 'A=B', to: 'C=D', isPseudotext: true }]);
|
||||
});
|
||||
|
||||
it('parses mixed quoted + pseudotext in one clause', () => {
|
||||
const result = parseReplacingClause(
|
||||
' "OLD-NAME" BY "NEW-NAME" ==DEL-PREFIX== BY ==== ',
|
||||
);
|
||||
expect(result).toEqual([
|
||||
{ type: 'EXACT', from: 'OLD-NAME', to: 'NEW-NAME' },
|
||||
{ type: 'EXACT', from: 'DEL-PREFIX', to: '', isPseudotext: true },
|
||||
]);
|
||||
});
|
||||
|
||||
it('LEADING modifier works alongside pseudotext', () => {
|
||||
const result = parseReplacingClause(
|
||||
' LEADING "ESP-" BY "LK-ESP-" ==OLD-EXACT== BY ==NEW-EXACT== ',
|
||||
);
|
||||
expect(result).toEqual([
|
||||
{ type: 'LEADING', from: 'ESP-', to: 'LK-ESP-' },
|
||||
{ type: 'EXACT', from: 'OLD-EXACT', to: 'NEW-EXACT', isPseudotext: true },
|
||||
]);
|
||||
});
|
||||
|
||||
it('returns empty array for empty input', () => {
|
||||
expect(parseReplacingClause('')).toEqual([]);
|
||||
expect(parseReplacingClause(' ')).toEqual([]);
|
||||
});
|
||||
});
|
||||
2746
gitnexus/test/unit/cobol-preprocessor.test.ts
Normal file
2746
gitnexus/test/unit/cobol-preprocessor.test.ts
Normal file
File diff suppressed because it is too large
Load diff
|
|
@ -186,4 +186,61 @@ describe('createKnowledgeGraph', () => {
|
|||
g.forEachRelationship(r => types.push(r.type));
|
||||
expect(types).toEqual(['CALLS']);
|
||||
});
|
||||
|
||||
// ─── removeRelationship ─────────────────────────────────────────────
|
||||
|
||||
it('removes a relationship by id', () => {
|
||||
const g = createKnowledgeGraph();
|
||||
g.addNode(makeNode('fn:a', 'a'));
|
||||
g.addNode(makeNode('fn:b', 'b'));
|
||||
g.addRelationship(makeRel('fn:a', 'fn:b'));
|
||||
expect(g.relationshipCount).toBe(1);
|
||||
|
||||
const removed = g.removeRelationship('fn:a-CALLS-fn:b');
|
||||
expect(removed).toBe(true);
|
||||
expect(g.relationshipCount).toBe(0);
|
||||
});
|
||||
|
||||
it('removeRelationship returns false for unknown id', () => {
|
||||
const g = createKnowledgeGraph();
|
||||
expect(g.removeRelationship('nonexistent')).toBe(false);
|
||||
});
|
||||
|
||||
it('removeRelationship returns false on second call with same id', () => {
|
||||
const g = createKnowledgeGraph();
|
||||
g.addNode(makeNode('fn:a', 'a'));
|
||||
g.addNode(makeNode('fn:b', 'b'));
|
||||
g.addRelationship(makeRel('fn:a', 'fn:b'));
|
||||
|
||||
expect(g.removeRelationship('fn:a-CALLS-fn:b')).toBe(true);
|
||||
expect(g.removeRelationship('fn:a-CALLS-fn:b')).toBe(false);
|
||||
});
|
||||
|
||||
it('removeRelationship does not affect nodes', () => {
|
||||
const g = createKnowledgeGraph();
|
||||
g.addNode(makeNode('fn:a', 'a'));
|
||||
g.addNode(makeNode('fn:b', 'b'));
|
||||
g.addRelationship(makeRel('fn:a', 'fn:b'));
|
||||
|
||||
g.removeRelationship('fn:a-CALLS-fn:b');
|
||||
expect(g.nodeCount).toBe(2);
|
||||
expect(g.getNode('fn:a')).toBeDefined();
|
||||
expect(g.getNode('fn:b')).toBeDefined();
|
||||
});
|
||||
|
||||
it('removeRelationship leaves other relationships intact', () => {
|
||||
const g = createKnowledgeGraph();
|
||||
g.addNode(makeNode('fn:a', 'a'));
|
||||
g.addNode(makeNode('fn:b', 'b'));
|
||||
g.addNode(makeNode('fn:c', 'c'));
|
||||
g.addRelationship(makeRel('fn:a', 'fn:b'));
|
||||
g.addRelationship(makeRel('fn:b', 'fn:c'));
|
||||
expect(g.relationshipCount).toBe(2);
|
||||
|
||||
g.removeRelationship('fn:a-CALLS-fn:b');
|
||||
expect(g.relationshipCount).toBe(1);
|
||||
const remaining = [...g.iterRelationships()];
|
||||
expect(remaining[0].sourceId).toBe('fn:b');
|
||||
expect(remaining[0].targetId).toBe('fn:c');
|
||||
});
|
||||
});
|
||||
|
|
|
|||
230
gitnexus/test/unit/impact-batching-grouping.test.ts
Normal file
230
gitnexus/test/unit/impact-batching-grouping.test.ts
Normal file
|
|
@ -0,0 +1,230 @@
|
|||
import { describe, it, expect, vi, beforeEach } from 'vitest';
|
||||
|
||||
// Mock the lbug-adapter module before importing LocalBackend so the class
|
||||
// uses the mocked implementations of executeQuery / executeParameterized.
|
||||
const executeQueryMock = vi.fn();
|
||||
const executeParameterizedMock = vi.fn();
|
||||
|
||||
// Use the exact import specifier including .js to match runtime imports
|
||||
vi.mock('../../src/mcp/core/lbug-adapter.js', async (importOriginal) => {
|
||||
const actual = await importOriginal();
|
||||
return {
|
||||
...actual,
|
||||
initLbug: vi.fn(),
|
||||
executeQuery: (...args: any[]) => executeQueryMock(...args),
|
||||
executeParameterized: (...args: any[]) => executeParameterizedMock(...args),
|
||||
closeLbug: vi.fn(),
|
||||
isLbugReady: vi.fn().mockReturnValue(true),
|
||||
};
|
||||
});
|
||||
|
||||
import { LocalBackend } from '../../src/mcp/local/local-backend';
|
||||
|
||||
describe('impact: batching and grouping', () => {
|
||||
beforeEach(() => {
|
||||
vi.clearAllMocks();
|
||||
});
|
||||
|
||||
it('batches 250 IDs into 3 chunked STEP_IN_PROCESS queries', async () => {
|
||||
// Prepare backend and a fake repo handle
|
||||
const backend = new LocalBackend();
|
||||
const repoHandle = {
|
||||
id: 'repo1', name: 'repo1', repoPath: '/tmp/repo', storagePath: '/tmp/repo/.gitnexus',
|
||||
lbugPath: '/tmp/repo/.gitnexus/lbug', indexedAt: 'now', lastCommit: 'c', stats: {},
|
||||
} as any;
|
||||
(backend as any).repos.set(repoHandle.id, repoHandle);
|
||||
(backend as any).ensureInitialized = vi.fn().mockResolvedValue(undefined);
|
||||
|
||||
// executeParameterized: resolve target -> return a symbol row (default)
|
||||
executeParameterizedMock.mockImplementation(async (...args: any[]) => {
|
||||
const query = typeof args[1] === 'string' ? args[1] : String(args[0] ?? '');
|
||||
// The initial target-resolution call will not contain STEP_IN_PROCESS
|
||||
if (!query.includes('STEP_IN_PROCESS')) return [{ id: 'sym1', name: 'Target', filePath: 'f' }];
|
||||
// For STEP_IN_PROCESS calls, fall through to test's executeQueryMock logic by returning [] here.
|
||||
return [];
|
||||
});
|
||||
|
||||
// Track chunk sizes
|
||||
const chunkSizes: number[] = [];
|
||||
let chunkCallIndex = 0;
|
||||
|
||||
executeQueryMock.mockImplementation(async (...args: any[]) => {
|
||||
const query = typeof args[1] === 'string' ? args[1] : String(args[0] ?? '');
|
||||
// Depth traversal query (find related nodes) -- return 250 impacted ids
|
||||
if (query.includes("r.type IN") && !query.includes('STEP_IN_PROCESS')) {
|
||||
const res: any[] = [];
|
||||
for (let i = 0; i < 250; i++) {
|
||||
res.push({ id: `node-${i}`, name: `n${i}`, filePath: `file-${i}.js`, relType: 'CALLS', confidence: null });
|
||||
}
|
||||
return res;
|
||||
}
|
||||
|
||||
// NOTE: process-chunk enrichment previously used executeQuery; our
|
||||
// implementation now calls executeParameterized for those chunks. We
|
||||
// still keep this branch to support any legacy calls, but primary
|
||||
// chunk tracking will be handled via executeParameterizedMock below.
|
||||
|
||||
return [];
|
||||
});
|
||||
|
||||
// Handle parameterized calls (including chunked STEP_IN_PROCESS queries)
|
||||
executeParameterizedMock.mockImplementation(async (...args: any[]) => {
|
||||
const query = typeof args[1] === 'string' ? args[1] : String(args[0] ?? '');
|
||||
const params = args[2] || {};
|
||||
if (query.includes('STEP_IN_PROCESS')) {
|
||||
// Count ids passed in as params.ids
|
||||
const ids = Array.isArray(params.ids) ? params.ids : [];
|
||||
const cnt = ids.length;
|
||||
chunkSizes.push(cnt);
|
||||
const idx = chunkCallIndex++;
|
||||
return [{ entryPointId: `ep-${Math.floor(idx)}`, epName: `epName-${idx}`, epType: 'Function', epFilePath: `/path/${idx}`, hits: cnt, minStep: 1 }];
|
||||
}
|
||||
// Default target resolution
|
||||
return [{ id: 'sym1', name: 'Target', filePath: 'f' }];
|
||||
});
|
||||
|
||||
const params = { target: 'Target', direction: 'downstream', maxDepth: 1 } as any;
|
||||
|
||||
const res = await (backend as any)._impactImpl(repoHandle, params);
|
||||
|
||||
// Expect 3 chunk calls: 100 + 100 + 50
|
||||
expect(chunkSizes.length).toBe(3);
|
||||
const total = chunkSizes.reduce((s, v) => s + v, 0);
|
||||
expect(total).toBe(250);
|
||||
|
||||
// Result impacted count should be 250
|
||||
expect(res.impactedCount).toBe(250);
|
||||
});
|
||||
|
||||
it('groups entry points across chunks and deduplicates correctly', async () => {
|
||||
const backend = new LocalBackend();
|
||||
const repoHandle = {
|
||||
id: 'repo2', name: 'repo2', repoPath: '/tmp/repo2', storagePath: '/tmp/repo2/.gitnexus',
|
||||
lbugPath: '/tmp/repo2/.gitnexus/lbug', indexedAt: 'now', lastCommit: 'c', stats: {},
|
||||
} as any;
|
||||
(backend as any).repos.set(repoHandle.id, repoHandle);
|
||||
(backend as any).ensureInitialized = vi.fn().mockResolvedValue(undefined);
|
||||
|
||||
executeParameterizedMock.mockImplementation(async (...args: any[]) => {
|
||||
const query = typeof args[1] === 'string' ? args[1] : String(args[0] ?? '');
|
||||
if (!query.includes('STEP_IN_PROCESS')) return [{ id: 'symA', name: 'TargetA', filePath: 'f' }];
|
||||
// For STEP_IN_PROCESS in this test, return grouping rows
|
||||
return [
|
||||
{ entryPointId: 'ep-1', epName: 'EP1', epType: 'Function', epFilePath: '/p/1', hits: 2, minStep: 1 },
|
||||
{ entryPointId: 'ep-2', epName: 'EP2', epType: 'Function', epFilePath: '/p/2', hits: 2, minStep: 2 },
|
||||
{ entryPointId: 'ep-1', epName: 'EP1', epType: 'Function', epFilePath: '/p/1', hits: 1, minStep: 3 },
|
||||
{ entryPointId: 'ep-3', epName: 'EP3', epType: 'Function', epFilePath: '/p/3', hits: 1, minStep: 1 },
|
||||
];
|
||||
});
|
||||
|
||||
// Prepare impacted nodes: smaller set for clarity (6 nodes -> chunk size default 100 so single chunk)
|
||||
executeQueryMock.mockImplementation(async (...args: any[]) => {
|
||||
const query = typeof args[1] === 'string' ? args[1] : String(args[0] ?? '');
|
||||
if (query.includes("r.type IN") && !query.includes('STEP_IN_PROCESS')) {
|
||||
// return 6 nodes
|
||||
const res: any[] = [];
|
||||
for (let i = 0; i < 6; i++) res.push({ id: `node-${i}`, name: `n${i}`, filePath: `file-${i}.js`, relType: 'CALLS', confidence: null });
|
||||
return res;
|
||||
}
|
||||
|
||||
return [];
|
||||
});
|
||||
|
||||
const params = { target: 'TargetA', direction: 'downstream', maxDepth: 1 } as any;
|
||||
const res = await (backend as any)._impactImpl(repoHandle, params);
|
||||
|
||||
// affected_processes should be grouped by entryPointId: ep-1, ep-2, ep-3 => 3 unique
|
||||
expect(Array.isArray(res.affected_processes)).toBe(true);
|
||||
const names = res.affected_processes.map((p: any) => p.name);
|
||||
expect(names.sort()).toEqual(['EP1', 'EP2', 'EP3'].sort());
|
||||
|
||||
const ep1 = res.affected_processes.find((p: any) => p.name === 'EP1');
|
||||
expect(ep1.total_hits).toBe(3);
|
||||
|
||||
const ep2 = res.affected_processes.find((p: any) => p.name === 'EP2');
|
||||
expect(ep2.total_hits).toBe(2);
|
||||
});
|
||||
|
||||
it('caps enrichment to MAX_CHUNKS and sets partial when capped', async () => {
|
||||
// Temporarily set MAX_CHUNKS small for deterministic test
|
||||
process.env.IMPACT_MAX_CHUNKS = '3'; // CHUNK_SIZE 100 => maxItems = 300
|
||||
|
||||
const backend = new LocalBackend();
|
||||
const repoHandle = {
|
||||
id: 'repo3', name: 'repo3', repoPath: '/tmp/repo3', storagePath: '/tmp/repo3/.gitnexus',
|
||||
lbugPath: '/tmp/repo3/.gitnexus/lbug', indexedAt: 'now', lastCommit: 'c', stats: {},
|
||||
} as any;
|
||||
(backend as any).repos.set(repoHandle.id, repoHandle);
|
||||
(backend as any).ensureInitialized = vi.fn().mockResolvedValue(undefined);
|
||||
|
||||
// Depth traversal returns 500 impacted nodes
|
||||
executeQueryMock.mockImplementation(async (...args: any[]) => {
|
||||
const query = typeof args[1] === 'string' ? args[1] : String(args[0] ?? '');
|
||||
if (query.includes("r.type IN") && !query.includes('STEP_IN_PROCESS')) {
|
||||
const res: any[] = [];
|
||||
for (let i = 0; i < 500; i++) res.push({ id: `node-${i}`, name: `n${i}`, filePath: `file-${i}.js`, relType: 'CALLS', confidence: null });
|
||||
return res;
|
||||
}
|
||||
return [];
|
||||
});
|
||||
|
||||
const chunkSizes: number[] = [];
|
||||
|
||||
executeParameterizedMock.mockImplementation(async (...args: any[]) => {
|
||||
const query = typeof args[1] === 'string' ? args[1] : String(args[0] ?? '');
|
||||
const params = args[2] || {};
|
||||
if (query.includes('STEP_IN_PROCESS')) {
|
||||
const ids = Array.isArray(params.ids) ? params.ids : [];
|
||||
chunkSizes.push(ids.length);
|
||||
return [{ entryPointId: 'ep-x', epName: 'EPX', epType: 'Function', epFilePath: '/p/x', hits: ids.length, minStep: 1 }];
|
||||
}
|
||||
|
||||
if (query.includes('COUNT(DISTINCT s.id)')) {
|
||||
// moduleQuery: return a module row
|
||||
return [{ name: 'ModuleA', hits: 42 }];
|
||||
}
|
||||
|
||||
if (query.includes('RETURN DISTINCT c.heuristicLabel')) {
|
||||
// directModuleQuery
|
||||
return [{ name: 'ModuleA' }];
|
||||
}
|
||||
|
||||
// Default: target resolution
|
||||
return [{ id: 'symX', name: 'TargetX', filePath: 'f' }];
|
||||
});
|
||||
|
||||
const params = { target: 'TargetX', direction: 'downstream', maxDepth: 1 } as any;
|
||||
const res = await (backend as any)._impactImpl(repoHandle, params);
|
||||
|
||||
// Expect we processed only MAX_CHUNKS chunks (3) -> total ids handled = 300
|
||||
expect(chunkSizes.length).toBe(3);
|
||||
const totalHandled = chunkSizes.reduce((s, v) => s + v, 0);
|
||||
expect(totalHandled).toBe(300);
|
||||
|
||||
// Because we capped enrichment, the result should include partial: true
|
||||
expect(res.partial).toBe(true);
|
||||
|
||||
// Module enrichment should have been called in chunks (3 calls, totaling 300 ids)
|
||||
const memberCalls = (executeParameterizedMock.mock.calls || []).filter((c: any[]) => {
|
||||
const q = typeof c[1] === 'string' ? c[1] : String(c[0] ?? '');
|
||||
// Only count the module-hits query (which returns COUNT(DISTINCT s.id)).
|
||||
// The process-chunk query also uses COUNT(DISTINCT s.id), so require MEMBER_OF
|
||||
// to avoid double-counting process-chunk calls.
|
||||
return q.includes('COUNT(DISTINCT s.id)') && q.includes('MEMBER_OF');
|
||||
});
|
||||
// MAX_CHUNKS = 3 in this test, so expect 3 module-enrichment chunk calls
|
||||
// DEBUG: print memberCalls and their ids lengths
|
||||
expect(memberCalls.length).toBe(3);
|
||||
const totalModuleIds = memberCalls.reduce((sum: number, call: any[]) => sum + ((Array.isArray(call[2]?.ids) ? call[2].ids.length : 0)), 0);
|
||||
// eslint-disable-next-line no-console
|
||||
expect(totalModuleIds).toBe(300);
|
||||
|
||||
// Affected modules should include ModuleA
|
||||
expect(Array.isArray(res.affected_modules)).toBe(true);
|
||||
const modNames = res.affected_modules.map((m: any) => m.name);
|
||||
expect(modNames).toContain('ModuleA');
|
||||
|
||||
// Cleanup env
|
||||
delete process.env.IMPACT_MAX_CHUNKS;
|
||||
});
|
||||
});
|
||||
|
|
@ -1,9 +1,9 @@
|
|||
import { describe, it, expect } from 'vitest';
|
||||
import { getLanguageFromFilename } from '../../src/core/ingestion/utils/language-detection.js';
|
||||
import { isBuiltInOrNoise } from '../../src/core/ingestion/utils/noise-filter.js';
|
||||
import { getProvider } from '../../src/core/ingestion/languages/index.js';
|
||||
import { SupportedLanguages } from '../../src/config/supported-languages.js';
|
||||
import { extractFunctionName } from '../../src/core/ingestion/utils/ast-helpers.js';
|
||||
import { getTreeSitterBufferSize, TREE_SITTER_BUFFER_SIZE, TREE_SITTER_MAX_BUFFER } from '../../src/core/ingestion/constants.js';
|
||||
import { SupportedLanguages } from '../../src/config/supported-languages.js';
|
||||
import Parser from 'tree-sitter';
|
||||
import C from 'tree-sitter-c';
|
||||
import CPP from 'tree-sitter-cpp';
|
||||
|
|
@ -140,203 +140,212 @@ describe('getLanguageFromFilename', () => {
|
|||
});
|
||||
|
||||
describe('isBuiltInOrNoise', () => {
|
||||
const js = getProvider(SupportedLanguages.JavaScript);
|
||||
const py = getProvider(SupportedLanguages.Python);
|
||||
const php = getProvider(SupportedLanguages.PHP);
|
||||
const c = getProvider(SupportedLanguages.C);
|
||||
const kt = getProvider(SupportedLanguages.Kotlin);
|
||||
const swift = getProvider(SupportedLanguages.Swift);
|
||||
const rust = getProvider(SupportedLanguages.Rust);
|
||||
const cs = getProvider(SupportedLanguages.CSharp);
|
||||
|
||||
describe('JavaScript/TypeScript', () => {
|
||||
it('filters console methods', () => {
|
||||
expect(isBuiltInOrNoise('console')).toBe(true);
|
||||
expect(isBuiltInOrNoise('log')).toBe(true);
|
||||
expect(isBuiltInOrNoise('warn')).toBe(true);
|
||||
expect(js.isBuiltInName('console')).toBe(true);
|
||||
expect(js.isBuiltInName('log')).toBe(true);
|
||||
expect(js.isBuiltInName('warn')).toBe(true);
|
||||
});
|
||||
|
||||
it('filters React hooks', () => {
|
||||
expect(isBuiltInOrNoise('useState')).toBe(true);
|
||||
expect(isBuiltInOrNoise('useEffect')).toBe(true);
|
||||
expect(isBuiltInOrNoise('useCallback')).toBe(true);
|
||||
expect(js.isBuiltInName('useState')).toBe(true);
|
||||
expect(js.isBuiltInName('useEffect')).toBe(true);
|
||||
expect(js.isBuiltInName('useCallback')).toBe(true);
|
||||
});
|
||||
|
||||
it('filters array methods', () => {
|
||||
expect(isBuiltInOrNoise('map')).toBe(true);
|
||||
expect(isBuiltInOrNoise('filter')).toBe(true);
|
||||
expect(isBuiltInOrNoise('reduce')).toBe(true);
|
||||
expect(js.isBuiltInName('map')).toBe(true);
|
||||
expect(js.isBuiltInName('filter')).toBe(true);
|
||||
expect(js.isBuiltInName('reduce')).toBe(true);
|
||||
});
|
||||
});
|
||||
|
||||
describe('Python', () => {
|
||||
it('filters built-in functions', () => {
|
||||
expect(isBuiltInOrNoise('print')).toBe(true);
|
||||
expect(isBuiltInOrNoise('len')).toBe(true);
|
||||
expect(isBuiltInOrNoise('range')).toBe(true);
|
||||
expect(py.isBuiltInName('print')).toBe(true);
|
||||
expect(py.isBuiltInName('len')).toBe(true);
|
||||
expect(py.isBuiltInName('range')).toBe(true);
|
||||
});
|
||||
});
|
||||
|
||||
describe('PHP', () => {
|
||||
it('filters PHP built-in functions', () => {
|
||||
expect(isBuiltInOrNoise('echo')).toBe(true);
|
||||
expect(isBuiltInOrNoise('isset')).toBe(true);
|
||||
expect(isBuiltInOrNoise('date')).toBe(true);
|
||||
expect(isBuiltInOrNoise('json_encode')).toBe(true);
|
||||
expect(isBuiltInOrNoise('array_map')).toBe(true);
|
||||
expect(php.isBuiltInName('echo')).toBe(true);
|
||||
expect(php.isBuiltInName('isset')).toBe(true);
|
||||
expect(php.isBuiltInName('date')).toBe(true);
|
||||
expect(php.isBuiltInName('json_encode')).toBe(true);
|
||||
expect(php.isBuiltInName('array_map')).toBe(true);
|
||||
});
|
||||
|
||||
it('filters PHP string functions', () => {
|
||||
expect(isBuiltInOrNoise('strlen')).toBe(true);
|
||||
expect(isBuiltInOrNoise('substr')).toBe(true);
|
||||
expect(isBuiltInOrNoise('str_replace')).toBe(true);
|
||||
expect(php.isBuiltInName('strlen')).toBe(true);
|
||||
expect(php.isBuiltInName('substr')).toBe(true);
|
||||
expect(php.isBuiltInName('str_replace')).toBe(true);
|
||||
});
|
||||
});
|
||||
|
||||
describe('C/C++', () => {
|
||||
it('filters standard library functions', () => {
|
||||
expect(isBuiltInOrNoise('printf')).toBe(true);
|
||||
expect(isBuiltInOrNoise('malloc')).toBe(true);
|
||||
expect(isBuiltInOrNoise('free')).toBe(true);
|
||||
expect(c.isBuiltInName('printf')).toBe(true);
|
||||
expect(c.isBuiltInName('malloc')).toBe(true);
|
||||
expect(c.isBuiltInName('free')).toBe(true);
|
||||
});
|
||||
|
||||
it('filters Linux kernel macros', () => {
|
||||
expect(isBuiltInOrNoise('container_of')).toBe(true);
|
||||
expect(isBuiltInOrNoise('ARRAY_SIZE')).toBe(true);
|
||||
expect(isBuiltInOrNoise('pr_info')).toBe(true);
|
||||
expect(c.isBuiltInName('container_of')).toBe(true);
|
||||
expect(c.isBuiltInName('ARRAY_SIZE')).toBe(true);
|
||||
expect(c.isBuiltInName('pr_info')).toBe(true);
|
||||
});
|
||||
});
|
||||
|
||||
describe('Kotlin', () => {
|
||||
it('filters stdlib functions', () => {
|
||||
expect(isBuiltInOrNoise('println')).toBe(true);
|
||||
expect(isBuiltInOrNoise('listOf')).toBe(true);
|
||||
expect(isBuiltInOrNoise('TODO')).toBe(true);
|
||||
expect(kt.isBuiltInName('println')).toBe(true);
|
||||
expect(kt.isBuiltInName('listOf')).toBe(true);
|
||||
expect(kt.isBuiltInName('TODO')).toBe(true);
|
||||
});
|
||||
|
||||
it('filters coroutine functions', () => {
|
||||
expect(isBuiltInOrNoise('launch')).toBe(true);
|
||||
expect(isBuiltInOrNoise('async')).toBe(true);
|
||||
expect(kt.isBuiltInName('launch')).toBe(true);
|
||||
expect(kt.isBuiltInName('async')).toBe(true);
|
||||
});
|
||||
});
|
||||
|
||||
describe('Swift', () => {
|
||||
it('filters built-in functions', () => {
|
||||
expect(isBuiltInOrNoise('print')).toBe(true);
|
||||
expect(isBuiltInOrNoise('fatalError')).toBe(true);
|
||||
expect(swift.isBuiltInName('print')).toBe(true);
|
||||
expect(swift.isBuiltInName('fatalError')).toBe(true);
|
||||
});
|
||||
|
||||
it('filters UIKit methods', () => {
|
||||
expect(isBuiltInOrNoise('addSubview')).toBe(true);
|
||||
expect(isBuiltInOrNoise('reloadData')).toBe(true);
|
||||
expect(swift.isBuiltInName('addSubview')).toBe(true);
|
||||
expect(swift.isBuiltInName('reloadData')).toBe(true);
|
||||
});
|
||||
});
|
||||
|
||||
describe('Rust', () => {
|
||||
it('filters Result/Option methods', () => {
|
||||
expect(isBuiltInOrNoise('unwrap')).toBe(true);
|
||||
expect(isBuiltInOrNoise('expect')).toBe(true);
|
||||
expect(isBuiltInOrNoise('unwrap_or')).toBe(true);
|
||||
expect(isBuiltInOrNoise('unwrap_or_else')).toBe(true);
|
||||
expect(isBuiltInOrNoise('unwrap_or_default')).toBe(true);
|
||||
expect(isBuiltInOrNoise('ok')).toBe(true);
|
||||
expect(isBuiltInOrNoise('err')).toBe(true);
|
||||
expect(isBuiltInOrNoise('is_ok')).toBe(true);
|
||||
expect(isBuiltInOrNoise('is_err')).toBe(true);
|
||||
expect(isBuiltInOrNoise('map_err')).toBe(true);
|
||||
expect(isBuiltInOrNoise('and_then')).toBe(true);
|
||||
expect(isBuiltInOrNoise('or_else')).toBe(true);
|
||||
expect(rust.isBuiltInName('unwrap')).toBe(true);
|
||||
expect(rust.isBuiltInName('expect')).toBe(true);
|
||||
expect(rust.isBuiltInName('unwrap_or')).toBe(true);
|
||||
expect(rust.isBuiltInName('unwrap_or_else')).toBe(true);
|
||||
expect(rust.isBuiltInName('unwrap_or_default')).toBe(true);
|
||||
expect(rust.isBuiltInName('ok')).toBe(true);
|
||||
expect(rust.isBuiltInName('err')).toBe(true);
|
||||
expect(rust.isBuiltInName('is_ok')).toBe(true);
|
||||
expect(rust.isBuiltInName('is_err')).toBe(true);
|
||||
expect(rust.isBuiltInName('map_err')).toBe(true);
|
||||
expect(rust.isBuiltInName('and_then')).toBe(true);
|
||||
expect(rust.isBuiltInName('or_else')).toBe(true);
|
||||
});
|
||||
|
||||
it('filters trait conversion methods', () => {
|
||||
expect(isBuiltInOrNoise('clone')).toBe(true);
|
||||
expect(isBuiltInOrNoise('to_string')).toBe(true);
|
||||
expect(isBuiltInOrNoise('to_owned')).toBe(true);
|
||||
expect(isBuiltInOrNoise('into')).toBe(true);
|
||||
expect(isBuiltInOrNoise('from')).toBe(true);
|
||||
expect(isBuiltInOrNoise('as_ref')).toBe(true);
|
||||
expect(isBuiltInOrNoise('as_mut')).toBe(true);
|
||||
expect(rust.isBuiltInName('clone')).toBe(true);
|
||||
expect(rust.isBuiltInName('to_string')).toBe(true);
|
||||
expect(rust.isBuiltInName('to_owned')).toBe(true);
|
||||
expect(rust.isBuiltInName('into')).toBe(true);
|
||||
expect(rust.isBuiltInName('from')).toBe(true);
|
||||
expect(rust.isBuiltInName('as_ref')).toBe(true);
|
||||
expect(rust.isBuiltInName('as_mut')).toBe(true);
|
||||
});
|
||||
|
||||
it('filters iterator methods', () => {
|
||||
expect(isBuiltInOrNoise('iter')).toBe(true);
|
||||
expect(isBuiltInOrNoise('into_iter')).toBe(true);
|
||||
expect(isBuiltInOrNoise('collect')).toBe(true);
|
||||
expect(isBuiltInOrNoise('fold')).toBe(true);
|
||||
expect(isBuiltInOrNoise('for_each')).toBe(true);
|
||||
expect(rust.isBuiltInName('iter')).toBe(true);
|
||||
expect(rust.isBuiltInName('into_iter')).toBe(true);
|
||||
expect(rust.isBuiltInName('collect')).toBe(true);
|
||||
expect(rust.isBuiltInName('fold')).toBe(true);
|
||||
expect(rust.isBuiltInName('for_each')).toBe(true);
|
||||
});
|
||||
|
||||
it('filters collection methods', () => {
|
||||
expect(isBuiltInOrNoise('len')).toBe(true);
|
||||
expect(isBuiltInOrNoise('is_empty')).toBe(true);
|
||||
expect(isBuiltInOrNoise('push')).toBe(true);
|
||||
expect(isBuiltInOrNoise('pop')).toBe(true);
|
||||
expect(isBuiltInOrNoise('insert')).toBe(true);
|
||||
expect(isBuiltInOrNoise('remove')).toBe(true);
|
||||
expect(isBuiltInOrNoise('contains')).toBe(true);
|
||||
expect(rust.isBuiltInName('len')).toBe(true);
|
||||
expect(rust.isBuiltInName('is_empty')).toBe(true);
|
||||
expect(rust.isBuiltInName('push')).toBe(true);
|
||||
expect(rust.isBuiltInName('pop')).toBe(true);
|
||||
expect(rust.isBuiltInName('insert')).toBe(true);
|
||||
expect(rust.isBuiltInName('remove')).toBe(true);
|
||||
expect(rust.isBuiltInName('contains')).toBe(true);
|
||||
});
|
||||
|
||||
it('filters macro-like and panic functions', () => {
|
||||
expect(isBuiltInOrNoise('format')).toBe(true);
|
||||
expect(isBuiltInOrNoise('panic')).toBe(true);
|
||||
expect(isBuiltInOrNoise('unreachable')).toBe(true);
|
||||
expect(isBuiltInOrNoise('todo')).toBe(true);
|
||||
expect(isBuiltInOrNoise('unimplemented')).toBe(true);
|
||||
expect(isBuiltInOrNoise('vec')).toBe(true);
|
||||
expect(isBuiltInOrNoise('println')).toBe(true);
|
||||
expect(isBuiltInOrNoise('eprintln')).toBe(true);
|
||||
expect(isBuiltInOrNoise('dbg')).toBe(true);
|
||||
expect(rust.isBuiltInName('format')).toBe(true);
|
||||
expect(rust.isBuiltInName('panic')).toBe(true);
|
||||
expect(rust.isBuiltInName('unreachable')).toBe(true);
|
||||
expect(rust.isBuiltInName('todo')).toBe(true);
|
||||
expect(rust.isBuiltInName('unimplemented')).toBe(true);
|
||||
expect(rust.isBuiltInName('vec')).toBe(true);
|
||||
expect(rust.isBuiltInName('println')).toBe(true);
|
||||
expect(rust.isBuiltInName('eprintln')).toBe(true);
|
||||
expect(rust.isBuiltInName('dbg')).toBe(true);
|
||||
});
|
||||
|
||||
it('filters sync primitives', () => {
|
||||
expect(isBuiltInOrNoise('lock')).toBe(true);
|
||||
expect(isBuiltInOrNoise('try_lock')).toBe(true);
|
||||
expect(isBuiltInOrNoise('spawn')).toBe(true);
|
||||
expect(isBuiltInOrNoise('join')).toBe(true);
|
||||
expect(isBuiltInOrNoise('sleep')).toBe(true);
|
||||
expect(rust.isBuiltInName('lock')).toBe(true);
|
||||
expect(rust.isBuiltInName('try_lock')).toBe(true);
|
||||
expect(rust.isBuiltInName('spawn')).toBe(true);
|
||||
expect(rust.isBuiltInName('join')).toBe(true);
|
||||
expect(rust.isBuiltInName('sleep')).toBe(true);
|
||||
});
|
||||
|
||||
it('filters enum constructors', () => {
|
||||
expect(isBuiltInOrNoise('Some')).toBe(true);
|
||||
expect(isBuiltInOrNoise('None')).toBe(true);
|
||||
expect(isBuiltInOrNoise('Ok')).toBe(true);
|
||||
expect(isBuiltInOrNoise('Err')).toBe(true);
|
||||
expect(rust.isBuiltInName('Some')).toBe(true);
|
||||
expect(rust.isBuiltInName('None')).toBe(true);
|
||||
expect(rust.isBuiltInName('Ok')).toBe(true);
|
||||
expect(rust.isBuiltInName('Err')).toBe(true);
|
||||
});
|
||||
|
||||
it('does not filter user-defined Rust functions', () => {
|
||||
expect(isBuiltInOrNoise('process_request')).toBe(false);
|
||||
expect(isBuiltInOrNoise('handle_connection')).toBe(false);
|
||||
expect(isBuiltInOrNoise('build_response')).toBe(false);
|
||||
expect(rust.isBuiltInName('process_request')).toBe(false);
|
||||
expect(rust.isBuiltInName('handle_connection')).toBe(false);
|
||||
expect(rust.isBuiltInName('build_response')).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe('C#/.NET', () => {
|
||||
it('filters Console I/O', () => {
|
||||
expect(isBuiltInOrNoise('Console')).toBe(true);
|
||||
expect(isBuiltInOrNoise('WriteLine')).toBe(true);
|
||||
expect(isBuiltInOrNoise('ReadLine')).toBe(true);
|
||||
expect(cs.isBuiltInName('Console')).toBe(true);
|
||||
expect(cs.isBuiltInName('WriteLine')).toBe(true);
|
||||
expect(cs.isBuiltInName('ReadLine')).toBe(true);
|
||||
});
|
||||
|
||||
it('filters LINQ methods', () => {
|
||||
expect(isBuiltInOrNoise('Where')).toBe(true);
|
||||
expect(isBuiltInOrNoise('Select')).toBe(true);
|
||||
expect(isBuiltInOrNoise('GroupBy')).toBe(true);
|
||||
expect(isBuiltInOrNoise('OrderBy')).toBe(true);
|
||||
expect(isBuiltInOrNoise('FirstOrDefault')).toBe(true);
|
||||
expect(isBuiltInOrNoise('ToList')).toBe(true);
|
||||
expect(cs.isBuiltInName('Where')).toBe(true);
|
||||
expect(cs.isBuiltInName('Select')).toBe(true);
|
||||
expect(cs.isBuiltInName('GroupBy')).toBe(true);
|
||||
expect(cs.isBuiltInName('OrderBy')).toBe(true);
|
||||
expect(cs.isBuiltInName('FirstOrDefault')).toBe(true);
|
||||
expect(cs.isBuiltInName('ToList')).toBe(true);
|
||||
});
|
||||
|
||||
it('filters Task async methods', () => {
|
||||
expect(isBuiltInOrNoise('Task')).toBe(true);
|
||||
expect(isBuiltInOrNoise('Run')).toBe(true);
|
||||
expect(isBuiltInOrNoise('WhenAll')).toBe(true);
|
||||
expect(isBuiltInOrNoise('ConfigureAwait')).toBe(true);
|
||||
expect(cs.isBuiltInName('Task')).toBe(true);
|
||||
expect(cs.isBuiltInName('Run')).toBe(true);
|
||||
expect(cs.isBuiltInName('WhenAll')).toBe(true);
|
||||
expect(cs.isBuiltInName('ConfigureAwait')).toBe(true);
|
||||
});
|
||||
|
||||
it('filters Object base methods', () => {
|
||||
expect(isBuiltInOrNoise('ToString')).toBe(true);
|
||||
expect(isBuiltInOrNoise('GetType')).toBe(true);
|
||||
expect(isBuiltInOrNoise('Equals')).toBe(true);
|
||||
expect(isBuiltInOrNoise('GetHashCode')).toBe(true);
|
||||
expect(cs.isBuiltInName('ToString')).toBe(true);
|
||||
expect(cs.isBuiltInName('GetType')).toBe(true);
|
||||
expect(cs.isBuiltInName('Equals')).toBe(true);
|
||||
expect(cs.isBuiltInName('GetHashCode')).toBe(true);
|
||||
});
|
||||
});
|
||||
|
||||
describe('user-defined functions', () => {
|
||||
it('does not filter custom function names', () => {
|
||||
expect(isBuiltInOrNoise('myCustomFunction')).toBe(false);
|
||||
expect(isBuiltInOrNoise('processData')).toBe(false);
|
||||
expect(isBuiltInOrNoise('handleUserRequest')).toBe(false);
|
||||
expect(js.isBuiltInName('myCustomFunction')).toBe(false);
|
||||
expect(py.isBuiltInName('processData')).toBe(false);
|
||||
expect(rust.isBuiltInName('handleUserRequest')).toBe(false);
|
||||
});
|
||||
});
|
||||
});
|
||||
|
|
|
|||
52
gitnexus/test/unit/isWriteQuery.test.ts
Normal file
52
gitnexus/test/unit/isWriteQuery.test.ts
Normal file
|
|
@ -0,0 +1,52 @@
|
|||
// ...existing code...
|
||||
import { describe, it, expect } from 'vitest';
|
||||
import { isWriteQuery as isWriteQueryAdapter } from '../../src/mcp/core/lbug-adapter';
|
||||
import { isWriteQuery as isWriteQueryBackend } from '../../src/mcp/local/local-backend';
|
||||
|
||||
describe('isWriteQuery regex tests', () => {
|
||||
const writeQueries = [
|
||||
'CREATE (n:Test {name: "x"})',
|
||||
'MATCH (n) SET n.x = 1',
|
||||
'MERGE (n:Foo {id: 1})',
|
||||
'DELETE n',
|
||||
'DROP INDEX ON :Foo(prop)',
|
||||
'ALTER TABLE Something',
|
||||
'COPY TO something',
|
||||
'DETACH DELETE n',
|
||||
];
|
||||
|
||||
const readQueries = [
|
||||
'MATCH (n:CreateHelpers) RETURN n',
|
||||
'MATCH (a)-[:CALLS]->(b) RETURN a, b',
|
||||
'MATCH (f:File)-[r:DEFINES]->(n) RETURN n',
|
||||
"MATCH (n) WHERE n.name = 'MERGEHelper' RETURN n", // word present as data
|
||||
'MATCH (n) RETURN n',
|
||||
'MATCH (n) WHERE n.content CONTAINS ":CREATE" RETURN n',
|
||||
'MATCH (n:SomethingWithSET) RETURN n',
|
||||
];
|
||||
|
||||
it('adapter isWriteQuery should detect real write queries', () => {
|
||||
for (const q of writeQueries) {
|
||||
expect(isWriteQueryAdapter(q), `adapter should detect write for: ${q}`).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
it('adapter isWriteQuery should not false-positive on label/rel or data', () => {
|
||||
for (const q of readQueries) {
|
||||
expect(isWriteQueryAdapter(q), `adapter false-positive on: ${q}`).toBe(false);
|
||||
}
|
||||
});
|
||||
|
||||
it('backend isWriteQuery should detect real write queries', () => {
|
||||
for (const q of writeQueries) {
|
||||
expect(isWriteQueryBackend(q), `backend should detect write for: ${q}`).toBe(true);
|
||||
}
|
||||
});
|
||||
|
||||
it('backend isWriteQuery should not false-positive on label/rel or data', () => {
|
||||
for (const q of readQueries) {
|
||||
expect(isWriteQueryBackend(q), `backend false-positive on: ${q}`).toBe(false);
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
338
gitnexus/test/unit/jcl-parser.test.ts
Normal file
338
gitnexus/test/unit/jcl-parser.test.ts
Normal file
|
|
@ -0,0 +1,338 @@
|
|||
import { describe, it, expect } from 'vitest';
|
||||
import { parseJcl } from '../../src/core/ingestion/cobol/jcl-parser.js';
|
||||
import type { JclParseResults } from '../../src/core/ingestion/cobol/jcl-parser.js';
|
||||
|
||||
describe('parseJcl', () => {
|
||||
// ── JOB statements ──────────────────────────────────────────────────
|
||||
|
||||
describe('JOB statements', () => {
|
||||
it('extracts job name', () => {
|
||||
const jcl = `//MYJOB JOB (ACCT),'MY JOB'`;
|
||||
const r = parseJcl(jcl, 'test.jcl');
|
||||
expect(r.jobs).toHaveLength(1);
|
||||
expect(r.jobs[0].name).toBe('MYJOB');
|
||||
expect(r.jobs[0].line).toBe(1);
|
||||
});
|
||||
|
||||
it('extracts CLASS and MSGCLASS parameters', () => {
|
||||
const jcl = `//PAYJOB JOB (ACCT),'PAYROLL',CLASS=A,MSGCLASS=X`;
|
||||
const r = parseJcl(jcl, 'test.jcl');
|
||||
expect(r.jobs).toHaveLength(1);
|
||||
expect(r.jobs[0].name).toBe('PAYJOB');
|
||||
expect(r.jobs[0].class).toBe('A');
|
||||
expect(r.jobs[0].msgclass).toBe('X');
|
||||
});
|
||||
|
||||
it('handles job with no CLASS or MSGCLASS', () => {
|
||||
const jcl = `//BAREJOB JOB (ACCT),'BARE'`;
|
||||
const r = parseJcl(jcl, 'test.jcl');
|
||||
expect(r.jobs).toHaveLength(1);
|
||||
expect(r.jobs[0].name).toBe('BAREJOB');
|
||||
expect(r.jobs[0].class).toBeUndefined();
|
||||
expect(r.jobs[0].msgclass).toBeUndefined();
|
||||
});
|
||||
});
|
||||
|
||||
// ── EXEC statements ─────────────────────────────────────────────────
|
||||
|
||||
describe('EXEC statements', () => {
|
||||
it('extracts step with PGM=program', () => {
|
||||
const jcl = [
|
||||
'//MYJOB JOB (ACCT)',
|
||||
'//STEP1 EXEC PGM=IEFBR14',
|
||||
].join('\n');
|
||||
const r = parseJcl(jcl, 'test.jcl');
|
||||
expect(r.steps).toHaveLength(1);
|
||||
expect(r.steps[0].name).toBe('STEP1');
|
||||
expect(r.steps[0].program).toBe('IEFBR14');
|
||||
expect(r.steps[0].proc).toBeUndefined();
|
||||
});
|
||||
|
||||
it('extracts step with proc name (no PGM= keyword)', () => {
|
||||
const jcl = [
|
||||
'//MYJOB JOB (ACCT)',
|
||||
'//STEP1 EXEC MYPROC',
|
||||
].join('\n');
|
||||
const r = parseJcl(jcl, 'test.jcl');
|
||||
expect(r.steps).toHaveLength(1);
|
||||
expect(r.steps[0].name).toBe('STEP1');
|
||||
expect(r.steps[0].program).toBeUndefined();
|
||||
expect(r.steps[0].proc).toBe('MYPROC');
|
||||
});
|
||||
|
||||
it('associates step with current job', () => {
|
||||
const jcl = [
|
||||
'//JOB1 JOB (ACCT)',
|
||||
'//STEPA EXEC PGM=PROG1',
|
||||
'//JOB2 JOB (ACCT)',
|
||||
'//STEPB EXEC PGM=PROG2',
|
||||
].join('\n');
|
||||
const r = parseJcl(jcl, 'test.jcl');
|
||||
expect(r.steps).toHaveLength(2);
|
||||
expect(r.steps[0].jobName).toBe('JOB1');
|
||||
expect(r.steps[1].jobName).toBe('JOB2');
|
||||
});
|
||||
});
|
||||
|
||||
// ── DD statements ───────────────────────────────────────────────────
|
||||
|
||||
describe('DD statements', () => {
|
||||
it('extracts DD name and dataset (DSN=)', () => {
|
||||
const jcl = [
|
||||
'//MYJOB JOB (ACCT)',
|
||||
'//STEP1 EXEC PGM=IEFBR14',
|
||||
'//INPUT DD DSN=MY.DATA.SET,DISP=SHR',
|
||||
].join('\n');
|
||||
const r = parseJcl(jcl, 'test.jcl');
|
||||
expect(r.ddStatements).toHaveLength(1);
|
||||
expect(r.ddStatements[0].ddName).toBe('INPUT');
|
||||
expect(r.ddStatements[0].dataset).toBe('MY.DATA.SET');
|
||||
});
|
||||
|
||||
it('extracts DISP parameter', () => {
|
||||
const jcl = [
|
||||
'//MYJOB JOB (ACCT)',
|
||||
'//STEP1 EXEC PGM=IEFBR14',
|
||||
'//OUTPUT DD DSN=MY.OUT,DISP=(NEW,CATLG,DELETE)',
|
||||
].join('\n');
|
||||
const r = parseJcl(jcl, 'test.jcl');
|
||||
expect(r.ddStatements).toHaveLength(1);
|
||||
expect(r.ddStatements[0].disp).toBe('NEW');
|
||||
});
|
||||
|
||||
it('associates DD with current step', () => {
|
||||
const jcl = [
|
||||
'//MYJOB JOB (ACCT)',
|
||||
'//STEP1 EXEC PGM=PROG1',
|
||||
'//DD1 DD DSN=DS1,DISP=SHR',
|
||||
'//STEP2 EXEC PGM=PROG2',
|
||||
'//DD2 DD DSN=DS2,DISP=SHR',
|
||||
].join('\n');
|
||||
const r = parseJcl(jcl, 'test.jcl');
|
||||
expect(r.ddStatements).toHaveLength(2);
|
||||
expect(r.ddStatements[0].stepName).toBe('STEP1');
|
||||
expect(r.ddStatements[1].stepName).toBe('STEP2');
|
||||
});
|
||||
});
|
||||
|
||||
// ── PROC definitions ────────────────────────────────────────────────
|
||||
|
||||
describe('PROC definitions', () => {
|
||||
it('extracts in-stream PROC with name', () => {
|
||||
const jcl = [
|
||||
'//MYPROC PROC',
|
||||
'//STEP1 EXEC PGM=IEFBR14',
|
||||
'// PEND',
|
||||
].join('\n');
|
||||
const r = parseJcl(jcl, 'test.jcl');
|
||||
expect(r.procs).toHaveLength(1);
|
||||
expect(r.procs[0].name).toBe('MYPROC');
|
||||
expect(r.procs[0].isInStream).toBe(true);
|
||||
});
|
||||
|
||||
it('handles PROC/PEND pairs', () => {
|
||||
const jcl = [
|
||||
'//PROC1 PROC',
|
||||
'//S1 EXEC PGM=PROG1',
|
||||
'// PEND',
|
||||
'//PROC2 PROC',
|
||||
'//S2 EXEC PGM=PROG2',
|
||||
'// PEND',
|
||||
].join('\n');
|
||||
const r = parseJcl(jcl, 'test.jcl');
|
||||
expect(r.procs).toHaveLength(2);
|
||||
expect(r.procs[0].name).toBe('PROC1');
|
||||
expect(r.procs[1].name).toBe('PROC2');
|
||||
});
|
||||
});
|
||||
|
||||
// ── INCLUDE / SET ───────────────────────────────────────────────────
|
||||
|
||||
describe('INCLUDE and SET', () => {
|
||||
it('extracts INCLUDE MEMBER=name', () => {
|
||||
const jcl = `// INCLUDE MEMBER=MYINCL`;
|
||||
const r = parseJcl(jcl, 'test.jcl');
|
||||
expect(r.includes).toHaveLength(1);
|
||||
expect(r.includes[0].member).toBe('MYINCL');
|
||||
expect(r.includes[0].line).toBe(1);
|
||||
});
|
||||
|
||||
it('extracts SET variable=value', () => {
|
||||
const jcl = `// SET ENV=PROD`;
|
||||
const r = parseJcl(jcl, 'test.jcl');
|
||||
expect(r.sets).toHaveLength(1);
|
||||
expect(r.sets[0].variable).toBe('ENV');
|
||||
expect(r.sets[0].value).toBe('PROD');
|
||||
});
|
||||
});
|
||||
|
||||
// ── Conditionals ────────────────────────────────────────────────────
|
||||
|
||||
describe('Conditionals', () => {
|
||||
it('extracts IF condition THEN', () => {
|
||||
const jcl = `// IF STEP1.RC = 0 THEN`;
|
||||
const r = parseJcl(jcl, 'test.jcl');
|
||||
expect(r.conditionals).toHaveLength(1);
|
||||
expect(r.conditionals[0].type).toBe('IF');
|
||||
expect(r.conditionals[0].condition).toBe('STEP1.RC = 0');
|
||||
});
|
||||
|
||||
it('extracts ELSE and ENDIF', () => {
|
||||
const jcl = [
|
||||
'// IF STEP1.RC = 0 THEN',
|
||||
'//GOOD EXEC PGM=GOODPGM',
|
||||
'// ELSE',
|
||||
'//BAD EXEC PGM=BADPGM',
|
||||
'// ENDIF',
|
||||
].join('\n');
|
||||
const r = parseJcl(jcl, 'test.jcl');
|
||||
expect(r.conditionals).toHaveLength(3);
|
||||
expect(r.conditionals[0].type).toBe('IF');
|
||||
expect(r.conditionals[1].type).toBe('ELSE');
|
||||
expect(r.conditionals[1].condition).toBeUndefined();
|
||||
expect(r.conditionals[2].type).toBe('ENDIF');
|
||||
expect(r.conditionals[2].condition).toBeUndefined();
|
||||
});
|
||||
});
|
||||
|
||||
// ── JCLLIB ──────────────────────────────────────────────────────────
|
||||
|
||||
describe('JCLLIB', () => {
|
||||
it('extracts JCLLIB ORDER=(lib1,lib2)', () => {
|
||||
const jcl = `// JCLLIB ORDER=(SYS1.PROCLIB,USER.PROCLIB)`;
|
||||
const r = parseJcl(jcl, 'test.jcl');
|
||||
expect(r.jcllib).toHaveLength(1);
|
||||
expect(r.jcllib[0].order).toEqual(['SYS1.PROCLIB', 'USER.PROCLIB']);
|
||||
expect(r.jcllib[0].line).toBe(1);
|
||||
});
|
||||
});
|
||||
|
||||
// ── Continuation lines ──────────────────────────────────────────────
|
||||
|
||||
describe('Continuation lines', () => {
|
||||
it('joins continuation lines (col 72 non-blank + next line starts with //)', () => {
|
||||
// Build a DD line that is exactly 72 chars with non-blank at col 72 (index 71).
|
||||
// The continuation line provides the DISP parameter.
|
||||
// "//DD1 DD DSN=MY.VERY.LONG.DATASET.NAME.THAT.KEEPS.GOING," is 60 chars.
|
||||
// Pad to 71 then add non-blank at col 72.
|
||||
const base = '//DD1 DD DSN=MY.VERY.LONG.DATASET.NAME.THAT.KEEPS.GOING,';
|
||||
const padding = ' '.repeat(71 - base.length);
|
||||
const line1 = base + padding + 'X'; // col 72 is 'X' (non-blank) -> continuation
|
||||
const line2 = '// DISP=SHR';
|
||||
const jcl = [
|
||||
'//MYJOB JOB (ACCT)',
|
||||
'//STEP1 EXEC PGM=IEFBR14',
|
||||
line1,
|
||||
line2,
|
||||
].join('\n');
|
||||
const r = parseJcl(jcl, 'test.jcl');
|
||||
// The continuation should join the DD line so both DSN and DISP are parsed
|
||||
expect(r.ddStatements).toHaveLength(1);
|
||||
expect(r.ddStatements[0].ddName).toBe('DD1');
|
||||
expect(r.ddStatements[0].dataset).toBe('MY.VERY.LONG.DATASET.NAME.THAT.KEEPS.GOING');
|
||||
expect(r.ddStatements[0].disp).toBe('SHR');
|
||||
});
|
||||
});
|
||||
|
||||
// ── Edge cases ──────────────────────────────────────────────────────
|
||||
|
||||
describe('Edge cases', () => {
|
||||
it('skips JCL comments (//*)', () => {
|
||||
const jcl = [
|
||||
'//* This is a comment',
|
||||
'//MYJOB JOB (ACCT)',
|
||||
'//* Another comment',
|
||||
].join('\n');
|
||||
const r = parseJcl(jcl, 'test.jcl');
|
||||
expect(r.jobs).toHaveLength(1);
|
||||
expect(r.jobs[0].name).toBe('MYJOB');
|
||||
});
|
||||
|
||||
it('skips non-JCL lines', () => {
|
||||
const jcl = [
|
||||
'This is not a JCL line',
|
||||
'//MYJOB JOB (ACCT)',
|
||||
' Some data',
|
||||
'//STEP1 EXEC PGM=IEFBR14',
|
||||
].join('\n');
|
||||
const r = parseJcl(jcl, 'test.jcl');
|
||||
expect(r.jobs).toHaveLength(1);
|
||||
expect(r.steps).toHaveLength(1);
|
||||
});
|
||||
|
||||
it('empty input returns empty results', () => {
|
||||
const r = parseJcl('', 'test.jcl');
|
||||
expect(r.jobs).toEqual([]);
|
||||
expect(r.steps).toEqual([]);
|
||||
expect(r.ddStatements).toEqual([]);
|
||||
expect(r.procs).toEqual([]);
|
||||
expect(r.includes).toEqual([]);
|
||||
expect(r.sets).toEqual([]);
|
||||
expect(r.jcllib).toEqual([]);
|
||||
expect(r.conditionals).toEqual([]);
|
||||
});
|
||||
|
||||
it('complete JCL job with multiple steps and DDs', () => {
|
||||
const jcl = [
|
||||
'//* Complete payroll job',
|
||||
'//PAYJOB JOB (ACCT123),\'PAYROLL RUN\',CLASS=A,MSGCLASS=X',
|
||||
'// JCLLIB ORDER=(PAY.PROCLIB,SYS1.PROCLIB)',
|
||||
'// SET ENV=PROD',
|
||||
'// INCLUDE MEMBER=STDPARMS',
|
||||
'//*',
|
||||
'// IF 1 = 1 THEN',
|
||||
'//STEP01 EXEC PGM=PAYEXT',
|
||||
'//INPUT DD DSN=PAY.MASTER,DISP=SHR',
|
||||
'//OUTPUT DD DSN=PAY.EXTRACT,DISP=(NEW,CATLG,DELETE)',
|
||||
'//SYSPRINT DD SYSOUT=*',
|
||||
'//*',
|
||||
'//STEP02 EXEC PAYCALC',
|
||||
'//INFILE DD DSN=PAY.EXTRACT,DISP=SHR',
|
||||
'// ELSE',
|
||||
'//STEP03 EXEC PGM=IEFBR14',
|
||||
'// ENDIF',
|
||||
].join('\n');
|
||||
const r = parseJcl(jcl, 'payroll.jcl');
|
||||
|
||||
// Jobs
|
||||
expect(r.jobs).toHaveLength(1);
|
||||
expect(r.jobs[0]).toEqual({
|
||||
name: 'PAYJOB',
|
||||
line: 2,
|
||||
class: 'A',
|
||||
msgclass: 'X',
|
||||
});
|
||||
|
||||
// JCLLIB
|
||||
expect(r.jcllib).toHaveLength(1);
|
||||
expect(r.jcllib[0].order).toEqual(['PAY.PROCLIB', 'SYS1.PROCLIB']);
|
||||
|
||||
// SET
|
||||
expect(r.sets).toHaveLength(1);
|
||||
expect(r.sets[0]).toEqual({ variable: 'ENV', value: 'PROD', line: 4 });
|
||||
|
||||
// INCLUDE
|
||||
expect(r.includes).toHaveLength(1);
|
||||
expect(r.includes[0].member).toBe('STDPARMS');
|
||||
|
||||
// Conditionals
|
||||
expect(r.conditionals).toHaveLength(3);
|
||||
expect(r.conditionals[0].type).toBe('IF');
|
||||
expect(r.conditionals[1].type).toBe('ELSE');
|
||||
expect(r.conditionals[2].type).toBe('ENDIF');
|
||||
|
||||
// Steps
|
||||
expect(r.steps).toHaveLength(3);
|
||||
expect(r.steps[0]).toMatchObject({ name: 'STEP01', program: 'PAYEXT', jobName: 'PAYJOB' });
|
||||
expect(r.steps[1]).toMatchObject({ name: 'STEP02', proc: 'PAYCALC', jobName: 'PAYJOB' });
|
||||
expect(r.steps[2]).toMatchObject({ name: 'STEP03', program: 'IEFBR14', jobName: 'PAYJOB' });
|
||||
|
||||
// DD statements
|
||||
expect(r.ddStatements).toHaveLength(4);
|
||||
expect(r.ddStatements[0]).toMatchObject({ ddName: 'INPUT', stepName: 'STEP01', dataset: 'PAY.MASTER', disp: 'SHR' });
|
||||
expect(r.ddStatements[1]).toMatchObject({ ddName: 'OUTPUT', stepName: 'STEP01', disp: 'NEW' });
|
||||
expect(r.ddStatements[2]).toMatchObject({ ddName: 'SYSPRINT', stepName: 'STEP01' });
|
||||
expect(r.ddStatements[3]).toMatchObject({ ddName: 'INFILE', stepName: 'STEP02', dataset: 'PAY.EXTRACT' });
|
||||
});
|
||||
});
|
||||
});
|
||||
92
gitnexus/test/unit/noise-filter.test.ts
Normal file
92
gitnexus/test/unit/noise-filter.test.ts
Normal file
|
|
@ -0,0 +1,92 @@
|
|||
import { describe, it, expect } from 'vitest';
|
||||
import { getProvider } from '../../src/core/ingestion/languages/index.js';
|
||||
import { SupportedLanguages } from '../../src/config/supported-languages.js';
|
||||
|
||||
const isBuiltIn = (name: string, lang: SupportedLanguages) => getProvider(lang).isBuiltInName(name);
|
||||
|
||||
describe('isBuiltInOrNoise (per-language)', () => {
|
||||
describe('language-specific filtering', () => {
|
||||
it('filters console for JS but not Python', () => {
|
||||
expect(isBuiltIn('console', SupportedLanguages.JavaScript)).toBe(true);
|
||||
expect(isBuiltIn('console', SupportedLanguages.Python)).toBe(false);
|
||||
});
|
||||
|
||||
it('filters println for Kotlin but not Java', () => {
|
||||
expect(isBuiltIn('println', SupportedLanguages.Kotlin)).toBe(true);
|
||||
expect(isBuiltIn('println', SupportedLanguages.Java)).toBe(false);
|
||||
});
|
||||
|
||||
it('filters malloc for C but not JavaScript', () => {
|
||||
expect(isBuiltIn('malloc', SupportedLanguages.C)).toBe(true);
|
||||
expect(isBuiltIn('malloc', SupportedLanguages.JavaScript)).toBe(false);
|
||||
});
|
||||
|
||||
it('filters setState for Dart but not TypeScript', () => {
|
||||
expect(isBuiltIn('setState', SupportedLanguages.Dart)).toBe(true);
|
||||
expect(isBuiltIn('setState', SupportedLanguages.TypeScript)).toBe(false);
|
||||
});
|
||||
|
||||
it('filters unwrap for Rust but not Go', () => {
|
||||
expect(isBuiltIn('unwrap', SupportedLanguages.Rust)).toBe(true);
|
||||
expect(isBuiltIn('unwrap', SupportedLanguages.Go)).toBe(false);
|
||||
});
|
||||
|
||||
it('filters puts for Ruby but not PHP', () => {
|
||||
expect(isBuiltIn('puts', SupportedLanguages.Ruby)).toBe(true);
|
||||
expect(isBuiltIn('puts', SupportedLanguages.PHP)).toBe(false);
|
||||
});
|
||||
|
||||
it('filters echo for PHP but not Python', () => {
|
||||
expect(isBuiltIn('echo', SupportedLanguages.PHP)).toBe(true);
|
||||
expect(isBuiltIn('echo', SupportedLanguages.Python)).toBe(false);
|
||||
});
|
||||
|
||||
it('filters NSLog for Swift but not C', () => {
|
||||
expect(isBuiltIn('NSLog', SupportedLanguages.Swift)).toBe(true);
|
||||
expect(isBuiltIn('NSLog', SupportedLanguages.C)).toBe(false);
|
||||
});
|
||||
|
||||
it('filters ToString for C# but not Rust', () => {
|
||||
expect(isBuiltIn('ToString', SupportedLanguages.CSharp)).toBe(true);
|
||||
expect(isBuiltIn('ToString', SupportedLanguages.Rust)).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe('cross-language pollution eliminated', () => {
|
||||
it('close is filtered for C# but not C (POSIX)', () => {
|
||||
expect(isBuiltIn('Close', SupportedLanguages.CSharp)).toBe(true);
|
||||
expect(isBuiltIn('close', SupportedLanguages.C)).toBe(false);
|
||||
});
|
||||
|
||||
it('then/catch are JS-specific, not filtered for Rust', () => {
|
||||
expect(isBuiltIn('then', SupportedLanguages.JavaScript)).toBe(true);
|
||||
expect(isBuiltIn('catch', SupportedLanguages.JavaScript)).toBe(true);
|
||||
expect(isBuiltIn('then', SupportedLanguages.Rust)).toBe(false);
|
||||
});
|
||||
|
||||
it('emit is Kotlin-specific, not filtered for Java', () => {
|
||||
expect(isBuiltIn('emit', SupportedLanguages.Kotlin)).toBe(true);
|
||||
expect(isBuiltIn('emit', SupportedLanguages.Java)).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe('languages without builtInNames', () => {
|
||||
it('Java has no language-specific noise', () => {
|
||||
expect(isBuiltIn('System', SupportedLanguages.Java)).toBe(false);
|
||||
expect(isBuiltIn('println', SupportedLanguages.Java)).toBe(false);
|
||||
});
|
||||
|
||||
it('Go has no language-specific noise', () => {
|
||||
expect(isBuiltIn('fmt', SupportedLanguages.Go)).toBe(false);
|
||||
expect(isBuiltIn('Println', SupportedLanguages.Go)).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe('domain names not filtered', () => {
|
||||
it('does not filter arbitrary names', () => {
|
||||
expect(isBuiltIn('processOrder', SupportedLanguages.TypeScript)).toBe(false);
|
||||
expect(isBuiltIn('UserService', SupportedLanguages.Java)).toBe(false);
|
||||
expect(isBuiltIn('handle_request', SupportedLanguages.Rust)).toBe(false);
|
||||
});
|
||||
});
|
||||
});
|
||||
|
|
@ -375,27 +375,27 @@ So return-type-aware receiver inference already exists in a constrained downstre
|
|||
|
||||
## Language Feature Matrix
|
||||
|
||||
| Feature | TS | JS | Java | Kotlin | C# | Go | Rust | Python | PHP | Ruby | Swift | C++ | C |
|
||||
|---------|:--:|:--:|:----:|:------:|:--:|:--:|:----:|:------:|:---:|:----:|:-----:|:---:|:-:|
|
||||
| Declarations | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
|
||||
| Parameters | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
|
||||
| Initializer / constructor inference | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
|
||||
| Constructor binding scan | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
|
||||
| For-loop element types | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes††† | Yes | Yes |
|
||||
| Pattern binding | Yes | Yes | Yes | Yes | No | Yes | Yes | No | No | No | Partial‡‡‡ | No | No |
|
||||
| Assignment chains | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | No | Yes | Yes | Yes |
|
||||
| Field/property type resolution | Yes | No† | Yes | Yes | Yes | Yes | Yes | Yes* | Yes | YARD | No | Yes | No‡ |
|
||||
| Comment-based types | JSDoc | JSDoc | No | No | No | No | No | No | PHPDoc | YARD | No | No | No |
|
||||
| Return type extraction | JSDoc | JSDoc | No | No | No | No | No | No | PHPDoc | YARD | No | No | No |
|
||||
| Call-result variable binding | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes¶ | Yes††† | Yes | No |
|
||||
| Field access binding | Yes | No† | Yes | Yes | Yes | Yes | Yes | No‖ | Yes | N/A | Yes††† | Yes | No |
|
||||
| Method-call-result binding | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes¶ | Yes††† | Yes | No |
|
||||
| Write access (ACCESSES write) | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes§ | Yes | Yes | Yes | No |
|
||||
| Parameter types extracted | Yes** | No | Yes | Yes | Yes | Yes | Yes | Partial†† | No | No | No | Yes | No |
|
||||
| Method overload disambiguation | Yes** | No | Yes | Yes | Yes | No | No | No | No | No | No | Yes | No |
|
||||
| Constructor-visible virtual dispatch | Yes | No | Yes | Yes‡‡ | Yes | No | No | No | No | No | No | Yes§§ | No |
|
||||
| Optional parameter arity resolution | Yes | No | No | Yes | Yes | No | No | Yes | Yes | Yes | No | Yes | No |
|
||||
| Cross-file binding propagation | Yes | Yes | Yes‖‖ | Yes | Yes¶¶ | Yes*** | Yes | Yes | Partial | Yes*** | Yes*** | Yes*** | Yes*** |
|
||||
| Feature | TS | JS | Java | Kotlin | C# | Go | Rust | Python | PHP | Ruby | Swift | C++ | C | Dart |
|
||||
|---------|:--:|:--:|:----:|:------:|:--:|:--:|:----:|:------:|:---:|:----:|:-----:|:---:|:-:|:----:|
|
||||
| Declarations | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
|
||||
| Parameters | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
|
||||
| Initializer / constructor inference | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
|
||||
| Constructor binding scan | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
|
||||
| For-loop element types | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes††† | Yes | Yes | Yes |
|
||||
| Pattern binding | Yes | Yes | Yes | Yes | No | Yes | Yes | No | No | No | Partial‡‡‡ | No | No | No |
|
||||
| Assignment chains | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | No | Yes | Yes | Yes | Yes |
|
||||
| Field/property type resolution | Yes | No† | Yes | Yes | Yes | Yes | Yes | Yes* | Yes | YARD | No | Yes | No‡ | No |
|
||||
| Comment-based types | JSDoc | JSDoc | No | No | No | No | No | No | PHPDoc | YARD | No | No | No | No |
|
||||
| Return type extraction | JSDoc | JSDoc | No | No | No | No | No | No | PHPDoc | YARD | No | No | No | No |
|
||||
| Call-result variable binding | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes¶ | Yes††† | Yes | No | Yes |
|
||||
| Field access binding | Yes | No† | Yes | Yes | Yes | Yes | Yes | No‖ | Yes | N/A | Yes††† | Yes | No | Yes |
|
||||
| Method-call-result binding | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes¶ | Yes††† | Yes | No | Yes |
|
||||
| Write access (ACCESSES write) | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes§ | Yes | Yes | Yes | No | Yes |
|
||||
| Parameter types extracted | Yes** | No | Yes | Yes | Yes | Yes | Yes | Partial†† | No | No | No | Yes | No | No |
|
||||
| Method overload disambiguation | Yes** | No | Yes | Yes | Yes | No | No | No | No | No | No | Yes | No | No |
|
||||
| Constructor-visible virtual dispatch | Yes | No | Yes | Yes‡‡ | Yes | No | No | No | No | No | No | Yes§§ | No | Yes |
|
||||
| Optional parameter arity resolution | Yes | No | No | Yes | Yes | No | No | Yes | Yes | Yes | No | Yes | No | No |
|
||||
| Cross-file binding propagation | Yes | Yes | Yes‖‖ | Yes | Yes¶¶ | Yes*** | Yes | Yes | Partial | Yes*** | Yes*** | Yes*** | Yes*** | Yes*** |
|
||||
|
||||
\* Python class-level annotated attributes (`address: Address`) now resolve `declaredType` correctly. The `self.x` instance attribute pattern is not yet supported.
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue