mirror of
https://github.com/abhigyanpatwari/GitNexus.git
synced 2026-10-09 03:17:54 +00:00
refactor(group): migrate topic-extractor from regex to tree-sitter queries
Addresses @magyargergo's feedback on #796 that regex-based lookups should use tree-sitter nodes instead, and that the top-level extractors must NOT carry language dependencies. This is phase 1 of a multi-step migration — topic-extractor first because its patterns are the most uniform (16 "call/annotation with first-arg string literal" variants), which makes it a clean proof of the approach before grpc-extractor and http-route-extractor get the same treatment. ## Architecture: language-agnostic orchestrator + per-language plugins The top-level extractor is a thin orchestrator that never imports a tree-sitter grammar or a query string. Per-language knowledge lives in a new `topic-patterns/` folder with one file per language plus a registry that maps file extensions to compiled plugins: ``` src/core/group/extractors/ ├── tree-sitter-scanner.ts # shared, language-agnostic scanning utilities ├── topic-extractor.ts # thin orchestrator (no grammar imports) └── topic-patterns/ ├── types.ts # TopicMeta, Broker ├── index.ts # registry: extension → compiled provider ├── java.ts # tree-sitter-java + JAVA_TOPIC_PROVIDER ├── go.ts # tree-sitter-go + GO_TOPIC_PROVIDER ├── python.ts # tree-sitter-python + PYTHON_TOPIC_PROVIDER └── node.ts # tree-sitter-javascript + tree-sitter-typescript # → JAVASCRIPT_/TYPESCRIPT_/TSX_TOPIC_PROVIDER ``` **Shared scanner (`tree-sitter-scanner.ts`)** — defines `PatternSpec<TMeta>`, `LanguagePatterns<TMeta>`, `CompiledPatterns<TMeta>` and the `scanFile(parser, plugin, content)` helper. Plugins compile their queries eagerly at module load via `compilePatterns()`, so a broken pattern fails loudly at import time instead of silently at scan time. `unquoteLiteral()` handles single/double/template quotes, Python triple-quoted strings, and Go raw backtick strings. **Per-language plugins** own: - the tree-sitter grammar import (this is the ONLY place in `src/core/group/` where tree-sitter grammars are imported), - the query S-expressions, - the `TopicMeta` payload (role, broker, confidence, symbolName) that the orchestrator receives back on every match. Each plugin uses a `@value` capture name to bind the topic literal node. The JavaScript and TypeScript grammars share AST node names for every construct we query, so `node.ts` defines the pattern sources once and compiles them against `JavaScript`, `TypeScript.typescript`, and `TypeScript.tsx` — exporting three providers because `Parser.Query` objects are NOT portable across grammar instances. **Registry (`topic-patterns/index.ts`)** — maps `.java` → Java provider, `.go` → Go, `.py` → Python, `.js`/`.jsx` → JS, `.ts` → TS, `.tsx` → TSX. Also exports `TOPIC_SCAN_GLOB` so adding a new language is a single file-level edit (drop `topic-patterns/<lang>.ts`, import + register it here — zero edits required in `topic-extractor.ts`). **Orchestrator (`topic-extractor.ts`)** — ~110 lines, no grammar or query imports. Per file: `getProviderForFile(rel)` → `scanFile(parser, provider, content)` → `unquoteLiteral(valueText)` → `makeContract(...)`. Reuses one `Parser` instance across files; the scanner calls `setLanguage` per plugin. ## Why this is better than regex 1. **Comments and strings are respected for free.** The old regex would match `// kafkaTemplate.send("fake.topic")` as a real producer; tree-sitter never visits comments or string literals as code nodes, so false positives from commented-out code are eliminated. 2. **Struct/object literal patterns are structural, not textual.** `sarama.ProducerMessage{Topic: "..."}` no longer needs a 300-char lookahead (which was a known cross-match bug partly mitigated by a loop regression test in the self-review). The new query matches a specific `composite_literal` with a specific `qualified_type` and `keyed_element` — exactly one struct literal per match. 3. **No order-of-operations fragility.** Regex for `channel.publish` vs `channel.consume` was independent and file-wide; the AST scopes matches to the specific `call_expression`. 4. **Language-agnostic extension.** Adding Ruby, Rust, or C# topic detection later means dropping one file in `topic-patterns/` — no changes to shared scanner or orchestrator, and no tree-sitter imports leak into top-level code. ## Per-file fault tolerance - Malformed files that tree-sitter can't parse are silently skipped (`parser.parse` is wrapped by `scanFile`). The ingestion pipeline already logs unparseable files at index time. - A syntactically invalid query is caught at `compilePatterns` time, not scan time — broken plugins fail loudly at import. - Per-pattern `matches()` failures are swallowed so one broken query in a plugin doesn't block the rest. ## Tests All 30 existing `topic-extractor.test.ts` tests pass **without any changes to the test file** — they were written as input/output contract tests (given this source file, expect these `ExtractedContract` objects) and that contract is unchanged. Regression coverage includes: - Kafka: Java `@KafkaListener` + `kafkaTemplate.send`; Node `producer.send` + `consumer.subscribe`; Go sarama producer/consumer (sync and async); kafka-go Writer/Reader; Python `KafkaConsumer` + `producer.send/produce` - RabbitMQ: Java `@RabbitListener` + `rabbitTemplate.convertAndSend`; Node `channel.consume/publish/sendToQueue`; Python `basic_consume/ basic_publish` with keyword args - NATS: Go and Node `nc.Subscribe/Publish`; Go and Node JetStream `js.Subscribe/Publish`; Python `await nc.subscribe/publish` Including the regression test for the sarama `ProducerMessage` in-loop case — the AST-based query captures every literal in the file independently, not just the first one after `NewSyncProducer`. ## Neighbor regression check - `topic-extractor.test.ts` — 30/30 pass (rewritten extractor) - `http-route-extractor.test.ts` — 18/18 pass (untouched) - `grpc-extractor.test.ts` — 43/43 pass (untouched) - `manifest-extractor.test.ts` — 8/8 pass (untouched) - Full `npx tsc --noEmit` clean ## Scope discipline (per GUARDRAILS.md) - Only files under `src/core/group/extractors/` are touched; no changes to other extractors, tests, MCP surface, or pipeline.ts. - No CI/release/security config changes, no secrets. - New tree-sitter imports all reference grammars that are already installed as dependencies (`tree-sitter`, `tree-sitter-javascript`, `tree-sitter-typescript`, `tree-sitter-python`, `tree-sitter-java`, `tree-sitter-go` — all in `package.json` for the existing pipeline). ## Phase 2 / phase 3 plan - **Phase 2 (next commit):** rewrite `http-route-extractor.ts` Strategy B (regex fallback) on the same plugin pattern. Graph-assisted Strategy A stays as-is (already uses pipeline-built tree-sitter data via `HANDLES_ROUTE` Cypher queries). - **Phase 3 (commit after):** rewrite `grpc-extractor.ts` for Java / Go / Python / TypeScript detection. `.proto` files are the one outstanding question — there is no `tree-sitter-proto` grammar installed; the in-tree string-sanitizing parser stays as a pragmatic exception with a comment, alternative being to add `tree-sitter-proto` as a dep (open for the maintainer). Co-authored-by: Claude <noreply@anthropic.com>
This commit is contained in:
parent
747937cca1
commit
55ed7443bc
8 changed files with 789 additions and 307 deletions
|
|
@ -1,13 +1,32 @@
|
|||
import * as fs from 'node:fs';
|
||||
import * as path from 'node:path';
|
||||
import { glob } from 'glob';
|
||||
import Parser from 'tree-sitter';
|
||||
import type { ContractExtractor, CypherExecutor } from '../contract-extractor.js';
|
||||
import type { ExtractedContract, RepoHandle } from '../types.js';
|
||||
import { scanFile, unquoteLiteral } from './tree-sitter-scanner.js';
|
||||
import {
|
||||
TOPIC_SCAN_GLOB,
|
||||
getProviderForFile,
|
||||
type Broker,
|
||||
type TopicMeta,
|
||||
} from './topic-patterns/index.js';
|
||||
|
||||
type Broker = 'kafka' | 'rabbitmq' | 'nats';
|
||||
|
||||
const KAFKAJS_CONSUMER_RUN_RE = /consumer\.run\s*\(\s*\{\s*eachMessage:/;
|
||||
const KAFKAJS_SUBSCRIBE_RE = /consumer\.subscribe\s*\(\s*\{\s*topic:\s*['"]([^'"]+)['"]/g;
|
||||
/**
|
||||
* Language-agnostic orchestrator for topic (message broker) contract
|
||||
* extraction. All grammar-specific knowledge lives in `topic-patterns/*`
|
||||
* — this file must not import any tree-sitter grammar directly.
|
||||
*
|
||||
* Flow per file:
|
||||
* 1. `getProviderForFile(rel)` → compiled plugin (or `undefined` if the
|
||||
* file's extension isn't registered, in which case we skip it).
|
||||
* 2. `scanFile(parser, provider, content)` → list of `{meta, valueText}`
|
||||
* pairs, one per matched literal.
|
||||
* 3. `unquoteLiteral(valueText)` → the raw topic string.
|
||||
* 4. `makeContract(topic, meta, relPath)` → `ExtractedContract`.
|
||||
*
|
||||
* Adding a new language is a one-file edit in `topic-patterns/index.ts`.
|
||||
*/
|
||||
|
||||
function readSafe(repoPath: string, rel: string): string | null {
|
||||
const abs = path.resolve(repoPath, rel);
|
||||
|
|
@ -21,281 +40,23 @@ function readSafe(repoPath: string, rel: string): string | null {
|
|||
}
|
||||
}
|
||||
|
||||
function makeContract(
|
||||
topicName: string,
|
||||
role: 'provider' | 'consumer',
|
||||
filePath: string,
|
||||
symbolName: string,
|
||||
confidence: number,
|
||||
broker: Broker,
|
||||
): ExtractedContract {
|
||||
function makeContract(topicName: string, meta: TopicMeta, filePath: string): ExtractedContract {
|
||||
return {
|
||||
contractId: `topic::${topicName}`,
|
||||
type: 'topic',
|
||||
role,
|
||||
role: meta.role,
|
||||
symbolUid: '',
|
||||
symbolRef: { filePath: filePath.replace(/\\/g, '/'), name: symbolName },
|
||||
symbolName,
|
||||
confidence,
|
||||
symbolRef: { filePath: filePath.replace(/\\/g, '/'), name: meta.symbolName },
|
||||
symbolName: meta.symbolName,
|
||||
confidence: meta.confidence,
|
||||
meta: {
|
||||
broker,
|
||||
broker: meta.broker satisfies Broker,
|
||||
topicName,
|
||||
extractionStrategy: 'source_scan',
|
||||
extractionStrategy: 'tree_sitter',
|
||||
},
|
||||
};
|
||||
}
|
||||
|
||||
interface PatternDef {
|
||||
regex: RegExp;
|
||||
role: 'provider' | 'consumer';
|
||||
broker: Broker;
|
||||
confidence: number;
|
||||
topicGroup: number;
|
||||
symbolName: string;
|
||||
}
|
||||
|
||||
// --- Kafka patterns ---
|
||||
const KAFKA_PATTERNS: PatternDef[] = [
|
||||
// Java: @KafkaListener(topics = "xxx")
|
||||
{
|
||||
regex: /@KafkaListener\s*\(\s*topics\s*=\s*"([^"]+)"/g,
|
||||
role: 'consumer',
|
||||
broker: 'kafka',
|
||||
confidence: 0.8,
|
||||
topicGroup: 1,
|
||||
symbolName: 'kafkaListener',
|
||||
},
|
||||
// Java: kafkaTemplate.send("xxx"
|
||||
{
|
||||
regex: /kafkaTemplate\.send\s*\(\s*"([^"]+)"/gi,
|
||||
role: 'provider',
|
||||
broker: 'kafka',
|
||||
confidence: 0.8,
|
||||
topicGroup: 1,
|
||||
symbolName: 'kafkaTemplate.send',
|
||||
},
|
||||
// Node: producer.send({ topic: 'xxx'
|
||||
{
|
||||
regex: /producer\.send\s*\(\s*\{\s*topic:\s*['"]([^'"]+)['"]/g,
|
||||
role: 'provider',
|
||||
broker: 'kafka',
|
||||
confidence: 0.8,
|
||||
topicGroup: 1,
|
||||
symbolName: 'producer.send',
|
||||
},
|
||||
// Node: consumer.subscribe({ topic: 'xxx'
|
||||
{
|
||||
regex: /consumer\.subscribe\s*\(\s*\{\s*topic:\s*['"]([^'"]+)['"]/g,
|
||||
role: 'consumer',
|
||||
broker: 'kafka',
|
||||
confidence: 0.8,
|
||||
topicGroup: 1,
|
||||
symbolName: 'consumer.subscribe',
|
||||
},
|
||||
// Go: consumer.ConsumePartition("xxx"
|
||||
{
|
||||
regex: /\.ConsumePartition\s*\(\s*"([^"]+)"/g,
|
||||
role: 'consumer',
|
||||
broker: 'kafka',
|
||||
confidence: 0.7,
|
||||
topicGroup: 1,
|
||||
symbolName: 'ConsumePartition',
|
||||
},
|
||||
// Python: KafkaConsumer('xxx'
|
||||
{
|
||||
regex: /KafkaConsumer\s*\(\s*['"]([^'"]+)['"]/g,
|
||||
role: 'consumer',
|
||||
broker: 'kafka',
|
||||
confidence: 0.7,
|
||||
topicGroup: 1,
|
||||
symbolName: 'KafkaConsumer',
|
||||
},
|
||||
// Python: producer.send('xxx' or producer.produce('xxx'
|
||||
{
|
||||
regex: /producer\.(?:send|produce)\s*\(\s*['"]([^'"]+)['"]/g,
|
||||
role: 'provider',
|
||||
broker: 'kafka',
|
||||
confidence: 0.7,
|
||||
topicGroup: 1,
|
||||
symbolName: 'producer.send',
|
||||
},
|
||||
// Go: sarama.ProducerMessage{Topic: "xxx"} struct literal (emitted by
|
||||
// both NewSyncProducer and NewAsyncProducer client code paths).
|
||||
//
|
||||
// Previous pattern was `sarama.NewSyncProducer[\s\S]{0,300}?Topic:...`
|
||||
// which anchored to the producer constructor and used a 300-char
|
||||
// lookahead. In a loop like
|
||||
// producer := sarama.NewSyncProducer(...)
|
||||
// for _, item := range items {
|
||||
// msg1 := &sarama.ProducerMessage{Topic: "order.created"}
|
||||
// msg2 := &sarama.ProducerMessage{Topic: "order.shipped"}
|
||||
// }
|
||||
// the regex captured only "order.created" (first Topic after the
|
||||
// constructor) and silently missed "order.shipped". Matching on the
|
||||
// struct literal directly fixes both the false negative in loops and
|
||||
// the spurious cross-message capture when multiple unrelated messages
|
||||
// sit within 300 chars of the constructor.
|
||||
{
|
||||
regex: /sarama\.ProducerMessage\s*\{[\s\S]{0,200}?Topic:\s*"([^"]+)"/g,
|
||||
role: 'provider',
|
||||
broker: 'kafka',
|
||||
confidence: 0.75,
|
||||
topicGroup: 1,
|
||||
symbolName: 'sarama.ProducerMessage',
|
||||
},
|
||||
// Go: kafka-go writer construction. kafka-go does NOT wrap messages in
|
||||
// a struct with a Topic field (the writer owns the topic), so we match
|
||||
// the Writer itself. A 200-char window bridges the gap between
|
||||
// `kafka.NewWriter(...)` / `kafka.Writer{` and the Topic field inside
|
||||
// the config literal — kafka-go writer configs are small and rarely
|
||||
// contain more than one Topic field, so the risk of cross-message
|
||||
// capture is low here.
|
||||
{
|
||||
regex: /kafka\.(?:NewWriter|Writer)\b[\s\S]{0,200}?Topic:\s*"([^"]+)"/g,
|
||||
role: 'provider',
|
||||
broker: 'kafka',
|
||||
confidence: 0.75,
|
||||
topicGroup: 1,
|
||||
symbolName: 'kafka.Writer',
|
||||
},
|
||||
// Go: kafka-go reader construction, mirrors Writer above.
|
||||
{
|
||||
regex: /kafka\.(?:NewReader|Reader)\b[\s\S]{0,200}?Topic:\s*"([^"]+)"/g,
|
||||
role: 'consumer',
|
||||
broker: 'kafka',
|
||||
confidence: 0.75,
|
||||
topicGroup: 1,
|
||||
symbolName: 'kafka.Reader',
|
||||
},
|
||||
];
|
||||
|
||||
// --- RabbitMQ patterns ---
|
||||
const RABBITMQ_PATTERNS: PatternDef[] = [
|
||||
// Java: @RabbitListener(queues = "xxx")
|
||||
{
|
||||
regex: /@RabbitListener\s*\(\s*queues\s*=\s*"([^"]+)"/g,
|
||||
role: 'consumer',
|
||||
broker: 'rabbitmq',
|
||||
confidence: 0.8,
|
||||
topicGroup: 1,
|
||||
symbolName: 'rabbitListener',
|
||||
},
|
||||
// Java: rabbitTemplate.convertAndSend("xxx"
|
||||
{
|
||||
regex: /rabbitTemplate\.convertAndSend\s*\(\s*"([^"]+)"/gi,
|
||||
role: 'provider',
|
||||
broker: 'rabbitmq',
|
||||
confidence: 0.8,
|
||||
topicGroup: 1,
|
||||
symbolName: 'rabbitTemplate.convertAndSend',
|
||||
},
|
||||
// Node: channel.consume("xxx"
|
||||
{
|
||||
regex: /channel\.consume\s*\(\s*"([^"]+)"/g,
|
||||
role: 'consumer',
|
||||
broker: 'rabbitmq',
|
||||
confidence: 0.8,
|
||||
topicGroup: 1,
|
||||
symbolName: 'channel.consume',
|
||||
},
|
||||
// Node: channel.publish("xxx"
|
||||
{
|
||||
regex: /channel\.publish\s*\(\s*"([^"]+)"/g,
|
||||
role: 'provider',
|
||||
broker: 'rabbitmq',
|
||||
confidence: 0.8,
|
||||
topicGroup: 1,
|
||||
symbolName: 'channel.publish',
|
||||
},
|
||||
// Node: channel.sendToQueue("xxx"
|
||||
{
|
||||
regex: /channel\.sendToQueue\s*\(\s*"([^"]+)"/g,
|
||||
role: 'provider',
|
||||
broker: 'rabbitmq',
|
||||
confidence: 0.8,
|
||||
topicGroup: 1,
|
||||
symbolName: 'channel.sendToQueue',
|
||||
},
|
||||
// Python: channel.basic_consume(queue='xxx'
|
||||
{
|
||||
regex: /channel\.basic_consume\s*\(\s*queue\s*=\s*['"]([^'"]+)['"]/g,
|
||||
role: 'consumer',
|
||||
broker: 'rabbitmq',
|
||||
confidence: 0.7,
|
||||
topicGroup: 1,
|
||||
symbolName: 'basic_consume',
|
||||
},
|
||||
// Python: channel.basic_publish(exchange='xxx'
|
||||
{
|
||||
regex: /channel\.basic_publish\s*\([^)]*exchange\s*=\s*['"]([^'"]+)['"]/g,
|
||||
role: 'provider',
|
||||
broker: 'rabbitmq',
|
||||
confidence: 0.7,
|
||||
topicGroup: 1,
|
||||
symbolName: 'basic_publish',
|
||||
},
|
||||
];
|
||||
|
||||
// --- NATS patterns ---
|
||||
const NATS_PATTERNS: PatternDef[] = [
|
||||
// Go/Node: nc.Subscribe("xxx" or nc.subscribe("xxx"
|
||||
{
|
||||
regex: /nc\.(?:S|s)ubscribe\s*\(\s*"([^"]+)"/g,
|
||||
role: 'consumer',
|
||||
broker: 'nats',
|
||||
confidence: 0.8,
|
||||
topicGroup: 1,
|
||||
symbolName: 'nc.Subscribe',
|
||||
},
|
||||
// Go/Node: nc.Publish("xxx" or nc.publish("xxx"
|
||||
{
|
||||
regex: /nc\.(?:P|p)ublish\s*\(\s*"([^"]+)"/g,
|
||||
role: 'provider',
|
||||
broker: 'nats',
|
||||
confidence: 0.8,
|
||||
topicGroup: 1,
|
||||
symbolName: 'nc.Publish',
|
||||
},
|
||||
// Go/Node JetStream: js.Subscribe("xxx"
|
||||
{
|
||||
regex: /js\.(?:S|s)ubscribe\s*\(\s*"([^"]+)"/g,
|
||||
role: 'consumer',
|
||||
broker: 'nats',
|
||||
confidence: 0.8,
|
||||
topicGroup: 1,
|
||||
symbolName: 'js.Subscribe',
|
||||
},
|
||||
// Go/Node JetStream: js.Publish("xxx"
|
||||
{
|
||||
regex: /js\.(?:P|p)ublish\s*\(\s*"([^"]+)"/g,
|
||||
role: 'provider',
|
||||
broker: 'nats',
|
||||
confidence: 0.8,
|
||||
topicGroup: 1,
|
||||
symbolName: 'js.Publish',
|
||||
},
|
||||
// Python: await nc.subscribe("xxx")
|
||||
{
|
||||
regex: /await\s+nc\.subscribe\s*\(\s*['"]([^'"]+)['"]/g,
|
||||
role: 'consumer',
|
||||
broker: 'nats',
|
||||
confidence: 0.75,
|
||||
topicGroup: 1,
|
||||
symbolName: 'nc.subscribe',
|
||||
},
|
||||
// Python: await nc.publish("xxx")
|
||||
{
|
||||
regex: /await\s+nc\.publish\s*\(\s*['"]([^'"]+)['"]/g,
|
||||
role: 'provider',
|
||||
broker: 'nats',
|
||||
confidence: 0.75,
|
||||
topicGroup: 1,
|
||||
symbolName: 'nc.publish',
|
||||
},
|
||||
];
|
||||
|
||||
const ALL_PATTERNS: PatternDef[] = [...KAFKA_PATTERNS, ...RABBITMQ_PATTERNS, ...NATS_PATTERNS];
|
||||
|
||||
export class TopicExtractor implements ContractExtractor {
|
||||
type = 'topic' as const;
|
||||
|
||||
|
|
@ -308,59 +69,37 @@ export class TopicExtractor implements ContractExtractor {
|
|||
repoPath: string,
|
||||
_repo: RepoHandle,
|
||||
): Promise<ExtractedContract[]> {
|
||||
const files = await glob('**/*.{ts,tsx,js,jsx,java,go,py}', {
|
||||
const files = await glob(TOPIC_SCAN_GLOB, {
|
||||
cwd: repoPath,
|
||||
ignore: ['**/node_modules/**', '**/.git/**', '**/vendor/**', '**/dist/**', '**/build/**'],
|
||||
nodir: true,
|
||||
});
|
||||
|
||||
// One parser reused across files; the scanner calls `setLanguage` per
|
||||
// file based on which plugin the registry returns.
|
||||
const parser = new Parser();
|
||||
const out: ExtractedContract[] = [];
|
||||
|
||||
for (const rel of files) {
|
||||
if (rel.endsWith('_test.go')) continue;
|
||||
|
||||
const provider = getProviderForFile(rel);
|
||||
if (!provider) continue;
|
||||
|
||||
const content = readSafe(repoPath, rel);
|
||||
if (!content) continue;
|
||||
out.push(...this.scanFile(content, rel));
|
||||
|
||||
const matches = scanFile(parser, provider, content);
|
||||
for (const match of matches) {
|
||||
const topicName = unquoteLiteral(match.valueText);
|
||||
if (!topicName) continue;
|
||||
out.push(makeContract(topicName, match.meta, rel));
|
||||
}
|
||||
}
|
||||
|
||||
return this.dedupe(out);
|
||||
}
|
||||
|
||||
private scanFile(content: string, filePath: string): ExtractedContract[] {
|
||||
const out: ExtractedContract[] = [];
|
||||
|
||||
for (const pattern of ALL_PATTERNS) {
|
||||
// Reset regex state for each file
|
||||
const re = new RegExp(pattern.regex.source, pattern.regex.flags);
|
||||
let m: RegExpExecArray | null;
|
||||
while ((m = re.exec(content)) !== null) {
|
||||
const topicName = m[pattern.topicGroup];
|
||||
if (!topicName) continue;
|
||||
out.push(
|
||||
makeContract(
|
||||
topicName,
|
||||
pattern.role,
|
||||
filePath,
|
||||
pattern.symbolName,
|
||||
pattern.confidence,
|
||||
pattern.broker,
|
||||
),
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
if (KAFKAJS_CONSUMER_RUN_RE.test(content)) {
|
||||
const subscribeRe = new RegExp(KAFKAJS_SUBSCRIBE_RE.source, KAFKAJS_SUBSCRIBE_RE.flags);
|
||||
let subscribeMatch: RegExpExecArray | null;
|
||||
while ((subscribeMatch = subscribeRe.exec(content)) !== null) {
|
||||
const topicName = subscribeMatch[1];
|
||||
if (!topicName) continue;
|
||||
out.push(makeContract(topicName, 'consumer', filePath, 'consumer.run', 0.75, 'kafka'));
|
||||
}
|
||||
}
|
||||
|
||||
return out;
|
||||
}
|
||||
|
||||
private dedupe(items: ExtractedContract[]): ExtractedContract[] {
|
||||
const seen = new Set<string>();
|
||||
const out: ExtractedContract[] = [];
|
||||
|
|
|
|||
123
gitnexus/src/core/group/extractors/topic-patterns/go.ts
Normal file
123
gitnexus/src/core/group/extractors/topic-patterns/go.ts
Normal file
|
|
@ -0,0 +1,123 @@
|
|||
import Go from 'tree-sitter-go';
|
||||
import { compilePatterns, type LanguagePatterns } from '../tree-sitter-scanner.js';
|
||||
import type { TopicMeta } from './types.js';
|
||||
|
||||
/**
|
||||
* Go topic extraction patterns.
|
||||
*
|
||||
* Detects Sarama, segmentio/kafka-go and nats.go producer/consumer APIs:
|
||||
* - `X.ConsumePartition("topic", ...)`
|
||||
* - `sarama.ProducerMessage{Topic: "xxx"}`
|
||||
* - `kafka.Writer{Topic: "xxx"}` / `kafka.WriterConfig{Topic: ...}`
|
||||
* - `kafka.Reader{Topic: "xxx"}` / `kafka.ReaderConfig{Topic: ...}`
|
||||
* - `nc.Subscribe("topic", ...)` / `js.Subscribe("topic", ...)`
|
||||
* - `nc.Publish("topic", ...)` / `js.Publish("topic", ...)`
|
||||
*
|
||||
* Every query MUST bind `@value` to the topic literal node.
|
||||
*/
|
||||
const GO_TOPIC_SPEC: LanguagePatterns<TopicMeta> = {
|
||||
name: 'go-topic',
|
||||
language: Go,
|
||||
patterns: [
|
||||
{
|
||||
meta: {
|
||||
role: 'consumer',
|
||||
broker: 'kafka',
|
||||
confidence: 0.7,
|
||||
symbolName: 'ConsumePartition',
|
||||
},
|
||||
query: `
|
||||
(call_expression
|
||||
function: (selector_expression
|
||||
field: (field_identifier) @method (#eq? @method "ConsumePartition"))
|
||||
arguments: (argument_list . (interpreted_string_literal) @value))
|
||||
`,
|
||||
},
|
||||
{
|
||||
meta: {
|
||||
role: 'provider',
|
||||
broker: 'kafka',
|
||||
confidence: 0.75,
|
||||
symbolName: 'sarama.ProducerMessage',
|
||||
},
|
||||
query: `
|
||||
(composite_literal
|
||||
type: (qualified_type
|
||||
package: (package_identifier) @pkg (#eq? @pkg "sarama")
|
||||
name: (type_identifier) @ty (#eq? @ty "ProducerMessage"))
|
||||
body: (literal_value
|
||||
(keyed_element
|
||||
(literal_element (identifier) @field (#eq? @field "Topic"))
|
||||
(literal_element (interpreted_string_literal) @value))))
|
||||
`,
|
||||
},
|
||||
{
|
||||
meta: {
|
||||
role: 'provider',
|
||||
broker: 'kafka',
|
||||
confidence: 0.75,
|
||||
symbolName: 'kafka.Writer',
|
||||
},
|
||||
query: `
|
||||
(composite_literal
|
||||
type: (qualified_type
|
||||
package: (package_identifier) @pkg (#eq? @pkg "kafka")
|
||||
name: (type_identifier) @ty (#match? @ty "^(Writer|WriterConfig)$"))
|
||||
body: (literal_value
|
||||
(keyed_element
|
||||
(literal_element (identifier) @field (#eq? @field "Topic"))
|
||||
(literal_element (interpreted_string_literal) @value))))
|
||||
`,
|
||||
},
|
||||
{
|
||||
meta: {
|
||||
role: 'consumer',
|
||||
broker: 'kafka',
|
||||
confidence: 0.75,
|
||||
symbolName: 'kafka.Reader',
|
||||
},
|
||||
query: `
|
||||
(composite_literal
|
||||
type: (qualified_type
|
||||
package: (package_identifier) @pkg (#eq? @pkg "kafka")
|
||||
name: (type_identifier) @ty (#match? @ty "^(Reader|ReaderConfig)$"))
|
||||
body: (literal_value
|
||||
(keyed_element
|
||||
(literal_element (identifier) @field (#eq? @field "Topic"))
|
||||
(literal_element (interpreted_string_literal) @value))))
|
||||
`,
|
||||
},
|
||||
{
|
||||
meta: {
|
||||
role: 'consumer',
|
||||
broker: 'nats',
|
||||
confidence: 0.8,
|
||||
symbolName: 'nc.Subscribe',
|
||||
},
|
||||
query: `
|
||||
(call_expression
|
||||
function: (selector_expression
|
||||
operand: (identifier) @obj (#match? @obj "^(nc|js)$")
|
||||
field: (field_identifier) @method (#match? @method "^[Ss]ubscribe$"))
|
||||
arguments: (argument_list . (interpreted_string_literal) @value))
|
||||
`,
|
||||
},
|
||||
{
|
||||
meta: {
|
||||
role: 'provider',
|
||||
broker: 'nats',
|
||||
confidence: 0.8,
|
||||
symbolName: 'nc.Publish',
|
||||
},
|
||||
query: `
|
||||
(call_expression
|
||||
function: (selector_expression
|
||||
operand: (identifier) @obj (#match? @obj "^(nc|js)$")
|
||||
field: (field_identifier) @method (#match? @method "^[Pp]ublish$"))
|
||||
arguments: (argument_list . (interpreted_string_literal) @value))
|
||||
`,
|
||||
},
|
||||
],
|
||||
};
|
||||
|
||||
export const GO_TOPIC_PROVIDER = compilePatterns(GO_TOPIC_SPEC);
|
||||
49
gitnexus/src/core/group/extractors/topic-patterns/index.ts
Normal file
49
gitnexus/src/core/group/extractors/topic-patterns/index.ts
Normal file
|
|
@ -0,0 +1,49 @@
|
|||
import * as path from 'node:path';
|
||||
import type { CompiledPatterns } from '../tree-sitter-scanner.js';
|
||||
import type { TopicMeta } from './types.js';
|
||||
import { JAVA_TOPIC_PROVIDER } from './java.js';
|
||||
import { GO_TOPIC_PROVIDER } from './go.js';
|
||||
import { PYTHON_TOPIC_PROVIDER } from './python.js';
|
||||
import {
|
||||
JAVASCRIPT_TOPIC_PROVIDER,
|
||||
TYPESCRIPT_TOPIC_PROVIDER,
|
||||
TSX_TOPIC_PROVIDER,
|
||||
} from './node.js';
|
||||
|
||||
export type { TopicMeta, Broker } from './types.js';
|
||||
|
||||
/**
|
||||
* File-extension → compiled-plugin registry for topic extraction. The
|
||||
* top-level orchestrator (`topic-extractor.ts`) looks up the plugin for
|
||||
* each file it visits and delegates the scanning to `tree-sitter-scanner`.
|
||||
*
|
||||
* Keys are lowercase extensions including the leading dot. To add a new
|
||||
* language, drop a `topic-patterns/<lang>.ts` that exports a compiled
|
||||
* provider, import it here and register the extension(s). No edits to
|
||||
* `topic-extractor.ts` are required.
|
||||
*/
|
||||
const REGISTRY: Record<string, CompiledPatterns<TopicMeta>> = {
|
||||
'.java': JAVA_TOPIC_PROVIDER,
|
||||
'.go': GO_TOPIC_PROVIDER,
|
||||
'.py': PYTHON_TOPIC_PROVIDER,
|
||||
'.js': JAVASCRIPT_TOPIC_PROVIDER,
|
||||
'.jsx': JAVASCRIPT_TOPIC_PROVIDER,
|
||||
'.ts': TYPESCRIPT_TOPIC_PROVIDER,
|
||||
'.tsx': TSX_TOPIC_PROVIDER,
|
||||
};
|
||||
|
||||
/**
|
||||
* Glob pattern for files worth scanning. Kept here so adding a new
|
||||
* language to the registry also widens the glob automatically via a
|
||||
* single edit.
|
||||
*/
|
||||
export const TOPIC_SCAN_GLOB = '**/*.{ts,tsx,js,jsx,java,go,py}';
|
||||
|
||||
/**
|
||||
* Return the compiled provider registered for the given file's
|
||||
* extension, or `undefined` if the extension is not registered.
|
||||
*/
|
||||
export function getProviderForFile(rel: string): CompiledPatterns<TopicMeta> | undefined {
|
||||
const ext = path.extname(rel).toLowerCase();
|
||||
return REGISTRY[ext];
|
||||
}
|
||||
83
gitnexus/src/core/group/extractors/topic-patterns/java.ts
Normal file
83
gitnexus/src/core/group/extractors/topic-patterns/java.ts
Normal file
|
|
@ -0,0 +1,83 @@
|
|||
import Java from 'tree-sitter-java';
|
||||
import { compilePatterns, type LanguagePatterns } from '../tree-sitter-scanner.js';
|
||||
import type { TopicMeta } from './types.js';
|
||||
|
||||
/**
|
||||
* Java topic extraction patterns.
|
||||
*
|
||||
* Detects Kafka and RabbitMQ (Spring conventions) producer/consumer APIs:
|
||||
* - `@KafkaListener(topics = "xxx")`
|
||||
* - `@RabbitListener(queues = "xxx")`
|
||||
* - `kafkaTemplate.send("xxx", ...)`
|
||||
* - `rabbitTemplate.convertAndSend("xxx", ...)`
|
||||
*
|
||||
* Every query MUST bind `@value` to the topic literal node.
|
||||
*/
|
||||
const JAVA_TOPIC_SPEC: LanguagePatterns<TopicMeta> = {
|
||||
name: 'java-topic',
|
||||
language: Java,
|
||||
patterns: [
|
||||
{
|
||||
meta: {
|
||||
role: 'consumer',
|
||||
broker: 'kafka',
|
||||
confidence: 0.8,
|
||||
symbolName: 'kafkaListener',
|
||||
},
|
||||
query: `
|
||||
(annotation
|
||||
name: (identifier) @name (#eq? @name "KafkaListener")
|
||||
arguments: (annotation_argument_list
|
||||
(element_value_pair
|
||||
key: (identifier) @key (#eq? @key "topics")
|
||||
value: (string_literal) @value)))
|
||||
`,
|
||||
},
|
||||
{
|
||||
meta: {
|
||||
role: 'consumer',
|
||||
broker: 'rabbitmq',
|
||||
confidence: 0.8,
|
||||
symbolName: 'rabbitListener',
|
||||
},
|
||||
query: `
|
||||
(annotation
|
||||
name: (identifier) @name (#eq? @name "RabbitListener")
|
||||
arguments: (annotation_argument_list
|
||||
(element_value_pair
|
||||
key: (identifier) @key (#eq? @key "queues")
|
||||
value: (string_literal) @value)))
|
||||
`,
|
||||
},
|
||||
{
|
||||
meta: {
|
||||
role: 'provider',
|
||||
broker: 'kafka',
|
||||
confidence: 0.8,
|
||||
symbolName: 'kafkaTemplate.send',
|
||||
},
|
||||
query: `
|
||||
(method_invocation
|
||||
object: (identifier) @obj (#eq? @obj "kafkaTemplate")
|
||||
name: (identifier) @method (#eq? @method "send")
|
||||
arguments: (argument_list . (string_literal) @value))
|
||||
`,
|
||||
},
|
||||
{
|
||||
meta: {
|
||||
role: 'provider',
|
||||
broker: 'rabbitmq',
|
||||
confidence: 0.8,
|
||||
symbolName: 'rabbitTemplate.convertAndSend',
|
||||
},
|
||||
query: `
|
||||
(method_invocation
|
||||
object: (identifier) @obj (#eq? @obj "rabbitTemplate")
|
||||
name: (identifier) @method (#eq? @method "convertAndSend")
|
||||
arguments: (argument_list . (string_literal) @value))
|
||||
`,
|
||||
},
|
||||
],
|
||||
};
|
||||
|
||||
export const JAVA_TOPIC_PROVIDER = compilePatterns(JAVA_TOPIC_SPEC);
|
||||
165
gitnexus/src/core/group/extractors/topic-patterns/node.ts
Normal file
165
gitnexus/src/core/group/extractors/topic-patterns/node.ts
Normal file
|
|
@ -0,0 +1,165 @@
|
|||
import JavaScript from 'tree-sitter-javascript';
|
||||
import TypeScript from 'tree-sitter-typescript';
|
||||
import {
|
||||
compilePatterns,
|
||||
type LanguagePatterns,
|
||||
type PatternSpec,
|
||||
} from '../tree-sitter-scanner.js';
|
||||
import type { TopicMeta } from './types.js';
|
||||
|
||||
/**
|
||||
* Node.js / TypeScript topic extraction patterns.
|
||||
*
|
||||
* Detects kafkajs, amqplib (RabbitMQ), and nats.js producer/consumer APIs:
|
||||
* - `producer.send({ topic: 'xxx', ... })` (kafkajs)
|
||||
* - `consumer.subscribe({ topic: 'xxx', ... })` (kafkajs)
|
||||
* - `channel.consume("queue", ...)` / `channel.publish(...)` / `channel.sendToQueue(...)`
|
||||
* - `nc.subscribe("topic")` / `js.subscribe("topic")`
|
||||
* - `nc.publish("topic", ...)` / `js.publish("topic", ...)`
|
||||
*
|
||||
* The JavaScript and TypeScript tree-sitter grammars share node type
|
||||
* names for every construct we query here, so the pattern sources are
|
||||
* defined once and compiled against each grammar variant. We export three
|
||||
* providers because Parser.Query objects are NOT portable across grammar
|
||||
* instances — `.js` files use the JavaScript grammar, `.ts` uses
|
||||
* TypeScript.typescript, and `.tsx` uses TypeScript.tsx.
|
||||
*
|
||||
* Every query MUST bind `@value` to the topic literal node.
|
||||
*/
|
||||
const NODE_TOPIC_PATTERNS: PatternSpec<TopicMeta>[] = [
|
||||
{
|
||||
meta: {
|
||||
role: 'provider',
|
||||
broker: 'kafka',
|
||||
confidence: 0.8,
|
||||
symbolName: 'producer.send',
|
||||
},
|
||||
query: `
|
||||
(call_expression
|
||||
function: (member_expression
|
||||
object: (identifier) @obj (#eq? @obj "producer")
|
||||
property: (property_identifier) @prop (#eq? @prop "send"))
|
||||
arguments: (arguments
|
||||
(object
|
||||
(pair
|
||||
key: (property_identifier) @key (#eq? @key "topic")
|
||||
value: [(string) (template_string)] @value))))
|
||||
`,
|
||||
},
|
||||
{
|
||||
meta: {
|
||||
role: 'consumer',
|
||||
broker: 'kafka',
|
||||
confidence: 0.8,
|
||||
symbolName: 'consumer.subscribe',
|
||||
},
|
||||
query: `
|
||||
(call_expression
|
||||
function: (member_expression
|
||||
object: (identifier) @obj (#eq? @obj "consumer")
|
||||
property: (property_identifier) @prop (#eq? @prop "subscribe"))
|
||||
arguments: (arguments
|
||||
(object
|
||||
(pair
|
||||
key: (property_identifier) @key (#eq? @key "topic")
|
||||
value: [(string) (template_string)] @value))))
|
||||
`,
|
||||
},
|
||||
{
|
||||
meta: {
|
||||
role: 'consumer',
|
||||
broker: 'rabbitmq',
|
||||
confidence: 0.8,
|
||||
symbolName: 'channel.consume',
|
||||
},
|
||||
query: `
|
||||
(call_expression
|
||||
function: (member_expression
|
||||
object: (identifier) @obj (#eq? @obj "channel")
|
||||
property: (property_identifier) @prop (#eq? @prop "consume"))
|
||||
arguments: (arguments . [(string) (template_string)] @value))
|
||||
`,
|
||||
},
|
||||
{
|
||||
meta: {
|
||||
role: 'provider',
|
||||
broker: 'rabbitmq',
|
||||
confidence: 0.8,
|
||||
symbolName: 'channel.publish',
|
||||
},
|
||||
query: `
|
||||
(call_expression
|
||||
function: (member_expression
|
||||
object: (identifier) @obj (#eq? @obj "channel")
|
||||
property: (property_identifier) @prop (#eq? @prop "publish"))
|
||||
arguments: (arguments . [(string) (template_string)] @value))
|
||||
`,
|
||||
},
|
||||
{
|
||||
meta: {
|
||||
role: 'provider',
|
||||
broker: 'rabbitmq',
|
||||
confidence: 0.8,
|
||||
symbolName: 'channel.sendToQueue',
|
||||
},
|
||||
query: `
|
||||
(call_expression
|
||||
function: (member_expression
|
||||
object: (identifier) @obj (#eq? @obj "channel")
|
||||
property: (property_identifier) @prop (#eq? @prop "sendToQueue"))
|
||||
arguments: (arguments . [(string) (template_string)] @value))
|
||||
`,
|
||||
},
|
||||
{
|
||||
meta: {
|
||||
role: 'consumer',
|
||||
broker: 'nats',
|
||||
confidence: 0.8,
|
||||
symbolName: 'nc.subscribe',
|
||||
},
|
||||
query: `
|
||||
(call_expression
|
||||
function: (member_expression
|
||||
object: (identifier) @obj (#match? @obj "^(nc|js)$")
|
||||
property: (property_identifier) @prop (#match? @prop "^[Ss]ubscribe$"))
|
||||
arguments: (arguments . [(string) (template_string)] @value))
|
||||
`,
|
||||
},
|
||||
{
|
||||
meta: {
|
||||
role: 'provider',
|
||||
broker: 'nats',
|
||||
confidence: 0.8,
|
||||
symbolName: 'nc.publish',
|
||||
},
|
||||
query: `
|
||||
(call_expression
|
||||
function: (member_expression
|
||||
object: (identifier) @obj (#match? @obj "^(nc|js)$")
|
||||
property: (property_identifier) @prop (#match? @prop "^[Pp]ublish$"))
|
||||
arguments: (arguments . [(string) (template_string)] @value))
|
||||
`,
|
||||
},
|
||||
];
|
||||
|
||||
const JAVASCRIPT_TOPIC_SPEC: LanguagePatterns<TopicMeta> = {
|
||||
name: 'javascript-topic',
|
||||
language: JavaScript,
|
||||
patterns: NODE_TOPIC_PATTERNS,
|
||||
};
|
||||
|
||||
const TYPESCRIPT_TOPIC_SPEC: LanguagePatterns<TopicMeta> = {
|
||||
name: 'typescript-topic',
|
||||
language: TypeScript.typescript,
|
||||
patterns: NODE_TOPIC_PATTERNS,
|
||||
};
|
||||
|
||||
const TSX_TOPIC_SPEC: LanguagePatterns<TopicMeta> = {
|
||||
name: 'tsx-topic',
|
||||
language: TypeScript.tsx,
|
||||
patterns: NODE_TOPIC_PATTERNS,
|
||||
};
|
||||
|
||||
export const JAVASCRIPT_TOPIC_PROVIDER = compilePatterns(JAVASCRIPT_TOPIC_SPEC);
|
||||
export const TYPESCRIPT_TOPIC_PROVIDER = compilePatterns(TYPESCRIPT_TOPIC_SPEC);
|
||||
export const TSX_TOPIC_PROVIDER = compilePatterns(TSX_TOPIC_SPEC);
|
||||
119
gitnexus/src/core/group/extractors/topic-patterns/python.ts
Normal file
119
gitnexus/src/core/group/extractors/topic-patterns/python.ts
Normal file
|
|
@ -0,0 +1,119 @@
|
|||
import Python from 'tree-sitter-python';
|
||||
import { compilePatterns, type LanguagePatterns } from '../tree-sitter-scanner.js';
|
||||
import type { TopicMeta } from './types.js';
|
||||
|
||||
/**
|
||||
* Python topic extraction patterns.
|
||||
*
|
||||
* Detects kafka-python, pika (RabbitMQ), and nats-py producer/consumer APIs:
|
||||
* - `KafkaConsumer('topic', ...)`
|
||||
* - `producer.send('topic', ...)` / `producer.produce('topic', ...)`
|
||||
* - `channel.basic_consume(queue='xxx', ...)`
|
||||
* - `channel.basic_publish(exchange='xxx', ...)`
|
||||
* - `await nc.subscribe('topic')`
|
||||
* - `await nc.publish('topic', ...)`
|
||||
*
|
||||
* Every query MUST bind `@value` to the topic literal node.
|
||||
*/
|
||||
const PYTHON_TOPIC_SPEC: LanguagePatterns<TopicMeta> = {
|
||||
name: 'python-topic',
|
||||
language: Python,
|
||||
patterns: [
|
||||
{
|
||||
meta: {
|
||||
role: 'consumer',
|
||||
broker: 'kafka',
|
||||
confidence: 0.7,
|
||||
symbolName: 'KafkaConsumer',
|
||||
},
|
||||
query: `
|
||||
(call
|
||||
function: (identifier) @func (#eq? @func "KafkaConsumer")
|
||||
arguments: (argument_list . (string) @value))
|
||||
`,
|
||||
},
|
||||
{
|
||||
meta: {
|
||||
role: 'provider',
|
||||
broker: 'kafka',
|
||||
confidence: 0.7,
|
||||
symbolName: 'producer.send',
|
||||
},
|
||||
query: `
|
||||
(call
|
||||
function: (attribute
|
||||
object: (identifier) @obj (#eq? @obj "producer")
|
||||
attribute: (identifier) @method (#match? @method "^(send|produce)$"))
|
||||
arguments: (argument_list . (string) @value))
|
||||
`,
|
||||
},
|
||||
{
|
||||
meta: {
|
||||
role: 'consumer',
|
||||
broker: 'rabbitmq',
|
||||
confidence: 0.7,
|
||||
symbolName: 'basic_consume',
|
||||
},
|
||||
query: `
|
||||
(call
|
||||
function: (attribute
|
||||
object: (identifier) @obj (#eq? @obj "channel")
|
||||
attribute: (identifier) @method (#eq? @method "basic_consume"))
|
||||
arguments: (argument_list
|
||||
(keyword_argument
|
||||
name: (identifier) @kw (#eq? @kw "queue")
|
||||
value: (string) @value)))
|
||||
`,
|
||||
},
|
||||
{
|
||||
meta: {
|
||||
role: 'provider',
|
||||
broker: 'rabbitmq',
|
||||
confidence: 0.7,
|
||||
symbolName: 'basic_publish',
|
||||
},
|
||||
query: `
|
||||
(call
|
||||
function: (attribute
|
||||
object: (identifier) @obj (#eq? @obj "channel")
|
||||
attribute: (identifier) @method (#eq? @method "basic_publish"))
|
||||
arguments: (argument_list
|
||||
(keyword_argument
|
||||
name: (identifier) @kw (#eq? @kw "exchange")
|
||||
value: (string) @value)))
|
||||
`,
|
||||
},
|
||||
{
|
||||
meta: {
|
||||
role: 'consumer',
|
||||
broker: 'nats',
|
||||
confidence: 0.75,
|
||||
symbolName: 'nc.subscribe',
|
||||
},
|
||||
query: `
|
||||
(call
|
||||
function: (attribute
|
||||
object: (identifier) @obj (#eq? @obj "nc")
|
||||
attribute: (identifier) @method (#eq? @method "subscribe"))
|
||||
arguments: (argument_list . (string) @value))
|
||||
`,
|
||||
},
|
||||
{
|
||||
meta: {
|
||||
role: 'provider',
|
||||
broker: 'nats',
|
||||
confidence: 0.75,
|
||||
symbolName: 'nc.publish',
|
||||
},
|
||||
query: `
|
||||
(call
|
||||
function: (attribute
|
||||
object: (identifier) @obj (#eq? @obj "nc")
|
||||
attribute: (identifier) @method (#eq? @method "publish"))
|
||||
arguments: (argument_list . (string) @value))
|
||||
`,
|
||||
},
|
||||
],
|
||||
};
|
||||
|
||||
export const PYTHON_TOPIC_PROVIDER = compilePatterns(PYTHON_TOPIC_SPEC);
|
||||
27
gitnexus/src/core/group/extractors/topic-patterns/types.ts
Normal file
27
gitnexus/src/core/group/extractors/topic-patterns/types.ts
Normal file
|
|
@ -0,0 +1,27 @@
|
|||
/**
|
||||
* Shared types for the topic-extractor language plugins.
|
||||
*
|
||||
* Each plugin lives in its own file (java.ts, go.ts, ...) and owns the
|
||||
* tree-sitter grammar import + query sources. The top-level
|
||||
* `topic-extractor.ts` orchestrator only knows about this type module and
|
||||
* the plugin registry (`./index.ts`). It MUST NOT import any grammar or
|
||||
* query text directly — that's the whole point of the split.
|
||||
*/
|
||||
|
||||
export type Broker = 'kafka' | 'rabbitmq' | 'nats';
|
||||
|
||||
/**
|
||||
* Per-pattern payload every topic plugin attaches to its query. Whatever
|
||||
* the pattern matches, the orchestrator receives this object verbatim
|
||||
* and uses it to build an `ExtractedContract`.
|
||||
*
|
||||
* Plugins produce one `TopicMeta` per pattern (not per match) because a
|
||||
* single query uniquely identifies its broker/role/confidence triple.
|
||||
*/
|
||||
export interface TopicMeta {
|
||||
role: 'provider' | 'consumer';
|
||||
broker: Broker;
|
||||
confidence: number;
|
||||
/** Short human-readable label of the API being detected. */
|
||||
symbolName: string;
|
||||
}
|
||||
177
gitnexus/src/core/group/extractors/tree-sitter-scanner.ts
Normal file
177
gitnexus/src/core/group/extractors/tree-sitter-scanner.ts
Normal file
|
|
@ -0,0 +1,177 @@
|
|||
import Parser from 'tree-sitter';
|
||||
|
||||
/**
|
||||
* Shared, language-agnostic tree-sitter scanning utilities used by group
|
||||
* extractors (topic, http, grpc, ...).
|
||||
*
|
||||
* Design goals:
|
||||
* - The top-level extractors must not import any tree-sitter grammar.
|
||||
* - Per-language plugins own their grammar import, their query sources,
|
||||
* and the mapping from capture → meta.
|
||||
* - This module provides the plumbing: compile queries once per plugin,
|
||||
* parse a file with a given grammar, run all patterns, and return the
|
||||
* captured `string_literal`-style nodes together with the plugin's meta.
|
||||
*/
|
||||
|
||||
/**
|
||||
* One pattern owned by a language plugin. Each pattern owns a tree-sitter
|
||||
* S-expression query. The query MUST contain a capture named `@value`
|
||||
* whose node text is the literal we want to extract (string/template/etc).
|
||||
*
|
||||
* `TMeta` is the plugin-specific payload the orchestrator receives back
|
||||
* when this pattern matches — e.g. for topic extraction it carries the
|
||||
* broker name, role, confidence, symbol name.
|
||||
*/
|
||||
export interface PatternSpec<TMeta> {
|
||||
/** Tree-sitter S-expression. MUST contain a `@value` capture. */
|
||||
query: string;
|
||||
/** Plugin-specific payload returned on every match. */
|
||||
meta: TMeta;
|
||||
}
|
||||
|
||||
/**
|
||||
* A set of patterns owned by one language plugin, bound to a specific
|
||||
* tree-sitter grammar.
|
||||
*
|
||||
* `language` is typed as `unknown` because tree-sitter's TypeScript
|
||||
* declarations use `any` for the grammar object, and the grammar modules
|
||||
* export different shapes (plain grammar vs. namespace with `typescript`
|
||||
* / `tsx` members). Callers pass the concrete grammar object; this
|
||||
* module forwards it to `parser.setLanguage` / `new Parser.Query`.
|
||||
*/
|
||||
export interface LanguagePatterns<TMeta> {
|
||||
/** Human-readable plugin name for diagnostics. */
|
||||
name: string;
|
||||
/** tree-sitter grammar object. */
|
||||
language: unknown;
|
||||
/** Patterns authored against `language`. */
|
||||
patterns: PatternSpec<TMeta>[];
|
||||
}
|
||||
|
||||
/**
|
||||
* Compiled form of a `LanguagePatterns` bundle. Queries are compiled
|
||||
* eagerly at module load time so a broken grammar/query pair fails
|
||||
* loudly the first time the plugin is imported, instead of silently
|
||||
* at scan time when no contract is produced.
|
||||
*/
|
||||
export interface CompiledPatterns<TMeta> {
|
||||
name: string;
|
||||
language: unknown;
|
||||
patterns: CompiledPattern<TMeta>[];
|
||||
}
|
||||
|
||||
export interface CompiledPattern<TMeta> {
|
||||
query: Parser.Query;
|
||||
meta: TMeta;
|
||||
}
|
||||
|
||||
/**
|
||||
* One match returned by `scanFile`. The orchestrator receives the raw
|
||||
* literal text (still including any surrounding quotes) together with
|
||||
* the plugin meta, and is responsible for calling `unquoteLiteral` /
|
||||
* emitting a domain object (ExtractedContract, Route, ...).
|
||||
*/
|
||||
export interface ScanMatch<TMeta> {
|
||||
meta: TMeta;
|
||||
/** The node captured as `@value` (the literal). */
|
||||
valueNode: Parser.SyntaxNode;
|
||||
/** Raw text of the captured value node — caller must unquote. */
|
||||
valueText: string;
|
||||
}
|
||||
|
||||
/**
|
||||
* Compile a LanguagePatterns bundle. Call this once per plugin, at
|
||||
* module load time, and export the result. Throws if any pattern
|
||||
* fails to compile against the grammar — that's a bug in the plugin
|
||||
* author's query, not a runtime condition.
|
||||
*/
|
||||
export function compilePatterns<TMeta>(bundle: LanguagePatterns<TMeta>): CompiledPatterns<TMeta> {
|
||||
const compiled: CompiledPattern<TMeta>[] = [];
|
||||
for (const spec of bundle.patterns) {
|
||||
try {
|
||||
const query = new Parser.Query(bundle.language, spec.query);
|
||||
compiled.push({ query, meta: spec.meta });
|
||||
} catch (err) {
|
||||
const message = err instanceof Error ? err.message : String(err);
|
||||
throw new Error(
|
||||
`[tree-sitter-scanner] Failed to compile pattern in ${bundle.name}: ${message}\n` +
|
||||
`Query source:\n${spec.query}`,
|
||||
);
|
||||
}
|
||||
}
|
||||
return { name: bundle.name, language: bundle.language, patterns: compiled };
|
||||
}
|
||||
|
||||
/**
|
||||
* Parse `content` as source code of the plugin's language and run every
|
||||
* compiled pattern against the resulting AST. Returns one `ScanMatch` per
|
||||
* matched `@value` capture, carrying the plugin's meta payload.
|
||||
*
|
||||
* Errors are swallowed at the file level (malformed file must not abort
|
||||
* the whole extract). Individual pattern failures are swallowed too so
|
||||
* a single unusable query doesn't block the rest of the plugin.
|
||||
*/
|
||||
export function scanFile<TMeta>(
|
||||
parser: Parser,
|
||||
plugin: CompiledPatterns<TMeta>,
|
||||
content: string,
|
||||
): ScanMatch<TMeta>[] {
|
||||
const out: ScanMatch<TMeta>[] = [];
|
||||
let tree: Parser.Tree;
|
||||
try {
|
||||
parser.setLanguage(plugin.language);
|
||||
tree = parser.parse(content);
|
||||
} catch {
|
||||
return out;
|
||||
}
|
||||
|
||||
for (const compiled of plugin.patterns) {
|
||||
let matches: Parser.QueryMatch[];
|
||||
try {
|
||||
matches = compiled.query.matches(tree.rootNode);
|
||||
} catch {
|
||||
continue;
|
||||
}
|
||||
for (const match of matches) {
|
||||
const valueCapture = match.captures.find((c) => c.name === 'value');
|
||||
if (!valueCapture) continue;
|
||||
out.push({
|
||||
meta: compiled.meta,
|
||||
valueNode: valueCapture.node,
|
||||
valueText: valueCapture.node.text,
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
return out;
|
||||
}
|
||||
|
||||
/**
|
||||
* Strip enclosing quotes from a tree-sitter string literal node's text.
|
||||
* Handles single / double / template quotes, Python triple-quoted strings,
|
||||
* and Go raw string literals (backticks).
|
||||
*
|
||||
* Returns null for empty/nullish input so callers can uniformly skip
|
||||
* captures whose value is missing.
|
||||
*/
|
||||
export function unquoteLiteral(raw: string): string | null {
|
||||
if (!raw) return null;
|
||||
|
||||
// Python triple-quoted
|
||||
if (
|
||||
(raw.startsWith('"""') && raw.endsWith('"""')) ||
|
||||
(raw.startsWith("'''") && raw.endsWith("'''"))
|
||||
) {
|
||||
return raw.slice(3, -3);
|
||||
}
|
||||
|
||||
const first = raw[0];
|
||||
const last = raw[raw.length - 1];
|
||||
if ((first === '"' || first === "'" || first === '`') && last === first && raw.length >= 2) {
|
||||
return raw.slice(1, -1);
|
||||
}
|
||||
|
||||
// Some grammars expose the string content without quotes already (e.g.
|
||||
// Python `string_content` child). Return as-is.
|
||||
return raw;
|
||||
}
|
||||
Loading…
Add table
Reference in a new issue