refactor(group): migrate topic-extractor from regex to tree-sitter queries

Addresses @magyargergo's feedback on #796 that regex-based lookups
should use tree-sitter nodes instead, and that the top-level
extractors must NOT carry language dependencies. This is phase 1 of
a multi-step migration — topic-extractor first because its patterns
are the most uniform (16 "call/annotation with first-arg string
literal" variants), which makes it a clean proof of the approach
before grpc-extractor and http-route-extractor get the same treatment.

## Architecture: language-agnostic orchestrator + per-language plugins

The top-level extractor is a thin orchestrator that never imports a
tree-sitter grammar or a query string. Per-language knowledge lives
in a new `topic-patterns/` folder with one file per language plus a
registry that maps file extensions to compiled plugins:

```
src/core/group/extractors/
├── tree-sitter-scanner.ts         # shared, language-agnostic scanning utilities
├── topic-extractor.ts              # thin orchestrator (no grammar imports)
└── topic-patterns/
    ├── types.ts                    # TopicMeta, Broker
    ├── index.ts                    # registry: extension → compiled provider
    ├── java.ts                     # tree-sitter-java + JAVA_TOPIC_PROVIDER
    ├── go.ts                       # tree-sitter-go + GO_TOPIC_PROVIDER
    ├── python.ts                   # tree-sitter-python + PYTHON_TOPIC_PROVIDER
    └── node.ts                     # tree-sitter-javascript + tree-sitter-typescript
                                    # → JAVASCRIPT_/TYPESCRIPT_/TSX_TOPIC_PROVIDER
```

**Shared scanner (`tree-sitter-scanner.ts`)** — defines
`PatternSpec<TMeta>`, `LanguagePatterns<TMeta>`, `CompiledPatterns<TMeta>`
and the `scanFile(parser, plugin, content)` helper. Plugins compile their
queries eagerly at module load via `compilePatterns()`, so a broken
pattern fails loudly at import time instead of silently at scan time.
`unquoteLiteral()` handles single/double/template quotes, Python
triple-quoted strings, and Go raw backtick strings.

**Per-language plugins** own:
- the tree-sitter grammar import (this is the ONLY place in
  `src/core/group/` where tree-sitter grammars are imported),
- the query S-expressions,
- the `TopicMeta` payload (role, broker, confidence, symbolName) that
  the orchestrator receives back on every match.

Each plugin uses a `@value` capture name to bind the topic literal node.
The JavaScript and TypeScript grammars share AST node names for every
construct we query, so `node.ts` defines the pattern sources once and
compiles them against `JavaScript`, `TypeScript.typescript`, and
`TypeScript.tsx` — exporting three providers because `Parser.Query`
objects are NOT portable across grammar instances.

**Registry (`topic-patterns/index.ts`)** — maps `.java` → Java provider,
`.go` → Go, `.py` → Python, `.js`/`.jsx` → JS, `.ts` → TS, `.tsx` → TSX.
Also exports `TOPIC_SCAN_GLOB` so adding a new language is a single
file-level edit (drop `topic-patterns/<lang>.ts`, import + register it
here — zero edits required in `topic-extractor.ts`).

**Orchestrator (`topic-extractor.ts`)** — ~110 lines, no grammar or
query imports. Per file: `getProviderForFile(rel)` → `scanFile(parser,
provider, content)` → `unquoteLiteral(valueText)` → `makeContract(...)`.
Reuses one `Parser` instance across files; the scanner calls
`setLanguage` per plugin.

## Why this is better than regex

1. **Comments and strings are respected for free.** The old regex
   would match `// kafkaTemplate.send("fake.topic")` as a real
   producer; tree-sitter never visits comments or string literals as
   code nodes, so false positives from commented-out code are
   eliminated.
2. **Struct/object literal patterns are structural, not textual.**
   `sarama.ProducerMessage{Topic: "..."}` no longer needs a 300-char
   lookahead (which was a known cross-match bug partly mitigated by a
   loop regression test in the self-review). The new query matches a
   specific `composite_literal` with a specific `qualified_type` and
   `keyed_element` — exactly one struct literal per match.
3. **No order-of-operations fragility.** Regex for
   `channel.publish` vs `channel.consume` was independent and
   file-wide; the AST scopes matches to the specific `call_expression`.
4. **Language-agnostic extension.** Adding Ruby, Rust, or C# topic
   detection later means dropping one file in `topic-patterns/` — no
   changes to shared scanner or orchestrator, and no tree-sitter
   imports leak into top-level code.

## Per-file fault tolerance

- Malformed files that tree-sitter can't parse are silently skipped
  (`parser.parse` is wrapped by `scanFile`). The ingestion pipeline
  already logs unparseable files at index time.
- A syntactically invalid query is caught at `compilePatterns` time,
  not scan time — broken plugins fail loudly at import.
- Per-pattern `matches()` failures are swallowed so one broken query
  in a plugin doesn't block the rest.

## Tests

All 30 existing `topic-extractor.test.ts` tests pass **without any
changes to the test file** — they were written as input/output contract
tests (given this source file, expect these `ExtractedContract` objects)
and that contract is unchanged. Regression coverage includes:

- Kafka: Java `@KafkaListener` + `kafkaTemplate.send`; Node
  `producer.send` + `consumer.subscribe`; Go sarama producer/consumer
  (sync and async); kafka-go Writer/Reader; Python `KafkaConsumer` +
  `producer.send/produce`
- RabbitMQ: Java `@RabbitListener` + `rabbitTemplate.convertAndSend`;
  Node `channel.consume/publish/sendToQueue`; Python `basic_consume/
  basic_publish` with keyword args
- NATS: Go and Node `nc.Subscribe/Publish`; Go and Node JetStream
  `js.Subscribe/Publish`; Python `await nc.subscribe/publish`

Including the regression test for the sarama `ProducerMessage`
in-loop case — the AST-based query captures every literal in the
file independently, not just the first one after `NewSyncProducer`.

## Neighbor regression check

- `topic-extractor.test.ts` — 30/30 pass (rewritten extractor)
- `http-route-extractor.test.ts` — 18/18 pass (untouched)
- `grpc-extractor.test.ts` — 43/43 pass (untouched)
- `manifest-extractor.test.ts` — 8/8 pass (untouched)
- Full `npx tsc --noEmit` clean

## Scope discipline (per GUARDRAILS.md)

- Only files under `src/core/group/extractors/` are touched; no
  changes to other extractors, tests, MCP surface, or pipeline.ts.
- No CI/release/security config changes, no secrets.
- New tree-sitter imports all reference grammars that are already
  installed as dependencies (`tree-sitter`, `tree-sitter-javascript`,
  `tree-sitter-typescript`, `tree-sitter-python`, `tree-sitter-java`,
  `tree-sitter-go` — all in `package.json` for the existing pipeline).

## Phase 2 / phase 3 plan

- **Phase 2 (next commit):** rewrite `http-route-extractor.ts`
  Strategy B (regex fallback) on the same plugin pattern. Graph-assisted
  Strategy A stays as-is (already uses pipeline-built tree-sitter data
  via `HANDLES_ROUTE` Cypher queries).
- **Phase 3 (commit after):** rewrite `grpc-extractor.ts` for Java /
  Go / Python / TypeScript detection. `.proto` files are the one
  outstanding question — there is no `tree-sitter-proto` grammar
  installed; the in-tree string-sanitizing parser stays as a pragmatic
  exception with a comment, alternative being to add
  `tree-sitter-proto` as a dep (open for the maintainer).

Co-authored-by: Claude <noreply@anthropic.com>
This commit is contained in:
ivkond 2026-04-11 21:41:17 +00:00 • committed by Claude
parent 747937cca1
commit 55ed7443bc
No known key found for this signature in database
8 changed files with 789 additions and 307 deletions

View file

@ -1,13 +1,32 @@
import * as fs from 'node:fs';
import * as path from 'node:path';
import { glob } from 'glob';
import Parser from 'tree-sitter';
import type { ContractExtractor, CypherExecutor } from '../contract-extractor.js';
import type { ExtractedContract, RepoHandle } from '../types.js';
import { scanFile, unquoteLiteral } from './tree-sitter-scanner.js';
import {
TOPIC_SCAN_GLOB,
getProviderForFile,
type Broker,
type TopicMeta,
} from './topic-patterns/index.js';
type Broker = 'kafka' | 'rabbitmq' | 'nats';
const KAFKAJS_CONSUMER_RUN_RE = /consumer\.run\s*\(\s*\{\s*eachMessage:/;
const KAFKAJS_SUBSCRIBE_RE = /consumer\.subscribe\s*\(\s*\{\s*topic:\s*['"]([^'"]+)['"]/g;
/**
* Language-agnostic orchestrator for topic (message broker) contract
* extraction. All grammar-specific knowledge lives in `topic-patterns/*`
* — this file must not import any tree-sitter grammar directly.
*
* Flow per file:
* 1. `getProviderForFile(rel)` → compiled plugin (or `undefined` if the
* file's extension isn't registered, in which case we skip it).
* 2. `scanFile(parser, provider, content)` → list of `{meta, valueText}`
* pairs, one per matched literal.
* 3. `unquoteLiteral(valueText)` → the raw topic string.
* 4. `makeContract(topic, meta, relPath)` → `ExtractedContract`.
*
* Adding a new language is a one-file edit in `topic-patterns/index.ts`.
*/
function readSafe(repoPath: string, rel: string): string | null {
const abs = path.resolve(repoPath, rel);
@ -21,281 +40,23 @@ function readSafe(repoPath: string, rel: string): string | null {
}
}
function makeContract(
topicName: string,
role: 'provider' | 'consumer',
filePath: string,
symbolName: string,
confidence: number,
broker: Broker,
): ExtractedContract {
function makeContract(topicName: string, meta: TopicMeta, filePath: string): ExtractedContract {
return {
contractId: `topic::${topicName}`,
type: 'topic',
role,
role: meta.role,
symbolUid: '',
symbolRef: { filePath: filePath.replace(/\\/g, '/'), name: symbolName },
symbolName,
confidence,
symbolRef: { filePath: filePath.replace(/\\/g, '/'), name: meta.symbolName },
symbolName: meta.symbolName,
confidence: meta.confidence,
meta: {
broker,
broker: meta.broker satisfies Broker,
topicName,
extractionStrategy: 'source_scan',
extractionStrategy: 'tree_sitter',
},
};
}
interface PatternDef {
regex: RegExp;
role: 'provider' | 'consumer';
broker: Broker;
confidence: number;
topicGroup: number;
symbolName: string;
}
// --- Kafka patterns ---
const KAFKA_PATTERNS: PatternDef[] = [
// Java: @KafkaListener(topics = "xxx")
{
regex: /@KafkaListener\s*\(\s*topics\s*=\s*"([^"]+)"/g,
role: 'consumer',
broker: 'kafka',
confidence: 0.8,
topicGroup: 1,
symbolName: 'kafkaListener',
},
// Java: kafkaTemplate.send("xxx"
{
regex: /kafkaTemplate\.send\s*\(\s*"([^"]+)"/gi,
role: 'provider',
broker: 'kafka',
confidence: 0.8,
topicGroup: 1,
symbolName: 'kafkaTemplate.send',
},
// Node: producer.send({ topic: 'xxx'
{
regex: /producer\.send\s*\(\s*\{\s*topic:\s*['"]([^'"]+)['"]/g,
role: 'provider',
broker: 'kafka',
confidence: 0.8,
topicGroup: 1,
symbolName: 'producer.send',
},
// Node: consumer.subscribe({ topic: 'xxx'
{
regex: /consumer\.subscribe\s*\(\s*\{\s*topic:\s*['"]([^'"]+)['"]/g,
role: 'consumer',
broker: 'kafka',
confidence: 0.8,
topicGroup: 1,
symbolName: 'consumer.subscribe',
},
// Go: consumer.ConsumePartition("xxx"
{
regex: /\.ConsumePartition\s*\(\s*"([^"]+)"/g,
role: 'consumer',
broker: 'kafka',
confidence: 0.7,
topicGroup: 1,
symbolName: 'ConsumePartition',
},
// Python: KafkaConsumer('xxx'
{
regex: /KafkaConsumer\s*\(\s*['"]([^'"]+)['"]/g,
role: 'consumer',
broker: 'kafka',
confidence: 0.7,
topicGroup: 1,
symbolName: 'KafkaConsumer',
},
// Python: producer.send('xxx' or producer.produce('xxx'
{
regex: /producer\.(?:send|produce)\s*\(\s*['"]([^'"]+)['"]/g,
role: 'provider',
broker: 'kafka',
confidence: 0.7,
topicGroup: 1,
symbolName: 'producer.send',
},
// Go: sarama.ProducerMessage{Topic: "xxx"} struct literal (emitted by
// both NewSyncProducer and NewAsyncProducer client code paths).
//
// Previous pattern was `sarama.NewSyncProducer[\s\S]{0,300}?Topic:...`
// which anchored to the producer constructor and used a 300-char
// lookahead. In a loop like
// producer := sarama.NewSyncProducer(...)
// for _, item := range items {
// msg1 := &sarama.ProducerMessage{Topic: "order.created"}
// msg2 := &sarama.ProducerMessage{Topic: "order.shipped"}
// }
// the regex captured only "order.created" (first Topic after the
// constructor) and silently missed "order.shipped". Matching on the
// struct literal directly fixes both the false negative in loops and
// the spurious cross-message capture when multiple unrelated messages
// sit within 300 chars of the constructor.
{
regex: /sarama\.ProducerMessage\s*\{[\s\S]{0,200}?Topic:\s*"([^"]+)"/g,
role: 'provider',
broker: 'kafka',
confidence: 0.75,
topicGroup: 1,
symbolName: 'sarama.ProducerMessage',
},
// Go: kafka-go writer construction. kafka-go does NOT wrap messages in
// a struct with a Topic field (the writer owns the topic), so we match
// the Writer itself. A 200-char window bridges the gap between
// `kafka.NewWriter(...)` / `kafka.Writer{` and the Topic field inside
// the config literal — kafka-go writer configs are small and rarely
// contain more than one Topic field, so the risk of cross-message
// capture is low here.
{
regex: /kafka\.(?:NewWriter|Writer)\b[\s\S]{0,200}?Topic:\s*"([^"]+)"/g,
role: 'provider',
broker: 'kafka',
confidence: 0.75,
topicGroup: 1,
symbolName: 'kafka.Writer',
},
// Go: kafka-go reader construction, mirrors Writer above.
{
regex: /kafka\.(?:NewReader|Reader)\b[\s\S]{0,200}?Topic:\s*"([^"]+)"/g,
role: 'consumer',
broker: 'kafka',
confidence: 0.75,
topicGroup: 1,
symbolName: 'kafka.Reader',
},
];
// --- RabbitMQ patterns ---
const RABBITMQ_PATTERNS: PatternDef[] = [
// Java: @RabbitListener(queues = "xxx")
{
regex: /@RabbitListener\s*\(\s*queues\s*=\s*"([^"]+)"/g,
role: 'consumer',
broker: 'rabbitmq',
confidence: 0.8,
topicGroup: 1,
symbolName: 'rabbitListener',
},
// Java: rabbitTemplate.convertAndSend("xxx"
{
regex: /rabbitTemplate\.convertAndSend\s*\(\s*"([^"]+)"/gi,
role: 'provider',
broker: 'rabbitmq',
confidence: 0.8,
topicGroup: 1,
symbolName: 'rabbitTemplate.convertAndSend',
},
// Node: channel.consume("xxx"
{
regex: /channel\.consume\s*\(\s*"([^"]+)"/g,
role: 'consumer',
broker: 'rabbitmq',
confidence: 0.8,
topicGroup: 1,
symbolName: 'channel.consume',
},
// Node: channel.publish("xxx"
{
regex: /channel\.publish\s*\(\s*"([^"]+)"/g,
role: 'provider',
broker: 'rabbitmq',
confidence: 0.8,
topicGroup: 1,
symbolName: 'channel.publish',
},
// Node: channel.sendToQueue("xxx"
{
regex: /channel\.sendToQueue\s*\(\s*"([^"]+)"/g,
role: 'provider',
broker: 'rabbitmq',
confidence: 0.8,
topicGroup: 1,
symbolName: 'channel.sendToQueue',
},
// Python: channel.basic_consume(queue='xxx'
{
regex: /channel\.basic_consume\s*\(\s*queue\s*=\s*['"]([^'"]+)['"]/g,
role: 'consumer',
broker: 'rabbitmq',
confidence: 0.7,
topicGroup: 1,
symbolName: 'basic_consume',
},
// Python: channel.basic_publish(exchange='xxx'
{
regex: /channel\.basic_publish\s*\([^)]*exchange\s*=\s*['"]([^'"]+)['"]/g,
role: 'provider',
broker: 'rabbitmq',
confidence: 0.7,
topicGroup: 1,
symbolName: 'basic_publish',
},
];
// --- NATS patterns ---
const NATS_PATTERNS: PatternDef[] = [
// Go/Node: nc.Subscribe("xxx" or nc.subscribe("xxx"
{
regex: /nc\.(?:S|s)ubscribe\s*\(\s*"([^"]+)"/g,
role: 'consumer',
broker: 'nats',
confidence: 0.8,
topicGroup: 1,
symbolName: 'nc.Subscribe',
},
// Go/Node: nc.Publish("xxx" or nc.publish("xxx"
{
regex: /nc\.(?:P|p)ublish\s*\(\s*"([^"]+)"/g,
role: 'provider',
broker: 'nats',
confidence: 0.8,
topicGroup: 1,
symbolName: 'nc.Publish',
},
// Go/Node JetStream: js.Subscribe("xxx"
{
regex: /js\.(?:S|s)ubscribe\s*\(\s*"([^"]+)"/g,
role: 'consumer',
broker: 'nats',
confidence: 0.8,
topicGroup: 1,
symbolName: 'js.Subscribe',
},
// Go/Node JetStream: js.Publish("xxx"
{
regex: /js\.(?:P|p)ublish\s*\(\s*"([^"]+)"/g,
role: 'provider',
broker: 'nats',
confidence: 0.8,
topicGroup: 1,
symbolName: 'js.Publish',
},
// Python: await nc.subscribe("xxx")
{
regex: /await\s+nc\.subscribe\s*\(\s*['"]([^'"]+)['"]/g,
role: 'consumer',
broker: 'nats',
confidence: 0.75,
topicGroup: 1,
symbolName: 'nc.subscribe',
},
// Python: await nc.publish("xxx")
{
regex: /await\s+nc\.publish\s*\(\s*['"]([^'"]+)['"]/g,
role: 'provider',
broker: 'nats',
confidence: 0.75,
topicGroup: 1,
symbolName: 'nc.publish',
},
];
const ALL_PATTERNS: PatternDef[] = [...KAFKA_PATTERNS, ...RABBITMQ_PATTERNS, ...NATS_PATTERNS];
export class TopicExtractor implements ContractExtractor {
type = 'topic' as const;
@ -308,59 +69,37 @@ export class TopicExtractor implements ContractExtractor {
repoPath: string,
_repo: RepoHandle,
): Promise<ExtractedContract[]> {
const files = await glob('**/*.{ts,tsx,js,jsx,java,go,py}', {
const files = await glob(TOPIC_SCAN_GLOB, {
cwd: repoPath,
ignore: ['**/node_modules/**', '**/.git/**', '**/vendor/**', '**/dist/**', '**/build/**'],
nodir: true,
});
// One parser reused across files; the scanner calls `setLanguage` per
// file based on which plugin the registry returns.
const parser = new Parser();
const out: ExtractedContract[] = [];
for (const rel of files) {
if (rel.endsWith('_test.go')) continue;
const provider = getProviderForFile(rel);
if (!provider) continue;
const content = readSafe(repoPath, rel);
if (!content) continue;
out.push(...this.scanFile(content, rel));
const matches = scanFile(parser, provider, content);
for (const match of matches) {
const topicName = unquoteLiteral(match.valueText);
if (!topicName) continue;
out.push(makeContract(topicName, match.meta, rel));
}
}
return this.dedupe(out);
}
private scanFile(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
for (const pattern of ALL_PATTERNS) {
// Reset regex state for each file
const re = new RegExp(pattern.regex.source, pattern.regex.flags);
let m: RegExpExecArray | null;
while ((m = re.exec(content)) !== null) {
const topicName = m[pattern.topicGroup];
if (!topicName) continue;
out.push(
makeContract(
topicName,
pattern.role,
filePath,
pattern.symbolName,
pattern.confidence,
pattern.broker,
),
);
}
}
if (KAFKAJS_CONSUMER_RUN_RE.test(content)) {
const subscribeRe = new RegExp(KAFKAJS_SUBSCRIBE_RE.source, KAFKAJS_SUBSCRIBE_RE.flags);
let subscribeMatch: RegExpExecArray | null;
while ((subscribeMatch = subscribeRe.exec(content)) !== null) {
const topicName = subscribeMatch[1];
if (!topicName) continue;
out.push(makeContract(topicName, 'consumer', filePath, 'consumer.run', 0.75, 'kafka'));
}
}
return out;
}
private dedupe(items: ExtractedContract[]): ExtractedContract[] {
const seen = new Set<string>();
const out: ExtractedContract[] = [];

View file

@ -0,0 +1,123 @@
import Go from 'tree-sitter-go';
import { compilePatterns, type LanguagePatterns } from '../tree-sitter-scanner.js';
import type { TopicMeta } from './types.js';
/**
* Go topic extraction patterns.
*
* Detects Sarama, segmentio/kafka-go and nats.go producer/consumer APIs:
* - `X.ConsumePartition("topic", ...)`
* - `sarama.ProducerMessage{Topic: "xxx"}`
* - `kafka.Writer{Topic: "xxx"}` / `kafka.WriterConfig{Topic: ...}`
* - `kafka.Reader{Topic: "xxx"}` / `kafka.ReaderConfig{Topic: ...}`
* - `nc.Subscribe("topic", ...)` / `js.Subscribe("topic", ...)`
* - `nc.Publish("topic", ...)` / `js.Publish("topic", ...)`
*
* Every query MUST bind `@value` to the topic literal node.
*/
const GO_TOPIC_SPEC: LanguagePatterns<TopicMeta> = {
name: 'go-topic',
language: Go,
patterns: [
{
meta: {
role: 'consumer',
broker: 'kafka',
confidence: 0.7,
symbolName: 'ConsumePartition',
},
query: `
(call_expression
function: (selector_expression
field: (field_identifier) @method (#eq? @method "ConsumePartition"))
arguments: (argument_list . (interpreted_string_literal) @value))
`,
},
{
meta: {
role: 'provider',
broker: 'kafka',
confidence: 0.75,
symbolName: 'sarama.ProducerMessage',
},
query: `
(composite_literal
type: (qualified_type
package: (package_identifier) @pkg (#eq? @pkg "sarama")
name: (type_identifier) @ty (#eq? @ty "ProducerMessage"))
body: (literal_value
(keyed_element
(literal_element (identifier) @field (#eq? @field "Topic"))
(literal_element (interpreted_string_literal) @value))))
`,
},
{
meta: {
role: 'provider',
broker: 'kafka',
confidence: 0.75,
symbolName: 'kafka.Writer',
},
query: `
(composite_literal
type: (qualified_type
package: (package_identifier) @pkg (#eq? @pkg "kafka")
name: (type_identifier) @ty (#match? @ty "^(Writer|WriterConfig)$"))
body: (literal_value
(keyed_element
(literal_element (identifier) @field (#eq? @field "Topic"))
(literal_element (interpreted_string_literal) @value))))
`,
},
{
meta: {
role: 'consumer',
broker: 'kafka',
confidence: 0.75,
symbolName: 'kafka.Reader',
},
query: `
(composite_literal
type: (qualified_type
package: (package_identifier) @pkg (#eq? @pkg "kafka")
name: (type_identifier) @ty (#match? @ty "^(Reader|ReaderConfig)$"))
body: (literal_value
(keyed_element
(literal_element (identifier) @field (#eq? @field "Topic"))
(literal_element (interpreted_string_literal) @value))))
`,
},
{
meta: {
role: 'consumer',
broker: 'nats',
confidence: 0.8,
symbolName: 'nc.Subscribe',
},
query: `
(call_expression
function: (selector_expression
operand: (identifier) @obj (#match? @obj "^(nc|js)$")
field: (field_identifier) @method (#match? @method "^[Ss]ubscribe$"))
arguments: (argument_list . (interpreted_string_literal) @value))
`,
},
{
meta: {
role: 'provider',
broker: 'nats',
confidence: 0.8,
symbolName: 'nc.Publish',
},
query: `
(call_expression
function: (selector_expression
operand: (identifier) @obj (#match? @obj "^(nc|js)$")
field: (field_identifier) @method (#match? @method "^[Pp]ublish$"))
arguments: (argument_list . (interpreted_string_literal) @value))
`,
},
],
};
export const GO_TOPIC_PROVIDER = compilePatterns(GO_TOPIC_SPEC);

View file

@ -0,0 +1,49 @@
import * as path from 'node:path';
import type { CompiledPatterns } from '../tree-sitter-scanner.js';
import type { TopicMeta } from './types.js';
import { JAVA_TOPIC_PROVIDER } from './java.js';
import { GO_TOPIC_PROVIDER } from './go.js';
import { PYTHON_TOPIC_PROVIDER } from './python.js';
import {
JAVASCRIPT_TOPIC_PROVIDER,
TYPESCRIPT_TOPIC_PROVIDER,
TSX_TOPIC_PROVIDER,
} from './node.js';
export type { TopicMeta, Broker } from './types.js';
/**
* File-extension → compiled-plugin registry for topic extraction. The
* top-level orchestrator (`topic-extractor.ts`) looks up the plugin for
* each file it visits and delegates the scanning to `tree-sitter-scanner`.
*
* Keys are lowercase extensions including the leading dot. To add a new
* language, drop a `topic-patterns/<lang>.ts` that exports a compiled
* provider, import it here and register the extension(s). No edits to
* `topic-extractor.ts` are required.
*/
const REGISTRY: Record<string, CompiledPatterns<TopicMeta>> = {
'.java': JAVA_TOPIC_PROVIDER,
'.go': GO_TOPIC_PROVIDER,
'.py': PYTHON_TOPIC_PROVIDER,
'.js': JAVASCRIPT_TOPIC_PROVIDER,
'.jsx': JAVASCRIPT_TOPIC_PROVIDER,
'.ts': TYPESCRIPT_TOPIC_PROVIDER,
'.tsx': TSX_TOPIC_PROVIDER,
};
/**
* Glob pattern for files worth scanning. Kept here so adding a new
* language to the registry also widens the glob automatically via a
* single edit.
*/
export const TOPIC_SCAN_GLOB = '**/*.{ts,tsx,js,jsx,java,go,py}';
/**
* Return the compiled provider registered for the given file's
* extension, or `undefined` if the extension is not registered.
*/
export function getProviderForFile(rel: string): CompiledPatterns<TopicMeta> | undefined {
const ext = path.extname(rel).toLowerCase();
return REGISTRY[ext];
}

View file

@ -0,0 +1,83 @@
import Java from 'tree-sitter-java';
import { compilePatterns, type LanguagePatterns } from '../tree-sitter-scanner.js';
import type { TopicMeta } from './types.js';
/**
* Java topic extraction patterns.
*
* Detects Kafka and RabbitMQ (Spring conventions) producer/consumer APIs:
* - `@KafkaListener(topics = "xxx")`
* - `@RabbitListener(queues = "xxx")`
* - `kafkaTemplate.send("xxx", ...)`
* - `rabbitTemplate.convertAndSend("xxx", ...)`
*
* Every query MUST bind `@value` to the topic literal node.
*/
const JAVA_TOPIC_SPEC: LanguagePatterns<TopicMeta> = {
name: 'java-topic',
language: Java,
patterns: [
{
meta: {
role: 'consumer',
broker: 'kafka',
confidence: 0.8,
symbolName: 'kafkaListener',
},
query: `
(annotation
name: (identifier) @name (#eq? @name "KafkaListener")
arguments: (annotation_argument_list
(element_value_pair
key: (identifier) @key (#eq? @key "topics")
value: (string_literal) @value)))
`,
},
{
meta: {
role: 'consumer',
broker: 'rabbitmq',
confidence: 0.8,
symbolName: 'rabbitListener',
},
query: `
(annotation
name: (identifier) @name (#eq? @name "RabbitListener")
arguments: (annotation_argument_list
(element_value_pair
key: (identifier) @key (#eq? @key "queues")
value: (string_literal) @value)))
`,
},
{
meta: {
role: 'provider',
broker: 'kafka',
confidence: 0.8,
symbolName: 'kafkaTemplate.send',
},
query: `
(method_invocation
object: (identifier) @obj (#eq? @obj "kafkaTemplate")
name: (identifier) @method (#eq? @method "send")
arguments: (argument_list . (string_literal) @value))
`,
},
{
meta: {
role: 'provider',
broker: 'rabbitmq',
confidence: 0.8,
symbolName: 'rabbitTemplate.convertAndSend',
},
query: `
(method_invocation
object: (identifier) @obj (#eq? @obj "rabbitTemplate")
name: (identifier) @method (#eq? @method "convertAndSend")
arguments: (argument_list . (string_literal) @value))
`,
},
],
};
export const JAVA_TOPIC_PROVIDER = compilePatterns(JAVA_TOPIC_SPEC);

View file

@ -0,0 +1,165 @@
import JavaScript from 'tree-sitter-javascript';
import TypeScript from 'tree-sitter-typescript';
import {
compilePatterns,
type LanguagePatterns,
type PatternSpec,
} from '../tree-sitter-scanner.js';
import type { TopicMeta } from './types.js';
/**
* Node.js / TypeScript topic extraction patterns.
*
* Detects kafkajs, amqplib (RabbitMQ), and nats.js producer/consumer APIs:
* - `producer.send({ topic: 'xxx', ... })` (kafkajs)
* - `consumer.subscribe({ topic: 'xxx', ... })` (kafkajs)
* - `channel.consume("queue", ...)` / `channel.publish(...)` / `channel.sendToQueue(...)`
* - `nc.subscribe("topic")` / `js.subscribe("topic")`
* - `nc.publish("topic", ...)` / `js.publish("topic", ...)`
*
* The JavaScript and TypeScript tree-sitter grammars share node type
* names for every construct we query here, so the pattern sources are
* defined once and compiled against each grammar variant. We export three
* providers because Parser.Query objects are NOT portable across grammar
* instances — `.js` files use the JavaScript grammar, `.ts` uses
* TypeScript.typescript, and `.tsx` uses TypeScript.tsx.
*
* Every query MUST bind `@value` to the topic literal node.
*/
const NODE_TOPIC_PATTERNS: PatternSpec<TopicMeta>[] = [
{
meta: {
role: 'provider',
broker: 'kafka',
confidence: 0.8,
symbolName: 'producer.send',
},
query: `
(call_expression
function: (member_expression
object: (identifier) @obj (#eq? @obj "producer")
property: (property_identifier) @prop (#eq? @prop "send"))
arguments: (arguments
(object
(pair
key: (property_identifier) @key (#eq? @key "topic")
value: [(string) (template_string)] @value))))
`,
},
{
meta: {
role: 'consumer',
broker: 'kafka',
confidence: 0.8,
symbolName: 'consumer.subscribe',
},
query: `
(call_expression
function: (member_expression
object: (identifier) @obj (#eq? @obj "consumer")
property: (property_identifier) @prop (#eq? @prop "subscribe"))
arguments: (arguments
(object
(pair
key: (property_identifier) @key (#eq? @key "topic")
value: [(string) (template_string)] @value))))
`,
},
{
meta: {
role: 'consumer',
broker: 'rabbitmq',
confidence: 0.8,
symbolName: 'channel.consume',
},
query: `
(call_expression
function: (member_expression
object: (identifier) @obj (#eq? @obj "channel")
property: (property_identifier) @prop (#eq? @prop "consume"))
arguments: (arguments . [(string) (template_string)] @value))
`,
},
{
meta: {
role: 'provider',
broker: 'rabbitmq',
confidence: 0.8,
symbolName: 'channel.publish',
},
query: `
(call_expression
function: (member_expression
object: (identifier) @obj (#eq? @obj "channel")
property: (property_identifier) @prop (#eq? @prop "publish"))
arguments: (arguments . [(string) (template_string)] @value))
`,
},
{
meta: {
role: 'provider',
broker: 'rabbitmq',
confidence: 0.8,
symbolName: 'channel.sendToQueue',
},
query: `
(call_expression
function: (member_expression
object: (identifier) @obj (#eq? @obj "channel")
property: (property_identifier) @prop (#eq? @prop "sendToQueue"))
arguments: (arguments . [(string) (template_string)] @value))
`,
},
{
meta: {
role: 'consumer',
broker: 'nats',
confidence: 0.8,
symbolName: 'nc.subscribe',
},
query: `
(call_expression
function: (member_expression
object: (identifier) @obj (#match? @obj "^(nc|js)$")
property: (property_identifier) @prop (#match? @prop "^[Ss]ubscribe$"))
arguments: (arguments . [(string) (template_string)] @value))
`,
},
{
meta: {
role: 'provider',
broker: 'nats',
confidence: 0.8,
symbolName: 'nc.publish',
},
query: `
(call_expression
function: (member_expression
object: (identifier) @obj (#match? @obj "^(nc|js)$")
property: (property_identifier) @prop (#match? @prop "^[Pp]ublish$"))
arguments: (arguments . [(string) (template_string)] @value))
`,
},
];
const JAVASCRIPT_TOPIC_SPEC: LanguagePatterns<TopicMeta> = {
name: 'javascript-topic',
language: JavaScript,
patterns: NODE_TOPIC_PATTERNS,
};
const TYPESCRIPT_TOPIC_SPEC: LanguagePatterns<TopicMeta> = {
name: 'typescript-topic',
language: TypeScript.typescript,
patterns: NODE_TOPIC_PATTERNS,
};
const TSX_TOPIC_SPEC: LanguagePatterns<TopicMeta> = {
name: 'tsx-topic',
language: TypeScript.tsx,
patterns: NODE_TOPIC_PATTERNS,
};
export const JAVASCRIPT_TOPIC_PROVIDER = compilePatterns(JAVASCRIPT_TOPIC_SPEC);
export const TYPESCRIPT_TOPIC_PROVIDER = compilePatterns(TYPESCRIPT_TOPIC_SPEC);
export const TSX_TOPIC_PROVIDER = compilePatterns(TSX_TOPIC_SPEC);

View file

@ -0,0 +1,119 @@
import Python from 'tree-sitter-python';
import { compilePatterns, type LanguagePatterns } from '../tree-sitter-scanner.js';
import type { TopicMeta } from './types.js';
/**
* Python topic extraction patterns.
*
* Detects kafka-python, pika (RabbitMQ), and nats-py producer/consumer APIs:
* - `KafkaConsumer('topic', ...)`
* - `producer.send('topic', ...)` / `producer.produce('topic', ...)`
* - `channel.basic_consume(queue='xxx', ...)`
* - `channel.basic_publish(exchange='xxx', ...)`
* - `await nc.subscribe('topic')`
* - `await nc.publish('topic', ...)`
*
* Every query MUST bind `@value` to the topic literal node.
*/
const PYTHON_TOPIC_SPEC: LanguagePatterns<TopicMeta> = {
name: 'python-topic',
language: Python,
patterns: [
{
meta: {
role: 'consumer',
broker: 'kafka',
confidence: 0.7,
symbolName: 'KafkaConsumer',
},
query: `
(call
function: (identifier) @func (#eq? @func "KafkaConsumer")
arguments: (argument_list . (string) @value))
`,
},
{
meta: {
role: 'provider',
broker: 'kafka',
confidence: 0.7,
symbolName: 'producer.send',
},
query: `
(call
function: (attribute
object: (identifier) @obj (#eq? @obj "producer")
attribute: (identifier) @method (#match? @method "^(send|produce)$"))
arguments: (argument_list . (string) @value))
`,
},
{
meta: {
role: 'consumer',
broker: 'rabbitmq',
confidence: 0.7,
symbolName: 'basic_consume',
},
query: `
(call
function: (attribute
object: (identifier) @obj (#eq? @obj "channel")
attribute: (identifier) @method (#eq? @method "basic_consume"))
arguments: (argument_list
(keyword_argument
name: (identifier) @kw (#eq? @kw "queue")
value: (string) @value)))
`,
},
{
meta: {
role: 'provider',
broker: 'rabbitmq',
confidence: 0.7,
symbolName: 'basic_publish',
},
query: `
(call
function: (attribute
object: (identifier) @obj (#eq? @obj "channel")
attribute: (identifier) @method (#eq? @method "basic_publish"))
arguments: (argument_list
(keyword_argument
name: (identifier) @kw (#eq? @kw "exchange")
value: (string) @value)))
`,
},
{
meta: {
role: 'consumer',
broker: 'nats',
confidence: 0.75,
symbolName: 'nc.subscribe',
},
query: `
(call
function: (attribute
object: (identifier) @obj (#eq? @obj "nc")
attribute: (identifier) @method (#eq? @method "subscribe"))
arguments: (argument_list . (string) @value))
`,
},
{
meta: {
role: 'provider',
broker: 'nats',
confidence: 0.75,
symbolName: 'nc.publish',
},
query: `
(call
function: (attribute
object: (identifier) @obj (#eq? @obj "nc")
attribute: (identifier) @method (#eq? @method "publish"))
arguments: (argument_list . (string) @value))
`,
},
],
};
export const PYTHON_TOPIC_PROVIDER = compilePatterns(PYTHON_TOPIC_SPEC);

View file

@ -0,0 +1,27 @@
/**
* Shared types for the topic-extractor language plugins.
*
* Each plugin lives in its own file (java.ts, go.ts, ...) and owns the
* tree-sitter grammar import + query sources. The top-level
* `topic-extractor.ts` orchestrator only knows about this type module and
* the plugin registry (`./index.ts`). It MUST NOT import any grammar or
* query text directly — that's the whole point of the split.
*/
export type Broker = 'kafka' | 'rabbitmq' | 'nats';
/**
* Per-pattern payload every topic plugin attaches to its query. Whatever
* the pattern matches, the orchestrator receives this object verbatim
* and uses it to build an `ExtractedContract`.
*
* Plugins produce one `TopicMeta` per pattern (not per match) because a
* single query uniquely identifies its broker/role/confidence triple.
*/
export interface TopicMeta {
role: 'provider' | 'consumer';
broker: Broker;
confidence: number;
/** Short human-readable label of the API being detected. */
symbolName: string;
}

View file

@ -0,0 +1,177 @@
import Parser from 'tree-sitter';
/**
* Shared, language-agnostic tree-sitter scanning utilities used by group
* extractors (topic, http, grpc, ...).
*
* Design goals:
* - The top-level extractors must not import any tree-sitter grammar.
* - Per-language plugins own their grammar import, their query sources,
* and the mapping from capture → meta.
* - This module provides the plumbing: compile queries once per plugin,
* parse a file with a given grammar, run all patterns, and return the
* captured `string_literal`-style nodes together with the plugin's meta.
*/
/**
* One pattern owned by a language plugin. Each pattern owns a tree-sitter
* S-expression query. The query MUST contain a capture named `@value`
* whose node text is the literal we want to extract (string/template/etc).
*
* `TMeta` is the plugin-specific payload the orchestrator receives back
* when this pattern matches — e.g. for topic extraction it carries the
* broker name, role, confidence, symbol name.
*/
export interface PatternSpec<TMeta> {
/** Tree-sitter S-expression. MUST contain a `@value` capture. */
query: string;
/** Plugin-specific payload returned on every match. */
meta: TMeta;
}
/**
* A set of patterns owned by one language plugin, bound to a specific
* tree-sitter grammar.
*
* `language` is typed as `unknown` because tree-sitter's TypeScript
* declarations use `any` for the grammar object, and the grammar modules
* export different shapes (plain grammar vs. namespace with `typescript`
* / `tsx` members). Callers pass the concrete grammar object; this
* module forwards it to `parser.setLanguage` / `new Parser.Query`.
*/
export interface LanguagePatterns<TMeta> {
/** Human-readable plugin name for diagnostics. */
name: string;
/** tree-sitter grammar object. */
language: unknown;
/** Patterns authored against `language`. */
patterns: PatternSpec<TMeta>[];
}
/**
* Compiled form of a `LanguagePatterns` bundle. Queries are compiled
* eagerly at module load time so a broken grammar/query pair fails
* loudly the first time the plugin is imported, instead of silently
* at scan time when no contract is produced.
*/
export interface CompiledPatterns<TMeta> {
name: string;
language: unknown;
patterns: CompiledPattern<TMeta>[];
}
export interface CompiledPattern<TMeta> {
query: Parser.Query;
meta: TMeta;
}
/**
* One match returned by `scanFile`. The orchestrator receives the raw
* literal text (still including any surrounding quotes) together with
* the plugin meta, and is responsible for calling `unquoteLiteral` /
* emitting a domain object (ExtractedContract, Route, ...).
*/
export interface ScanMatch<TMeta> {
meta: TMeta;
/** The node captured as `@value` (the literal). */
valueNode: Parser.SyntaxNode;
/** Raw text of the captured value node — caller must unquote. */
valueText: string;
}
/**
* Compile a LanguagePatterns bundle. Call this once per plugin, at
* module load time, and export the result. Throws if any pattern
* fails to compile against the grammar — that's a bug in the plugin
* author's query, not a runtime condition.
*/
export function compilePatterns<TMeta>(bundle: LanguagePatterns<TMeta>): CompiledPatterns<TMeta> {
const compiled: CompiledPattern<TMeta>[] = [];
for (const spec of bundle.patterns) {
try {
const query = new Parser.Query(bundle.language, spec.query);
compiled.push({ query, meta: spec.meta });
} catch (err) {
const message = err instanceof Error ? err.message : String(err);
throw new Error(
`[tree-sitter-scanner] Failed to compile pattern in ${bundle.name}: ${message}\n` +
`Query source:\n${spec.query}`,
);
}
}
return { name: bundle.name, language: bundle.language, patterns: compiled };
}
/**
* Parse `content` as source code of the plugin's language and run every
* compiled pattern against the resulting AST. Returns one `ScanMatch` per
* matched `@value` capture, carrying the plugin's meta payload.
*
* Errors are swallowed at the file level (malformed file must not abort
* the whole extract). Individual pattern failures are swallowed too so
* a single unusable query doesn't block the rest of the plugin.
*/
export function scanFile<TMeta>(
parser: Parser,
plugin: CompiledPatterns<TMeta>,
content: string,
): ScanMatch<TMeta>[] {
const out: ScanMatch<TMeta>[] = [];
let tree: Parser.Tree;
try {
parser.setLanguage(plugin.language);
tree = parser.parse(content);
} catch {
return out;
}
for (const compiled of plugin.patterns) {
let matches: Parser.QueryMatch[];
try {
matches = compiled.query.matches(tree.rootNode);
} catch {
continue;
}
for (const match of matches) {
const valueCapture = match.captures.find((c) => c.name === 'value');
if (!valueCapture) continue;
out.push({
meta: compiled.meta,
valueNode: valueCapture.node,
valueText: valueCapture.node.text,
});
}
}
return out;
}
/**
* Strip enclosing quotes from a tree-sitter string literal node's text.
* Handles single / double / template quotes, Python triple-quoted strings,
* and Go raw string literals (backticks).
*
* Returns null for empty/nullish input so callers can uniformly skip
* captures whose value is missing.
*/
export function unquoteLiteral(raw: string): string | null {
if (!raw) return null;
// Python triple-quoted
if (
(raw.startsWith('"""') && raw.endsWith('"""')) ||
(raw.startsWith("'''") && raw.endsWith("'''"))
) {
return raw.slice(3, -3);
}
const first = raw[0];
const last = raw[raw.length - 1];
if ((first === '"' || first === "'" || first === '`') && last === first && raw.length >= 2) {
return raw.slice(1, -1);
}
// Some grammars expose the string content without quotes already (e.g.
// Python `string_content` child). Return as-is.
return raw;
}