diff --git a/.cursor/plans/enhance_523ca41c.plan.md b/.cursor/plans/enhance_523ca41c.plan.md new file mode 100644 index 000000000..471e13f67 --- /dev/null +++ b/.cursor/plans/enhance_523ca41c.plan.md @@ -0,0 +1,239 @@ +--- +name: Enhance +overview: Restructure GitNexus LLM tools to leverage clusters and processes for better code understanding. Remove unused highlight tool, add new tools (explore, overview), enhance existing tools with cluster/process context, and improve impact analysis reliability. +todos: [] +--- + +# Enhanced LLM Tools with Cluster and Process Integration + +## Summary + +Consolidate GitNexus from 6 tools to **7 focused tools** that leverage the pre-computed clusters (Communities) and processes for richer context. Remove the highlight tool, add `explore` and `overview` tools, and enhance `search` and `blastRadius` with cluster/process awareness. + +## Final Tool Set + +| Tool | Status | Purpose ||------|--------|---------|| `search` | Enhance | Hybrid search + group results by process/cluster || `grep` | Keep | Regex pattern search || `read` | Keep | Read file content || `explore` | **New** | Deep dive on one symbol, cluster, or process || `overview` | **New** | Codebase map (all clusters + all processes) || `impact` | Enhance | Rename from blastRadius, add process/cluster context, increase limits || `cypher` | Keep | Raw graph queries || `highlight` | **Remove** | No longer needed | + +## Architecture + +```mermaid +flowchart TD + subgraph tools [LLM Tools Layer] + search[search] + grep[grep] + read[read] + explore[explore] + overview[overview] + impact[impact] + cypher[cypher] + end + + subgraph graph [Knowledge Graph] + nodes[Nodes: File, Function, Class...] + communities[Community Nodes] + processes[Process Nodes] + edges[CodeRelation Edges] + memberOf[MEMBER_OF Edges] + stepIn[STEP_IN_PROCESS Edges] + end + + search --> edges + search --> communities + search --> processes + explore --> communities + explore --> processes + explore --> memberOf + explore --> stepIn + overview --> communities + overview --> processes + impact --> edges + impact --> communities + impact --> processes + cypher --> graph +``` + + + +## File Changes + +### 1. Remove Highlight Tool + +**File:** [gitnexus/src/core/llm/tools.ts](gitnexus/src/core/llm/tools.ts) + +- Delete the `highlightTool` definition (lines ~395-414) +- Remove `highlightTool` from the returned array (line ~862) +- Remove highlight marker logic from `blastRadius` output (line ~814-816) + +**File:** [gitnexus/src/core/llm/agent.ts](gitnexus/src/core/llm/agent.ts) + +- Remove highlight references from system prompt (lines 70, 77) +- Update tool list in prompt to reflect new tools + +**File:** [gitnexus/src/core/llm/types.ts](gitnexus/src/core/llm/types.ts) + +- Remove `'highlight'` from `AgentStreamChunk.type` union (line 180) +- Remove `highlightNodeIds` property (line 187-188) + +### 2. Add `explore` Tool + +**File:** [gitnexus/src/core/llm/tools.ts](gitnexus/src/core/llm/tools.ts)New tool that auto-detects target type and returns comprehensive context: + +```typescript +explore({ + target: string, // Name of symbol, cluster, or process + type?: 'symbol' | 'cluster' | 'process' // Optional, auto-detected +}) +``` + +**Functionality:** + +- For symbols: Query node, get MEMBER_OF cluster, get STEP_IN_PROCESS processes, get 1-hop connections +- For clusters: Query Community node, get members via MEMBER_OF, get processes that touch this cluster +- For processes: Query Process node, get steps via STEP_IN_PROCESS with step order, get clusters touched + +**Cypher queries needed:** + +```cypher +-- Symbol cluster membership +MATCH (s {name: $name})-[:CodeRelation {type: 'MEMBER_OF'}]->(c:Community) +RETURN c.label, c.description + +-- Symbol process participation +MATCH (s {name: $name})-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process) +RETURN p.label, r.step, p.stepCount + +-- Process steps in order +MATCH (s)-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process {id: $processId}) +RETURN s.name, s.filePath, r.step +ORDER BY r.step +``` + + + +### 3. Add `overview` Tool + +**File:** [gitnexus/src/core/llm/tools.ts](gitnexus/src/core/llm/tools.ts)New tool that returns codebase structure: + +```typescript +overview() // No parameters +``` + +**Functionality:** + +- Query all Community nodes with member counts +- Query all Process nodes with step counts and types +- Calculate cluster dependencies (cross-cluster CALLS) +- Identify critical paths (most connected processes) + +**Output format:** + +```javascript +CLUSTERS (N total): +| Cluster | Symbols | Cohesion | Description | +... + +PROCESSES (N total): +| Process | Steps | Type | Clusters | +... + +CRITICAL PATHS: +- LoginFlow (45 edges) +... +``` + + + +### 4. Enhance `search` Tool + +**File:** [gitnexus/src/core/llm/tools.ts](gitnexus/src/core/llm/tools.ts)Modify existing search to group results by process:**Current:** Returns flat list with 1-hop connections**Enhanced:** Groups results by process, adds cluster context**Changes:** + +- After hybrid search, query STEP_IN_PROCESS for each result +- Group results by process ID +- Sort processes by number of matching results (relevance) +- Add cluster label for each result via MEMBER_OF query +- Keep 1-hop connections as optional detail + +**New parameter:** + +```typescript +search({ + query: string, + groupByProcess?: boolean, // Default: true + limit?: number +}) +``` + + + +### 5. Enhance `impact` Tool (rename from blastRadius) + +**File:** [gitnexus/src/core/llm/tools.ts](gitnexus/src/core/llm/tools.ts)**Rename:** `blastRadiusTool` to `impactTool`**Enhancements:** + +1. Increase LIMIT clauses: 100 to 300 (depth 1), 100 to 200 (depth 2), 50 to 100 (depth 3) +2. Add affected processes section (query STEP_IN_PROCESS for all affected symbols) +3. Add affected clusters section (query MEMBER_OF for all affected symbols) +4. Add risk assessment summary +5. Surface confidence scores more prominently (group by confidence level) + +**New output sections:** + +```javascript +AFFECTED PROCESSES: +- LoginFlow - BROKEN at step 2 +- SignupFlow - BROKEN at step 1 + +AFFECTED CLUSTERS: +- Authentication (direct) +- API Routes (indirect) + +RISK: CRITICAL +- N direct callers +- N processes affected +- N clusters affected +``` + + + +### 6. Increase Process Detection Limits + +**File:** [gitnexus/src/core/ingestion/process-processor.ts](gitnexus/src/core/ingestion/process-processor.ts)Change default config (lines 27-32): + +```typescript +const DEFAULT_CONFIG: ProcessDetectionConfig = { + maxTraceDepth: 10, // Keep + maxBranching: 4, // Was 3 + maxProcesses: 75, // Was 50 + minSteps: 2, // Keep +}; +``` + + + +### 7. Update System Prompt + +**File:** [gitnexus/src/core/llm/agent.ts](gitnexus/src/core/llm/agent.ts)Update BASE_SYSTEM_PROMPT to reflect new tools: + +```javascript +## TOOLS +- **search** - Hybrid search. Results grouped by process with cluster context. +- **grep** - Regex pattern search for exact strings. +- **read** - Read file content. +- **explore** - Deep dive on a symbol, cluster, or process. Shows membership, participation, connections. +- **overview** - Codebase map showing all clusters and processes. +- **impact** - Impact analysis. Shows affected processes, clusters, and risk level. +- **cypher** - Raw Cypher queries against the graph. + +## GRAPH SCHEMA +Nodes: File, Folder, Function, Class, Interface, Method, Community, Process +Relations: CodeRelation with type: CONTAINS, DEFINES, IMPORTS, CALLS, EXTENDS, IMPLEMENTS, MEMBER_OF, STEP_IN_PROCESS +``` + + + +## Implementation Order + +1. Remove highlight tool (cleanup) +2. Increase process detection limits +3. Add overview tool (simplest new tool) +4. Add explore tool +5. Enhance impact tool \ No newline at end of file diff --git a/.sisyphus/drafts/gitnexus-brainstorming.md b/.sisyphus/drafts/gitnexus-brainstorming.md new file mode 100644 index 000000000..95865a68b --- /dev/null +++ b/.sisyphus/drafts/gitnexus-brainstorming.md @@ -0,0 +1,18 @@ +# Draft: Gitnexus Brainstorming - Clustering & Process Maps + +## Initial Context +- Project: **GitnexusV2** +- Structure: + - `gitnexus/` (Likely the core application) + - `gitnexus-mcp/` (Likely a Model Context Protocol server) +- Goal: Make it accurate and usable for smaller/dumber models. +- Current Focus: Implementing **Clustering** and **Process Maps**. + +## Findings +- **Clustering**: Found `gitnexus/src/core/ingestion/cluster-enricher.ts`. +- **Process Maps**: No files matched `*process*map*` yet. Searching content next. + +## Open Questions +- How is "process map" defined in this context? (Graph, mermaid diagram, flowchart?) +- What is the input for clustering? (Code chunks, files, commits?) +- What is the intended output for "smaller models"? (Simplified context, summaries?) diff --git a/.sisyphus/drafts/noodlbox-comparison.md b/.sisyphus/drafts/noodlbox-comparison.md new file mode 100644 index 000000000..3ef3adc48 --- /dev/null +++ b/.sisyphus/drafts/noodlbox-comparison.md @@ -0,0 +1,34 @@ +# Draft: Gitnexus vs Noodlbox Strategy + +## Objectives +- Understand GitnexusV2 current state and goals. +- Analyze Noodlbox capabilities from provided URL. +- Compare features, architecture, and value proposition. +- Provide strategic views and recommendations. + +## Research Findings +- [GitnexusV2]: Zero-server, browser-native (WASM), KuzuDB based. Graph + Vector hybrid search. +- [Noodlbox]: CLI-first, heavy install. Has "Session Hooks" and "Search Hooks" via plugins/CLI. + +## Comparison Points +- **Core Philosophy**: Both bet on "Knowledge Graph + MCP" as the future. Noodlbox validates Gitnexus's direction. +- **Architecture**: + - *Noodlbox*: CLI/Binary based. Likely local server management. + - *Gitnexus*: Zero-server, Browser-native (WASM). Lower friction, higher privacy. +- **Features**: + - *Communities/Processes*: Both have them. Noodlbox uses them for "context injection". Gitnexus uses them for "visual exploration + query". + - *Impact Analysis*: Noodlbox has polished workflows (e.g., `detect_impact staged`). Gitnexus has the engine (`blastRadius`) but maybe not the specific workflow wrappers yet. +- **UX/Integration**: + - *Noodlbox*: "Hooks" (Session/Search) are a killer feature. Proactively injecting context into the agent's session. + - *Gitnexus*: Powerful tools, but relies on agent *pulling* data? + +## Strategic Views +1. **Validation**: The market direction is confirmed. You are building the right thing. +2. **differentiation**: Lean into "Zero-Setup / Browser-Native". Noodlbox requires `noodl init` and CLI handling. Gitnexus could just *be*. +3. **Opportunity**: Steal the "Session/Search Hooks" pattern. Make the agent smarter *automatically* without the user asking "check impact". +4. **Workflow Polish**: Noodlbox's `/detect_impact staged` is a great specific use case. Gitnexus should wrap `blastRadius` into similar concrete workflows. + +## Technical Feasibility (Interception) +- **Cursor**: Use `.cursorrules` to "shadow" default tools. Instruct agent to ALWAYS use `gitnexus_search` instead of `grep`. +- **Claude Code**: Likely uses a private plugin API for `PreToolUse`. We can't match this exactly without an official plugin, but we can approximate it with strong prompt instructions in `AGENTS.md`. +- **MCP Shadowing**: Define tools with names that conflict (e.g., `grep`)? No, unsafe. Better to use "Virtual Hooks" via system prompt instructions. diff --git a/gitnexus/src/components/ActivityFeed.tsx b/gitnexus/src/components/ActivityFeed.tsx index 0f4444e0d..a802b859b 100644 --- a/gitnexus/src/components/ActivityFeed.tsx +++ b/gitnexus/src/components/ActivityFeed.tsx @@ -6,7 +6,7 @@ */ import { useState, useEffect, useRef } from 'react'; -import { Activity, Search, Database, Terminal, Eye, Loader2, CheckCircle, XCircle, Clock, FileText, Zap } from 'lucide-react'; +import { Activity, Search, Database, Terminal, Loader2, CheckCircle, XCircle, Clock, FileText, Zap, Map, Compass } from 'lucide-react'; import { getMCPClient, type ActivityEvent } from '../core/mcp/mcp-client'; // Tool icons @@ -16,8 +16,9 @@ const TOOL_ICONS: Record = { cypher: Database, grep: Terminal, read: FileText, - blastRadius: Activity, - highlight: Eye, + impact: Activity, + overview: Map, + explore: Compass, }; // Tool colors @@ -27,8 +28,9 @@ const TOOL_COLORS: Record = { cypher: 'text-purple-400', grep: 'text-green-400', read: 'text-blue-400', - blastRadius: 'text-rose-400', - highlight: 'text-teal-400', + impact: 'text-rose-400', + overview: 'text-indigo-400', + explore: 'text-teal-400', }; export function ActivityFeed() { diff --git a/gitnexus/src/components/Header.tsx b/gitnexus/src/components/Header.tsx index 9e4936223..c1926842a 100644 --- a/gitnexus/src/components/Header.tsx +++ b/gitnexus/src/components/Header.tsx @@ -238,27 +238,74 @@ export const Header = ({ onFocusNode }: HeaderProps) => { const results = await runQuery(query); return results; }} - onBlastRadius={async (nodeId, hops = 2) => { - // Run blast radius query + onImpact={async (nodeId: string, hops = 2) => { + // Run impact analysis query const query = ` MATCH (start)-[*1..${hops}]-(connected) WHERE start.id = '${nodeId}' OR start.name = '${nodeId}' RETURN DISTINCT connected.id AS id, connected.name AS name, labels(connected) AS labels `; const results = await runQuery(query); - // Trigger ripple animation on blast radius results + // Trigger ripple animation on impact results const nodeIds = results.map((r: any) => r.id).filter(Boolean); if (nodeIds.length > 0) { triggerNodeAnimation(nodeIds, 'ripple'); } return results; }} - onHighlight={(nodeIds) => { - // Highlight nodes in the graph - setHighlightedNodeIds(new Set(nodeIds)); - // Trigger glow animation on highlighted nodes - if (nodeIds.length > 0) { - triggerNodeAnimation(nodeIds, 'glow'); + onOverview={async () => { + // Return codebase overview: clusters + processes + const clustersQuery = ` + MATCH (c:Community) + OPTIONAL MATCH (c)<-[:CodeRelation {type: 'MEMBER_OF'}]-(m) + RETURN c.id AS id, c.label AS label, c.cohesion AS cohesion, c.description AS description, count(m) AS memberCount + ORDER BY memberCount DESC + LIMIT 50 + `; + const processesQuery = ` + MATCH (p:Process) + RETURN p.id AS id, p.label AS label, p.processType AS type, p.stepCount AS steps + ORDER BY p.stepCount DESC + LIMIT 50 + `; + const [clusters, processes] = await Promise.all([ + runQuery(clustersQuery), + runQuery(processesQuery), + ]); + return { clusters, processes }; + }} + onExplore={async (target: string, type?: 'symbol' | 'cluster' | 'process') => { + // Explore a specific target + if (type === 'cluster' || target.startsWith('comm_')) { + const query = ` + MATCH (c:Community) + WHERE c.id = '${target}' OR c.label CONTAINS '${target}' + OPTIONAL MATCH (c)<-[:CodeRelation {type: 'MEMBER_OF'}]-(m) + RETURN c.id AS id, c.label AS label, c.description AS description, collect(m.name)[0..10] AS members + LIMIT 1 + `; + return await runQuery(query); + } else if (type === 'process' || target.startsWith('proc_')) { + const query = ` + MATCH (p:Process) + WHERE p.id = '${target}' OR p.label CONTAINS '${target}' + OPTIONAL MATCH (s)-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p) + RETURN p.id AS id, p.label AS label, p.stepCount AS steps, collect({name: s.name, step: r.step})[0..20] AS trace + LIMIT 1 + `; + return await runQuery(query); + } else { + // Symbol exploration + const query = ` + MATCH (n) + WHERE n.name = '${target}' OR n.id ENDS WITH ':${target}' + OPTIONAL MATCH (n)-[:CodeRelation {type: 'MEMBER_OF'}]->(c:Community) + OPTIONAL MATCH (n)-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process) + RETURN n.id AS id, n.name AS name, n.filePath AS filePath, label(n) AS nodeType, + c.label AS cluster, collect({process: p.label, step: r.step}) AS processes + LIMIT 1 + `; + return await runQuery(query); } }} getContext={async () => { diff --git a/gitnexus/src/components/MCPToggle.tsx b/gitnexus/src/components/MCPToggle.tsx index 46b765889..87c2c3e94 100644 --- a/gitnexus/src/components/MCPToggle.tsx +++ b/gitnexus/src/components/MCPToggle.tsx @@ -14,10 +14,11 @@ type ConnectionState = 'disconnected' | 'connecting' | 'connected' | 'error'; interface MCPToggleProps { onSearch?: (query: string, limit?: number) => Promise; onCypher?: (query: string) => Promise; - onBlastRadius?: (nodeId: string, hops?: number) => Promise; - onHighlight?: (nodeIds: string[], color?: string) => void; + onImpact?: (nodeId: string, hops?: number) => Promise; onGrep?: (pattern: string, caseSensitive?: boolean, maxResults?: number) => Promise; onRead?: (filePath: string, startLine?: number, endLine?: number) => Promise; + onOverview?: () => Promise; + onExplore?: (target: string, type?: 'symbol' | 'cluster' | 'process') => Promise; showOnboardingTip?: boolean; getContext?: () => Promise; } @@ -37,10 +38,11 @@ const MCP_CONFIG = `{ export function MCPToggle({ onSearch, onCypher, - onBlastRadius, - onHighlight, + onImpact, onGrep, onRead, + onOverview, + onExplore, showOnboardingTip = false, getContext, }: MCPToggleProps = {}) { @@ -92,10 +94,11 @@ export function MCPToggle({ // Register tool handlers if (onSearch) client.registerHandler('search', async (params) => onSearch(params.query, params.limit)); if (onCypher) client.registerHandler('cypher', async (params) => onCypher(params.query)); - if (onBlastRadius) client.registerHandler('blastRadius', async (params) => onBlastRadius(params.nodeId, params.hops)); - if (onHighlight) client.registerHandler('highlight', async (params) => { onHighlight(params.nodeIds, params.color); return { highlighted: params.nodeIds.length }; }); + if (onImpact) client.registerHandler('impact', async (params) => onImpact(params.nodeId, params.hops)); if (onGrep) client.registerHandler('grep', async (params) => onGrep(params.pattern, params.caseSensitive, params.maxResults)); if (onRead) client.registerHandler('read', async (params) => onRead(params.filePath, params.startLine, params.endLine)); + if (onOverview) client.registerHandler('overview', async () => onOverview()); + if (onExplore) client.registerHandler('explore', async (params) => onExplore(params.target, params.type)); if (getContext) client.registerHandler('context', async () => getContext()); setStatus('connected'); @@ -114,7 +117,7 @@ export function MCPToggle({ } catch { setStatus('error'); } - }, [onSearch, onCypher, onBlastRadius, onHighlight, onGrep, onRead, getContext]); + }, [onSearch, onCypher, onImpact, onGrep, onRead, onOverview, onExplore, getContext]); const disconnect = useCallback(() => { const client = getMCPClient(); diff --git a/gitnexus/src/core/graph/types.ts b/gitnexus/src/core/graph/types.ts index 1e8fa2037..17a12388e 100644 --- a/gitnexus/src/core/graph/types.ts +++ b/gitnexus/src/core/graph/types.ts @@ -14,7 +14,8 @@ export type NodeLabel = | 'Import' | 'Type' | 'CodeElement' - | 'Community'; + | 'Community' + | 'Process'; export type NodeProperties = { @@ -31,6 +32,12 @@ export type NodeProperties = { keywords?: string[], description?: string, enrichedBy?: 'heuristic' | 'llm', + // Process-specific properties + processType?: 'intra_community' | 'cross_community', + stepCount?: number, + communities?: string[], + entryPointId?: string, + terminalId?: string, } export type RelationshipType = @@ -45,6 +52,7 @@ export type RelationshipType = | 'IMPLEMENTS' | 'EXTENDS' | 'MEMBER_OF' + | 'STEP_IN_PROCESS' export interface GraphNode { id: string, @@ -61,6 +69,8 @@ export interface GraphRelationship { confidence: number, /** Resolution reason: 'import-resolved', 'same-file', 'fuzzy-global', or empty for non-CALLS */ reason: string, + /** Step number for STEP_IN_PROCESS relationships (1-indexed) */ + step?: number, } export interface KnowledgeGraph { diff --git a/gitnexus/src/core/ingestion/call-processor.ts b/gitnexus/src/core/ingestion/call-processor.ts index d0ac9b22e..2b71c5aaa 100644 --- a/gitnexus/src/core/ingestion/call-processor.ts +++ b/gitnexus/src/core/ingestion/call-processor.ts @@ -34,6 +34,7 @@ const FUNCTION_NODE_TYPES = new Set([ 'local_function_statement', // Rust 'function_item', + 'impl_item', // Methods inside impl blocks ]); /** @@ -51,21 +52,46 @@ const findEnclosingFunction = ( if (FUNCTION_NODE_TYPES.has(current.type)) { // Found enclosing function - try to get its name let funcName: string | null = null; + let label = 'Function'; // Different node types have different name locations if (current.type === 'function_declaration' || current.type === 'function_definition' || current.type === 'async_function_declaration' || - current.type === 'generator_function_declaration') { + current.type === 'generator_function_declaration' || + current.type === 'function_item') { // Rust function // Named function: function foo() {} const nameNode = current.childForFieldName?.('name') || current.children?.find((c: any) => c.type === 'identifier' || c.type === 'property_identifier'); funcName = nameNode?.text; + } else if (current.type === 'impl_item') { + // Rust method inside impl block: wrapper around function_item or const_item + // We need to look inside for the function_item + const funcItem = current.children?.find((c: any) => c.type === 'function_item'); + if (funcItem) { + const nameNode = funcItem.childForFieldName?.('name') || + funcItem.children?.find((c: any) => c.type === 'identifier'); + funcName = nameNode?.text; + label = 'Method'; + } } else if (current.type === 'method_definition') { - // Method: foo() {} inside class + // Method: foo() {} inside class (JS/TS) const nameNode = current.childForFieldName?.('name') || current.children?.find((c: any) => c.type === 'property_identifier'); funcName = nameNode?.text; + label = 'Method'; + } else if (current.type === 'method_declaration') { + // Java method: public void foo() {} + const nameNode = current.childForFieldName?.('name') || + current.children?.find((c: any) => c.type === 'identifier'); + funcName = nameNode?.text; + label = 'Method'; + } else if (current.type === 'constructor_declaration') { + // Java constructor: public ClassName() {} + const nameNode = current.childForFieldName?.('name') || + current.children?.find((c: any) => c.type === 'identifier'); + funcName = nameNode?.text; + label = 'Method'; // Treat constructors as methods for process detection } else if (current.type === 'arrow_function' || current.type === 'function_expression') { // Arrow/expression: const foo = () => {} - check parent variable declarator const parent = current.parent; @@ -78,12 +104,18 @@ const findEnclosingFunction = ( if (funcName) { // Look up the function in symbol table to get its node ID + // Try exact match first const nodeId = symbolTable.lookupExact(filePath, funcName); if (nodeId) return nodeId; - // Fallback: generate ID based on name and file - const fallbackLabel = current.type === 'method_definition' ? 'Method' : 'Function'; - return generateId(fallbackLabel, `${filePath}:${funcName}`); + // Try construct ID manually if lookup fails (common for non-exported internal functions) + // Format should match what parsing-processor generates: "Function:path/to/file:funcName" + // Check if we already have a node with this ID in the symbol table to be safe + const generatedId = generateId(label, `${filePath}:${funcName}`); + + // Ideally we should verify this ID exists, but strictly speaking if we are inside it, + // it SHOULD exist. Returning it is better than falling back to File. + return generatedId; } // Couldn't determine function name - try parent (might be nested) diff --git a/gitnexus/src/core/ingestion/pipeline.ts b/gitnexus/src/core/ingestion/pipeline.ts index a273d1a5d..acdfb4f0c 100644 --- a/gitnexus/src/core/ingestion/pipeline.ts +++ b/gitnexus/src/core/ingestion/pipeline.ts @@ -6,6 +6,7 @@ import { processImports, createImportMap } from './import-processor'; import { processCalls } from './call-processor'; import { processHeritage } from './heritage-processor'; import { processCommunities, CommunityDetectionResult } from './community-processor'; +import { processProcesses, ProcessDetectionResult } from './process-processor'; import { createSymbolTable } from './symbol-table'; import { createASTCache } from './ast-cache'; import { PipelineProgress, PipelineResult } from '../../types/pipeline'; @@ -210,12 +211,70 @@ export const runPipelineFromFiles = async ( }); }); + // Phase 8: Process Detection (98-99%) + onProgress({ + phase: 'processes', + percent: 98, + message: 'Detecting execution flows...', + stats: { filesProcessed: files.length, totalFiles: files.length, nodesCreated: graph.nodeCount }, + }); + + const processResult = await processProcesses( + graph, + communityResult.memberships, + (message, progress) => { + const processProgress = 98 + (progress * 0.01); + onProgress({ + phase: 'processes', + percent: Math.round(processProgress), + message, + stats: { filesProcessed: files.length, totalFiles: files.length, nodesCreated: graph.nodeCount }, + }); + } + ); + + // Log process detection results + if (import.meta.env.DEV) { + console.log(`πŸ”„ Process detection: ${processResult.stats.totalProcesses} processes found (${processResult.stats.crossCommunityCount} cross-community)`); + } + + // Add Process nodes to the graph + processResult.processes.forEach(proc => { + graph.addNode({ + id: proc.id, + label: 'Process' as const, + properties: { + name: proc.label, + filePath: '', + heuristicLabel: proc.heuristicLabel, + processType: proc.processType, + stepCount: proc.stepCount, + communities: proc.communities, + entryPointId: proc.entryPointId, + terminalId: proc.terminalId, + } + }); + }); + + // Add STEP_IN_PROCESS relationships + processResult.steps.forEach(step => { + graph.addRelationship({ + id: `${step.nodeId}_step_${step.step}_${step.processId}`, + type: 'STEP_IN_PROCESS', + sourceId: step.nodeId, + targetId: step.processId, + confidence: 1.0, + reason: 'trace-detection', + step: step.step, + }); + }); + - // Phase 8: Complete (100%) + // Phase 9: Complete (100%) onProgress({ phase: 'complete', percent: 100, - message: `Graph complete! ${communityResult.stats.totalCommunities} communities detected.`, + message: `Graph complete! ${communityResult.stats.totalCommunities} communities, ${processResult.stats.totalProcesses} processes detected.`, stats: { filesProcessed: files.length, totalFiles: files.length, @@ -226,7 +285,7 @@ export const runPipelineFromFiles = async ( // Cleanup WASM memory before returning astCache.clear(); - return { graph, fileContents, communityResult }; + return { graph, fileContents, communityResult, processResult }; } catch (error) { cleanup(); diff --git a/gitnexus/src/core/ingestion/process-processor.ts b/gitnexus/src/core/ingestion/process-processor.ts new file mode 100644 index 000000000..654eae992 --- /dev/null +++ b/gitnexus/src/core/ingestion/process-processor.ts @@ -0,0 +1,393 @@ +/** + * Process Detection Processor + * + * Detects execution flows (Processes) in the code graph by: + * 1. Finding entry points (functions with no internal callers) + * 2. Tracing forward via CALLS edges (BFS) + * 3. Grouping and deduplicating similar paths + * 4. Labeling with heuristic names + * + * Processes help agents understand how features work through the codebase. + */ + +import { KnowledgeGraph, GraphNode, GraphRelationship, NodeLabel } from '../graph/types'; +import { CommunityMembership } from './community-processor'; + +// ============================================================================ +// CONFIGURATION +// ============================================================================ + +export interface ProcessDetectionConfig { + maxTraceDepth: number; // Maximum steps to trace (default: 10) + maxBranching: number; // Max branches to follow per node (default: 3) + maxProcesses: number; // Maximum processes to detect (default: 50) + minSteps: number; // Minimum steps for a valid process (default: 2) +} + +const DEFAULT_CONFIG: ProcessDetectionConfig = { + maxTraceDepth: 10, + maxBranching: 4, + maxProcesses: 75, + minSteps: 2, +}; + +// ============================================================================ +// TYPES +// ============================================================================ + +export interface ProcessNode { + id: string; // "proc_handleLogin_createSession" + label: string; // "HandleLogin β†’ CreateSession" + heuristicLabel: string; + processType: 'intra_community' | 'cross_community'; + stepCount: number; + communities: string[]; // Community IDs touched + entryPointId: string; + terminalId: string; + trace: string[]; // Ordered array of node IDs +} + +export interface ProcessStep { + nodeId: string; + processId: string; + step: number; // 1-indexed position in trace +} + +export interface ProcessDetectionResult { + processes: ProcessNode[]; + steps: ProcessStep[]; + stats: { + totalProcesses: number; + crossCommunityCount: number; + avgStepCount: number; + entryPointsFound: number; + }; +} + +// ============================================================================ +// MAIN PROCESSOR +// ============================================================================ + +/** + * Detect processes (execution flows) in the knowledge graph + * + * This runs AFTER community detection, using CALLS edges to trace flows. + */ +export const processProcesses = async ( + knowledgeGraph: KnowledgeGraph, + memberships: CommunityMembership[], + onProgress?: (message: string, progress: number) => void, + config: Partial = {} +): Promise => { + const cfg = { ...DEFAULT_CONFIG, ...config }; + + onProgress?.('Finding entry points...', 0); + + // Build lookup maps + const membershipMap = new Map(); + memberships.forEach(m => membershipMap.set(m.nodeId, m.communityId)); + + const callsEdges = buildCallsGraph(knowledgeGraph); + const reverseCallsEdges = buildReverseCallsGraph(knowledgeGraph); + const nodeMap = new Map(); + knowledgeGraph.nodes.forEach(n => nodeMap.set(n.id, n)); + + // Step 1: Find entry points (functions that call others but have few callers) + const entryPoints = findEntryPoints(knowledgeGraph, reverseCallsEdges, callsEdges); + + onProgress?.(`Found ${entryPoints.length} entry points, tracing flows...`, 20); + + onProgress?.(`Found ${entryPoints.length} entry points, tracing flows...`, 20); + + // Step 2: Trace processes from each entry point + const allTraces: string[][] = []; + + for (let i = 0; i < entryPoints.length && allTraces.length < cfg.maxProcesses * 2; i++) { + const entryId = entryPoints[i]; + const traces = traceFromEntryPoint(entryId, callsEdges, cfg); + + // Filter out traces that are too short + traces.filter(t => t.length >= cfg.minSteps).forEach(t => allTraces.push(t)); + + if (i % 10 === 0) { + onProgress?.(`Tracing entry point ${i + 1}/${entryPoints.length}...`, 20 + (i / entryPoints.length) * 40); + } + } + + onProgress?.(`Found ${allTraces.length} traces, deduplicating...`, 60); + + // Step 3: Deduplicate similar traces + const uniqueTraces = deduplicateTraces(allTraces); + + // Step 4: Limit to max processes (prioritize longer traces) + const limitedTraces = uniqueTraces + .sort((a, b) => b.length - a.length) + .slice(0, cfg.maxProcesses); + + onProgress?.(`Creating ${limitedTraces.length} process nodes...`, 80); + + // Step 5: Create process nodes + const processes: ProcessNode[] = []; + const steps: ProcessStep[] = []; + + limitedTraces.forEach((trace, idx) => { + const entryPointId = trace[0]; + const terminalId = trace[trace.length - 1]; + + // Get communities touched + const communitiesSet = new Set(); + trace.forEach(nodeId => { + const comm = membershipMap.get(nodeId); + if (comm) communitiesSet.add(comm); + }); + const communities = Array.from(communitiesSet); + + // Determine process type + const processType: 'intra_community' | 'cross_community' = + communities.length > 1 ? 'cross_community' : 'intra_community'; + + // Generate label + const entryNode = nodeMap.get(entryPointId); + const terminalNode = nodeMap.get(terminalId); + const entryName = entryNode?.properties.name || 'Unknown'; + const terminalName = terminalNode?.properties.name || 'Unknown'; + const heuristicLabel = `${capitalize(entryName)} β†’ ${capitalize(terminalName)}`; + + const processId = `proc_${idx}_${sanitizeId(entryName)}`; + + processes.push({ + id: processId, + label: heuristicLabel, + heuristicLabel, + processType, + stepCount: trace.length, + communities, + entryPointId, + terminalId, + trace, + }); + + // Create step relationships + trace.forEach((nodeId, stepIdx) => { + steps.push({ + nodeId, + processId, + step: stepIdx + 1, // 1-indexed + }); + }); + }); + + onProgress?.('Process detection complete!', 100); + + // Calculate stats + const crossCommunityCount = processes.filter(p => p.processType === 'cross_community').length; + const avgStepCount = processes.length > 0 + ? processes.reduce((sum, p) => sum + p.stepCount, 0) / processes.length + : 0; + + return { + processes, + steps, + stats: { + totalProcesses: processes.length, + crossCommunityCount, + avgStepCount: Math.round(avgStepCount * 10) / 10, + entryPointsFound: entryPoints.length, + }, + }; +}; + +// ============================================================================ +// HELPER: Build CALLS adjacency list +// ============================================================================ + +type AdjacencyList = Map; + +const buildCallsGraph = (graph: KnowledgeGraph): AdjacencyList => { + const adj = new Map(); + + graph.relationships.forEach(rel => { + if (rel.type === 'CALLS') { + if (!adj.has(rel.sourceId)) { + adj.set(rel.sourceId, []); + } + adj.get(rel.sourceId)!.push(rel.targetId); + } + }); + + return adj; +}; + +const buildReverseCallsGraph = (graph: KnowledgeGraph): AdjacencyList => { + const adj = new Map(); + + graph.relationships.forEach(rel => { + if (rel.type === 'CALLS') { + if (!adj.has(rel.targetId)) { + adj.set(rel.targetId, []); + } + adj.get(rel.targetId)!.push(rel.sourceId); + } + }); + + return adj; +}; + +/** + * Find functions/methods that are good entry points for tracing. + * + * An entry point is a function that: + * 1. Has outgoing CALLS (so we can trace forward) + * 2. Ranked by having high outgoing/incoming call ratio + * + * We prioritize functions that call many others but are called by few. + */ +const findEntryPoints = ( + graph: KnowledgeGraph, + reverseCallsEdges: AdjacencyList, + callsEdges: AdjacencyList +): string[] => { + const symbolTypes = new Set(['Function', 'Method']); + const entryPointCandidates: { id: string; score: number; callers: number; callees: number }[] = []; + + graph.nodes.forEach(node => { + if (symbolTypes.has(node.label)) { + const callers = reverseCallsEdges.get(node.id) || []; + const callees = callsEdges.get(node.id) || []; + + // Must have at least 1 outgoing call to trace forward + if (callees.length > 0) { + // Score: ratio of outgoing to incoming calls + // Higher ratio = better entry point (calls many, called by few) + // Add 1 to denominators to avoid division by zero + const score = callees.length / (callers.length + 1); + + entryPointCandidates.push({ + id: node.id, + score, + callers: callers.length, + callees: callees.length + }); + } + } + }); + + // Sort by score descending and return top candidates + const sorted = entryPointCandidates.sort((a, b) => b.score - a.score); + + // DEBUG: Log top candidates + if (sorted.length > 0) { + console.log(`[Process Debug] Top 5 entry point candidates:`); + sorted.slice(0, 5).forEach((c, i) => { + const node = graph.nodes.find(n => n.id === c.id); + console.log(` ${i+1}. ${node?.properties.name} - calls: ${c.callees}, callers: ${c.callers}, score: ${c.score.toFixed(2)}`); + }); + } + + return sorted + .slice(0, 200) // Limit to prevent explosion + .map(c => c.id); +}; + +// ============================================================================ +// HELPER: Trace from entry point (BFS) +// ============================================================================ + +/** + * Trace forward from an entry point using BFS. + * Returns all distinct paths up to maxDepth. + */ +const traceFromEntryPoint = ( + entryId: string, + callsEdges: AdjacencyList, + config: ProcessDetectionConfig +): string[][] => { + const traces: string[][] = []; + + // BFS with path tracking + // Each queue item: [currentNodeId, pathSoFar] + const queue: [string, string[]][] = [[entryId, [entryId]]]; + const visited = new Set(); + + while (queue.length > 0 && traces.length < config.maxBranching * 3) { + const [currentId, path] = queue.shift()!; + + // Get outgoing calls + const callees = callsEdges.get(currentId) || []; + + if (callees.length === 0) { + // Terminal node - this is a complete trace + if (path.length >= config.minSteps) { + traces.push([...path]); + } + } else if (path.length >= config.maxTraceDepth) { + // Max depth reached - save what we have + if (path.length >= config.minSteps) { + traces.push([...path]); + } + } else { + // Continue tracing - limit branching + const limitedCallees = callees.slice(0, config.maxBranching); + let addedBranch = false; + + for (const calleeId of limitedCallees) { + // Avoid cycles + if (!path.includes(calleeId)) { + queue.push([calleeId, [...path, calleeId]]); + addedBranch = true; + } + } + + // If all branches were cycles, save current path as terminal + if (!addedBranch && path.length >= config.minSteps) { + traces.push([...path]); + } + } + } + + return traces; +}; + +// ============================================================================ +// HELPER: Deduplicate traces +// ============================================================================ + +/** + * Merge traces that are subsets of other traces. + * Keep longer traces, remove redundant shorter ones. + */ +const deduplicateTraces = (traces: string[][]): string[][] => { + if (traces.length === 0) return []; + + // Sort by length descending + const sorted = [...traces].sort((a, b) => b.length - a.length); + const unique: string[][] = []; + + for (const trace of sorted) { + // Check if this trace is a subset of any already-added trace + const traceKey = trace.join('->'); + const isSubset = unique.some(existing => { + const existingKey = existing.join('->'); + return existingKey.includes(traceKey); + }); + + if (!isSubset) { + unique.push(trace); + } + } + + return unique; +}; + +// ============================================================================ +// HELPER: String utilities +// ============================================================================ + +const capitalize = (s: string): string => { + if (!s) return s; + return s.charAt(0).toUpperCase() + s.slice(1); +}; + +const sanitizeId = (s: string): string => { + return s.replace(/[^a-zA-Z0-9]/g, '_').substring(0, 20).toLowerCase(); +}; diff --git a/gitnexus/src/core/kuzu/csv-generator.ts b/gitnexus/src/core/kuzu/csv-generator.ts index 0b1689893..43df569cf 100644 --- a/gitnexus/src/core/kuzu/csv-generator.ts +++ b/gitnexus/src/core/kuzu/csv-generator.ts @@ -226,9 +226,35 @@ const generateCommunityCSV = (nodes: GraphNode[]): string => { return rows.join('\n'); }; -// ============================================================================ -// RELATIONSHIP CSV GENERATOR (Single Table) -// ============================================================================ +/** + * Generate CSV for Process nodes + * Headers: id,label,heuristicLabel,processType,stepCount,communities,entryPointId,terminalId + */ +const generateProcessCSV = (nodes: GraphNode[]): string => { + const headers = ['id', 'label', 'heuristicLabel', 'processType', 'stepCount', 'communities', 'entryPointId', 'terminalId']; + const rows: string[] = [headers.join(',')]; + + for (const node of nodes) { + if (node.label !== 'Process') continue; + + // Handle communities array (string[]) + const communities = (node.properties as any).communities || []; + const communitiesStr = `[${communities.map((c: string) => `'${c.replace(/'/g, "''")}'`).join(',')}]`; + + rows.push([ + escapeCSVField(node.id), + escapeCSVField(node.properties.name || ''), // label stores name + escapeCSVField((node.properties as any).heuristicLabel || ''), + escapeCSVField((node.properties as any).processType || ''), + escapeCSVNumber((node.properties as any).stepCount, 0), + escapeCSVField(communitiesStr), // Needs CSV escaping because it contains commas! + escapeCSVField((node.properties as any).entryPointId || ''), + escapeCSVField((node.properties as any).terminalId || ''), + ].join(',')); + } + + return rows.join('\n'); +}; /** * Generate CSV for the single CodeRelation table @@ -238,7 +264,7 @@ const generateCommunityCSV = (nodes: GraphNode[]): string => { * reason: 'import-resolved' | 'same-file' | 'fuzzy-global' (or empty for non-CALLS) */ const generateRelationCSV = (graph: KnowledgeGraph): string => { - const headers = ['from', 'to', 'type', 'confidence', 'reason']; + const headers = ['from', 'to', 'type', 'confidence', 'reason', 'step']; const rows: string[] = [headers.join(',')]; for (const rel of graph.relationships) { @@ -248,6 +274,7 @@ const generateRelationCSV = (graph: KnowledgeGraph): string => { escapeCSVField(rel.type), escapeCSVNumber(rel.confidence, 1.0), escapeCSVField(rel.reason), + escapeCSVNumber((rel as any).step, 0), ].join(',')); } @@ -278,6 +305,7 @@ export const generateAllCSVs = ( nodeCSVs.set('Method', generateCodeElementCSV(nodes, 'Method', fileContents)); nodeCSVs.set('CodeElement', generateCodeElementCSV(nodes, 'CodeElement', fileContents)); nodeCSVs.set('Community', generateCommunityCSV(nodes)); + nodeCSVs.set('Process', generateProcessCSV(nodes)); // Generate single relation CSV const relCSV = generateRelationCSV(graph); diff --git a/gitnexus/src/core/kuzu/kuzu-adapter.ts b/gitnexus/src/core/kuzu/kuzu-adapter.ts index 73aa40b25..c16e2edf3 100644 --- a/gitnexus/src/core/kuzu/kuzu-adapter.ts +++ b/gitnexus/src/core/kuzu/kuzu-adapter.ts @@ -120,12 +120,15 @@ export const loadGraphToKuzu = async ( for (const line of relLines) { try { // Parse CSV - handle quoted fields and numeric confidence - // Format: "from","to","type",confidence,"reason" - const match = line.match(/"([^"]*)","([^"]*)","([^"]*)",([0-9.]+),"([^"]*)"/); + // Parse CSV - handle quoted fields and numeric confidence + // Format: "from","to","type",confidence,"reason",step + // Note: step is unquoted numeric + const match = line.match(/"([^"]*)","([^"]*)","([^"]*)",([0-9.]+),"([^"]*)",([0-9-]+)/); if (!match) continue; - const [, fromId, toId, relType, confidenceStr, reason] = match; + const [, fromId, toId, relType, confidenceStr, reason, stepStr] = match; const confidence = parseFloat(confidenceStr) || 1.0; + const step = parseInt(stepStr) || 0; // Extract labels from node IDs // Community nodes have IDs like "comm_14" (no colon) @@ -134,6 +137,9 @@ export const loadGraphToKuzu = async ( if (nodeId.startsWith('comm_')) { return 'Community'; } + if (nodeId.startsWith('proc_')) { + return 'Process'; + } return nodeId.split(':')[0]; }; @@ -150,7 +156,7 @@ export const loadGraphToKuzu = async ( const insertQuery = ` MATCH (a:${fromLabel} {id: '${fromId.replace(/'/g, "''")}'}), (b:${toLabel} {id: '${toId.replace(/'/g, "''")}'}) - CREATE (a)-[:${REL_TABLE_NAME} {type: '${relType}', confidence: ${confidence}, reason: '${reason.replace(/'/g, "''")}'}]->(b) + CREATE (a)-[:${REL_TABLE_NAME} {type: '${relType}', confidence: ${confidence}, reason: '${reason.replace(/'/g, "''")}', step: ${step}}]->(b) `; await conn.query(insertQuery); insertedRels++; @@ -162,6 +168,7 @@ export const loadGraphToKuzu = async ( const [, fromId, toId, relType] = match; const getNodeLabel = (nodeId: string): string => { if (nodeId.startsWith('comm_')) return 'Community'; + if (nodeId.startsWith('proc_')) return 'Process'; return nodeId.split(':')[0]; }; const fromLabel = getNodeLabel(fromId); @@ -229,6 +236,9 @@ const getCopyQuery = (table: NodeTableName, path: string): string => { if (table === 'Community') { return `COPY Community(id, label, heuristicLabel, keywords, description, enrichedBy, cohesion, symbolCount) FROM "${path}" (HEADER=true, PARALLEL=false)`; } + if (table === 'Process') { + return `COPY Process(id, label, heuristicLabel, processType, stepCount, communities, entryPointId, terminalId) FROM "${path}" (HEADER=true, PARALLEL=false)`; + } // All code element tables: Function, Class, Interface, Method, CodeElement return `COPY ${table}(id, name, filePath, startLine, endLine, content) FROM "${path}" (HEADER=true, PARALLEL=false)`; }; diff --git a/gitnexus/src/core/kuzu/schema.ts b/gitnexus/src/core/kuzu/schema.ts index 437383c2d..a482d8e97 100644 --- a/gitnexus/src/core/kuzu/schema.ts +++ b/gitnexus/src/core/kuzu/schema.ts @@ -13,7 +13,7 @@ // NODE TABLE NAMES // ============================================================================ export const NODE_TABLES = [ - 'File', 'Folder', 'Function', 'Class', 'Interface', 'Method', 'CodeElement', 'Community', + 'File', 'Folder', 'Function', 'Class', 'Interface', 'Method', 'CodeElement', 'Community', 'Process', // Multi-language support 'Struct', 'Enum', 'Macro', 'Typedef', 'Union', 'Namespace', 'Trait', 'Impl', 'TypeAlias', 'Const', 'Static', 'Property', 'Record', 'Delegate', 'Annotation', 'Constructor', 'Template', 'Module' @@ -26,7 +26,7 @@ export type NodeTableName = typeof NODE_TABLES[number]; export const REL_TABLE_NAME = 'CodeRelation'; // Valid relation types -export const REL_TYPES = ['CONTAINS', 'DEFINES', 'IMPORTS', 'CALLS', 'EXTENDS', 'IMPLEMENTS', 'MEMBER_OF'] as const; +export const REL_TYPES = ['CONTAINS', 'DEFINES', 'IMPORTS', 'CALLS', 'EXTENDS', 'IMPLEMENTS', 'MEMBER_OF', 'STEP_IN_PROCESS'] as const; export type RelType = typeof REL_TYPES[number]; // ============================================================================ @@ -127,6 +127,23 @@ CREATE NODE TABLE Community ( PRIMARY KEY (id) )`; +// ============================================================================ +// PROCESS NODE TABLE (for execution flow detection) +// ============================================================================ + +export const PROCESS_SCHEMA = ` +CREATE NODE TABLE Process ( + id STRING, + label STRING, + heuristicLabel STRING, + processType STRING, + stepCount INT32, + communities STRING[], + entryPointId STRING, + terminalId STRING, + PRIMARY KEY (id) +)`; + // ============================================================================ // MULTI-LANGUAGE NODE TABLE SCHEMAS // ============================================================================ @@ -206,6 +223,8 @@ CREATE REL TABLE ${REL_TABLE_NAME} ( FROM Function TO \`Enum\`, FROM Function TO Namespace, FROM Function TO TypeAlias, + FROM Function TO \`Module\`, + FROM Function TO Impl, FROM Class TO Method, FROM Class TO Function, FROM Class TO Class, @@ -223,6 +242,7 @@ CREATE REL TABLE ${REL_TABLE_NAME} ( FROM Template TO Method, FROM Template TO Class, FROM Template TO \`Struct\`, + FROM \`Module\` TO \`Module\`, FROM CodeElement TO Community, FROM Interface TO Community, FROM \`Struct\` TO Community, @@ -231,6 +251,10 @@ CREATE REL TABLE ${REL_TABLE_NAME} ( FROM \`Struct\` TO Method, FROM \`Enum\` TO Community, FROM \`Macro\` TO Community, + FROM \`Macro\` TO Function, + FROM \`Macro\` TO Method, + FROM \`Module\` TO Function, + FROM \`Module\` TO Method, FROM Typedef TO Community, FROM \`Union\` TO Community, FROM Namespace TO Community, @@ -247,11 +271,28 @@ CREATE REL TABLE ${REL_TABLE_NAME} ( FROM Constructor TO Community, FROM Constructor TO Interface, FROM Constructor TO Class, + FROM Constructor TO Method, + FROM Constructor TO Function, + FROM Constructor TO Constructor, + FROM Constructor TO \`Struct\`, + FROM Constructor TO \`Macro\`, + FROM Constructor TO Template, FROM Template TO Community, FROM \`Module\` TO Community, + FROM Function TO Process, + FROM Method TO Process, + FROM Class TO Process, + FROM Interface TO Process, + FROM \`Struct\` TO Process, + FROM Constructor TO Process, + FROM \`Module\` TO Process, + FROM \`Macro\` TO Process, + FROM Impl TO Process, + FROM Typedef TO Process, type STRING, confidence DOUBLE, - reason STRING + reason STRING, + step INT32 )`; // ============================================================================ @@ -288,6 +329,7 @@ export const NODE_SCHEMA_QUERIES = [ METHOD_SCHEMA, CODE_ELEMENT_SCHEMA, COMMUNITY_SCHEMA, + PROCESS_SCHEMA, // Multi-language support STRUCT_SCHEMA, ENUM_SCHEMA, diff --git a/gitnexus/src/core/llm/agent.ts b/gitnexus/src/core/llm/agent.ts index 3584655c4..3e8b54cdd 100644 --- a/gitnexus/src/core/llm/agent.ts +++ b/gitnexus/src/core/llm/agent.ts @@ -67,26 +67,26 @@ You are an investigator. For each question: 3. **Trace** β†’ Use cypher to follow connections in the graph 4. **Cite** β†’ Ground every finding with [[file:line]] or [[Type:Name]] 5. **Validate** β†’ Use cypher to validate the results and confirm completeness of context before final output. ( MUST DO ) -6. **Highlight** β†’ Visualize key nodes with highlight ## πŸ› οΈ TOOLS -- **\`search\`** β€” Hybrid search (keyword + semantic). Returns code matches with graph connections. +- **\`search\`** β€” Hybrid search. Results grouped by process with cluster context. - **\`cypher\`** β€” Cypher queries against the graph. Use \`{{QUERY_VECTOR}}\` for vector search. - **\`grep\`** β€” Regex search. Best for exact strings, TODOs, error codes. - **\`read\`** β€” Read file content. Always use after search/grep to see full code. -- **\`highlight\`** β€” Highlight nodes in the visual graph. -- **\`blastRadius\`** β€” Impact analysis. Output is graph-verified (trusted). Run optional grep for dynamic patterns if thoroughness needed. +- **\`explore\`** β€” Deep dive on a symbol, cluster, or process. Shows membership, participation, connections. +- **\`overview\`** β€” Codebase map showing all clusters and processes. +- **\`impact\`** β€” Impact analysis. Shows affected processes, clusters, and risk level. ## πŸ“Š GRAPH SCHEMA -Nodes: File, Folder, Function, Class, Interface, Method, CodeElement -Relation: \`CodeRelation\` with \`type\` property: CONTAINS, DEFINES, IMPORTS, CALLS, EXTENDS, IMPLEMENTS +Nodes: File, Folder, Function, Class, Interface, Method, Community, Process +Relations: \`CodeRelation\` with \`type\` property: CONTAINS, DEFINES, IMPORTS, CALLS, EXTENDS, IMPLEMENTS, MEMBER_OF, STEP_IN_PROCESS Cypher examples: - \`MATCH (f:Function) RETURN f.name LIMIT 10\` - \`MATCH (f:File)-[:CodeRelation {type: 'IMPORTS'}]->(g:File) RETURN f.name, g.name\` ## πŸ“CRITICAL RULES -- **blastRadius output is trusted.** Do NOT re-validate with cypher. Optionally run the suggested grep commands for dynamic patterns. +- **impact output is trusted.** Do NOT re-validate with cypher. Optionally run the suggested grep commands for dynamic patterns. - **Cite or retract.** Never state something you can't ground. - **Read before concluding.** Don't guess from names alone. - **Retry on failure.** If a tool fails, fix the input and try again. diff --git a/gitnexus/src/core/llm/tools.ts b/gitnexus/src/core/llm/tools.ts index b86357ebe..ca75c7f63 100644 --- a/gitnexus/src/core/llm/tools.ts +++ b/gitnexus/src/core/llm/tools.ts @@ -1,13 +1,14 @@ /** * Graph RAG Tools for LangChain Agent * - * Consolidated tools (6 total): - * - search: Hybrid search (BM25 + semantic + RRF) with 1-hop expansion + * Consolidated tools (7 total): + * - search: Hybrid search (BM25 + semantic + RRF), grouped by process/cluster * - cypher: Execute Cypher queries (auto-embeds {{QUERY_VECTOR}} if present) * - grep: Regex pattern search across files * - read: Read file content by path - * - highlight: Highlight nodes in graph UI - * - blastRadius: Impact analysis (what depends on / is affected by changes) + * - overview: Codebase map (clusters + processes) + * - explore: Deep dive on a symbol, cluster, or process + * - impact: Impact analysis (what depends on / is affected by changes) */ import { tool } from '@langchain/core/tools'; @@ -36,8 +37,9 @@ export const createGraphRAGTools = ( * Unified search tool: BM25 + Semantic + RRF, with 1-hop graph context */ const searchTool = tool( - async ({ query, limit }: { query: string; limit?: number }) => { + async ({ query, limit, groupByProcess }: { query: string; limit?: number; groupByProcess?: boolean }) => { const k = limit ?? 10; + const shouldGroup = groupByProcess ?? true; // Step 1: Hybrid search (BM25 + semantic with RRF) let searchResults: any[] = []; @@ -62,12 +64,26 @@ export const createGraphRAGTools = ( return `No code found matching "${query}". Try different terms or use grep for exact patterns.`; } - // Step 2: Get 1-hop connections for each result - const resultsWithContext: string[] = []; + type ProcessInfo = { id: string; label: string; step?: number; stepCount?: number }; + type ResultInfo = { + idx: number; + nodeId: string; + name: string; + label: string; + filePath: string; + location: string; + sources: string; + score: string; + connections: string; + clusterLabel: string; + processes: ProcessInfo[]; + }; + + const results: ResultInfo[] = []; for (let i = 0; i < Math.min(searchResults.length, k); i++) { const r = searchResults[i]; - const nodeId = r.nodeId || r.id; + const nodeId = r.nodeId || r.id || ''; const name = r.name || r.filePath?.split('/').pop() || 'Unknown'; const label = r.label || 'File'; const filePath = r.filePath || ''; @@ -91,7 +107,6 @@ export const createGraphRAGTools = ( `; const connRes = await executeQuery(connectionsQuery); if (connRes.length > 0) { - // Result is nested array: [[outgoing], [incoming]] or {outgoing: [], incoming: []} const row = connRes[0]; const rawOutgoing = Array.isArray(row) ? row[0] : (row.outgoing || []); const rawIncoming = Array.isArray(row) ? row[1] : (row.incoming || []); @@ -116,18 +131,130 @@ export const createGraphRAGTools = ( } } - resultsWithContext.push( - `[${i + 1}] ${label}: ${name}${score}\n ID: ${nodeId}\n File: ${filePath}${location}\n Found by: ${sources}${connections}` - ); + // Cluster membership + let clusterLabel = 'Unclustered'; + if (nodeId) { + try { + const nodeLabel = nodeId.split(':')[0]; + const clusterQuery = ` + MATCH (n:${nodeLabel} {id: '${nodeId.replace(/'/g, "''")}'}) + MATCH (n)-[:CodeRelation {type: 'MEMBER_OF'}]->(c:Community) + RETURN c.label AS label + LIMIT 1 + `; + const clusterRes = await executeQuery(clusterQuery); + if (clusterRes.length > 0) { + const row = clusterRes[0]; + const labelValue = Array.isArray(row) ? row[0] : row.label; + if (labelValue) clusterLabel = labelValue; + } + } catch { + // Skip cluster lookup if query fails + } + } + + // Process participation + const processes: ProcessInfo[] = []; + if (nodeId) { + try { + const nodeLabel = nodeId.split(':')[0]; + const processQuery = ` + MATCH (n:${nodeLabel} {id: '${nodeId.replace(/'/g, "''")}'}) + MATCH (n)-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process) + RETURN p.id AS id, p.label AS label, r.step AS step, p.stepCount AS stepCount + ORDER BY r.step + `; + const procRes = await executeQuery(processQuery); + for (const row of procRes) { + const id = Array.isArray(row) ? row[0] : row.id; + const labelValue = Array.isArray(row) ? row[1] : row.label; + const step = Array.isArray(row) ? row[2] : row.step; + const stepCount = Array.isArray(row) ? row[3] : row.stepCount; + if (id && labelValue) { + processes.push({ id, label: labelValue, step, stepCount }); + } + } + } catch { + // Skip process lookup if query fails + } + } + + results.push({ + idx: i + 1, + nodeId, + name, + label, + filePath, + location, + sources, + score, + connections, + clusterLabel, + processes, + }); } - return `Found ${searchResults.length} matches:\n\n${resultsWithContext.join('\n\n')}`; + const formatResult = (r: ResultInfo, stepInfo?: ProcessInfo) => { + const stepLabel = stepInfo?.step ? ` (step ${stepInfo.step}/${stepInfo.stepCount ?? '?'})` : ''; + return `[${r.idx}] ${r.label}: ${r.name}${r.score}${stepLabel}\n ID: ${r.nodeId}\n File: ${r.filePath}${r.location}\n Cluster: ${r.clusterLabel}\n Found by: ${r.sources}${r.connections}`; + }; + + if (!shouldGroup) { + return `Found ${searchResults.length} matches:\n\n${results.map(r => formatResult(r)).join('\n\n')}`; + } + + // Group by process (or "No process") + const processMap = new Map(); + const noProcessKey = '__no_process__'; + + for (const r of results) { + if (r.processes.length === 0) { + if (!processMap.has(noProcessKey)) { + processMap.set(noProcessKey, { label: 'No process', entries: [] }); + } + processMap.get(noProcessKey)!.entries.push({ result: r }); + continue; + } + + for (const p of r.processes) { + if (!processMap.has(p.id)) { + processMap.set(p.id, { label: p.label, stepCount: p.stepCount, entries: [] }); + } + processMap.get(p.id)!.entries.push({ result: r, step: p.step, stepCount: p.stepCount }); + } + } + + const sortedProcesses = Array.from(processMap.entries()).sort((a, b) => { + const aCount = a[1].entries.length; + const bCount = b[1].entries.length; + return bCount - aCount; + }); + + const lines: string[] = []; + lines.push(`Found ${searchResults.length} matches grouped by process:`); + lines.push(''); + + for (const [pid, group] of sortedProcesses) { + const stepInfo = group.stepCount ? `, ${group.stepCount} steps` : ''; + const header = pid === noProcessKey + ? `NO PROCESS (${group.entries.length} matches)` + : `PROCESS: ${group.label} (${group.entries.length} matches${stepInfo})`; + lines.push(header); + group.entries.forEach(entry => { + const stepLabel = entry.step ? { id: pid, label: group.label, step: entry.step, stepCount: entry.stepCount } : undefined; + lines.push(formatResult(entry.result, stepLabel)); + }); + lines.push(''); + } + + return lines.join('\n').trim(); }, { name: 'search', - description: 'Search for code by keywords or concepts. Combines keyword matching and semantic understanding. Returns relevant code with their graph connections (what calls them, what they import, etc.).', + description: 'Search for code by keywords or concepts. Combines keyword matching and semantic understanding. Groups results by process with cluster context.', schema: z.object({ query: z.string().describe('What you are looking for (e.g., "authentication middleware", "database connection")'), + groupByProcess: z.boolean().optional().nullable().describe('Group results by process (default: true)'), limit: z.number().optional().nullable().describe('Max results to return (default: 10)'), }), } @@ -389,35 +516,364 @@ MATCH (n:Function {id: emb.nodeId}) RETURN n`, ); // ============================================================================ - // TOOL 5: HIGHLIGHT (Highlight nodes in graph UI) + // TOOL 5: OVERVIEW (Codebase map) // ============================================================================ - const highlightTool = tool( - async ({ nodeIds, description }: { nodeIds: string[]; description?: string }) => { - if (!nodeIds || nodeIds.length === 0) { - return 'No node IDs provided.'; + const overviewTool = tool( + async () => { + try { + const clustersQuery = ` + MATCH (c:Community) + RETURN c.id AS id, c.label AS label, c.cohesion AS cohesion, c.symbolCount AS symbolCount, c.description AS description + ORDER BY c.symbolCount DESC + LIMIT 200 + `; + const processesQuery = ` + MATCH (p:Process) + RETURN p.id AS id, p.label AS label, p.processType AS type, p.stepCount AS stepCount, p.communities AS communities + ORDER BY p.stepCount DESC + LIMIT 200 + `; + const depsQuery = ` + MATCH (a)-[:CodeRelation {type: 'CALLS'}]->(b) + MATCH (a)-[:CodeRelation {type: 'MEMBER_OF'}]->(c1:Community) + MATCH (b)-[:CodeRelation {type: 'MEMBER_OF'}]->(c2:Community) + WHERE c1.id <> c2.id + RETURN c1.label AS \`from\`, c2.label AS \`to\`, COUNT(*) AS calls + ORDER BY calls DESC + LIMIT 15 + `; + const criticalQuery = ` + MATCH (s)-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process) + RETURN p.label AS label, COUNT(r) AS steps + ORDER BY steps DESC + LIMIT 10 + `; + + const [clusters, processes, deps, critical] = await Promise.all([ + executeQuery(clustersQuery), + executeQuery(processesQuery), + executeQuery(depsQuery), + executeQuery(criticalQuery), + ]); + + const clusterLines = clusters.map((row: any) => { + const label = Array.isArray(row) ? row[1] : row.label; + const symbols = Array.isArray(row) ? row[3] : row.symbolCount; + const cohesion = Array.isArray(row) ? row[2] : row.cohesion; + const desc = Array.isArray(row) ? row[4] : row.description; + const cohesionText = cohesion !== null && cohesion !== undefined ? Number(cohesion).toFixed(2) : ''; + return `| ${label || ''} | ${symbols ?? ''} | ${cohesionText} | ${desc ?? ''} |`; + }); + + const processLines = processes.map((row: any) => { + const label = Array.isArray(row) ? row[1] : row.label; + const steps = Array.isArray(row) ? row[3] : row.stepCount; + const type = Array.isArray(row) ? row[2] : row.type; + const communities = Array.isArray(row) ? row[4] : row.communities; + const clusterText = Array.isArray(communities) ? communities.length : (communities ? 1 : 0); + return `| ${label || ''} | ${steps ?? ''} | ${type ?? ''} | ${clusterText} |`; + }); + + const depLines = deps.map((row: any) => { + const from = Array.isArray(row) ? row[0] : row.from; + const to = Array.isArray(row) ? row[1] : row.to; + const calls = Array.isArray(row) ? row[2] : row.calls; + return `- ${from} -> ${to} (${calls} calls)`; + }); + + const criticalLines = critical.map((row: any) => { + const label = Array.isArray(row) ? row[0] : row.label; + const steps = Array.isArray(row) ? row[1] : row.steps; + return `- ${label} (${steps} steps)`; + }); + + return [ + `CLUSTERS (${clusters.length} total):`, + `| Cluster | Symbols | Cohesion | Description |`, + `| --- | --- | --- | --- |`, + ...clusterLines, + ``, + `PROCESSES (${processes.length} total):`, + `| Process | Steps | Type | Clusters |`, + `| --- | --- | --- | --- |`, + ...processLines, + ``, + `CLUSTER DEPENDENCIES:`, + ...(depLines.length > 0 ? depLines : ['- None found']), + ``, + `CRITICAL PATHS:`, + ...(criticalLines.length > 0 ? criticalLines : ['- None found']), + ].join('\n'); + } catch (error) { + return `Overview error: ${error instanceof Error ? error.message : String(error)}`; } - - const marker = `[HIGHLIGHT_NODES:${nodeIds.join(',')}]`; - const desc = description || `Highlighting ${nodeIds.length} node(s)`; - - return `${desc}\n\n${marker}\n\nNodes highlighted in the graph.`; }, { - name: 'highlight', - description: 'Highlight nodes in the visual graph. Use node IDs from search/cypher results (format: Label:filepath:name).', + name: 'overview', + description: 'Codebase map showing all clusters and processes, plus cross-cluster dependencies.', + schema: z.object({}), + } + ); + + // ============================================================================ + // TOOL 6: EXPLORE (Deep dive on symbol, cluster, or process) + // ============================================================================ + + const exploreTool = tool( + async ({ target, type }: { target: string; type?: 'symbol' | 'cluster' | 'process' | null }) => { + const safeTarget = target.replace(/'/g, "''"); + let resolvedType = type ?? null; + let processRow: any | null = null; + let communityRow: any | null = null; + let symbolRow: any | null = null; + + const getRowValue = (row: any, idx: number, key: string) => Array.isArray(row) ? row[idx] : row[key]; + + if (!resolvedType || resolvedType === 'process') { + const processQuery = ` + MATCH (p:Process) + WHERE p.id = '${safeTarget}' OR p.label = '${safeTarget}' + RETURN p.id AS id, p.label AS label, p.processType AS type, p.stepCount AS stepCount + LIMIT 1 + `; + const processRes = await executeQuery(processQuery); + if (processRes.length > 0) { + processRow = processRes[0]; + resolvedType = 'process'; + } + } + + if (!resolvedType || resolvedType === 'cluster') { + const communityQuery = ` + MATCH (c:Community) + WHERE c.id = '${safeTarget}' OR c.label = '${safeTarget}' OR c.heuristicLabel = '${safeTarget}' + RETURN c.id AS id, c.label AS label, c.cohesion AS cohesion, c.symbolCount AS symbolCount, c.description AS description + LIMIT 1 + `; + const communityRes = await executeQuery(communityQuery); + if (communityRes.length > 0) { + communityRow = communityRes[0]; + resolvedType = 'cluster'; + } + } + + if (!resolvedType || resolvedType === 'symbol') { + const symbolQuery = ` + MATCH (n) + WHERE n.name = '${safeTarget}' OR n.id = '${safeTarget}' OR n.filePath = '${safeTarget}' + RETURN n.id AS id, n.name AS name, n.filePath AS filePath, label(n) AS nodeType + LIMIT 5 + `; + const symbolRes = await executeQuery(symbolQuery); + if (symbolRes.length > 0) { + symbolRow = symbolRes[0]; + resolvedType = 'symbol'; + } + } + + if (!resolvedType) { + return `Could not find "${target}" as a symbol, cluster, or process. Try search first.`; + } + + if (resolvedType === 'process') { + const pid = getRowValue(processRow, 0, 'id'); + const label = getRowValue(processRow, 1, 'label'); + const ptype = getRowValue(processRow, 2, 'type'); + const stepCount = getRowValue(processRow, 3, 'stepCount'); + + const stepsQuery = ` + MATCH (s)-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process {id: '${pid.replace(/'/g, "''")}'}) + RETURN s.name AS name, s.filePath AS filePath, r.step AS step + ORDER BY r.step + `; + const clustersQuery = ` + MATCH (c:Community)<-[:CodeRelation {type: 'MEMBER_OF'}]-(s) + MATCH (s)-[:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process {id: '${pid.replace(/'/g, "''")}'}) + RETURN DISTINCT c.id AS id, c.label AS label, c.description AS description + ORDER BY c.label + LIMIT 20 + `; + + const [steps, clusters] = await Promise.all([ + executeQuery(stepsQuery), + executeQuery(clustersQuery), + ]); + + const stepLines = steps.map((row: any) => { + const name = getRowValue(row, 0, 'name'); + const filePath = getRowValue(row, 1, 'filePath'); + const step = getRowValue(row, 2, 'step'); + return `- ${step}. ${name} (${filePath || 'n/a'})`; + }); + + const clusterLines = clusters.map((row: any) => { + const clabel = getRowValue(row, 1, 'label'); + const desc = getRowValue(row, 2, 'description'); + return `- ${clabel}${desc ? ` β€” ${desc}` : ''}`; + }); + + return [ + `PROCESS: ${label}`, + `Type: ${ptype || 'n/a'}`, + `Steps: ${stepCount ?? steps.length}`, + ``, + `STEPS:`, + ...(stepLines.length > 0 ? stepLines : ['- None found']), + ``, + `CLUSTERS TOUCHED:`, + ...(clusterLines.length > 0 ? clusterLines : ['- None found']), + ].join('\n'); + } + + if (resolvedType === 'cluster') { + const cid = getRowValue(communityRow, 0, 'id'); + const label = getRowValue(communityRow, 1, 'label'); + const cohesion = getRowValue(communityRow, 2, 'cohesion'); + const symbolCount = getRowValue(communityRow, 3, 'symbolCount'); + const description = getRowValue(communityRow, 4, 'description'); + + const membersQuery = ` + MATCH (c:Community {id: '${cid.replace(/'/g, "''")}'})<-[:CodeRelation {type: 'MEMBER_OF'}]-(m) + RETURN m.name AS name, m.filePath AS filePath, label(m) AS nodeType + LIMIT 50 + `; + const processesQuery = ` + MATCH (c:Community {id: '${cid.replace(/'/g, "''")}'})<-[:CodeRelation {type: 'MEMBER_OF'}]-(s) + MATCH (s)-[:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process) + RETURN DISTINCT p.id AS id, p.label AS label, p.stepCount AS stepCount + ORDER BY p.stepCount DESC + LIMIT 20 + `; + + const [members, processes] = await Promise.all([ + executeQuery(membersQuery), + executeQuery(processesQuery), + ]); + + const memberLines = members.map((row: any) => { + const name = getRowValue(row, 0, 'name'); + const filePath = getRowValue(row, 1, 'filePath'); + const nodeType = getRowValue(row, 2, 'nodeType'); + return `- ${nodeType}: ${name} (${filePath || 'n/a'})`; + }); + + const processLines = processes.map((row: any) => { + const plabel = getRowValue(row, 1, 'label'); + const steps = getRowValue(row, 2, 'stepCount'); + return `- ${plabel} (${steps} steps)`; + }); + + return [ + `CLUSTER: ${label}`, + `Symbols: ${symbolCount ?? members.length}`, + `Cohesion: ${cohesion !== null && cohesion !== undefined ? Number(cohesion).toFixed(2) : 'n/a'}`, + `Description: ${description || 'n/a'}`, + ``, + `TOP MEMBERS:`, + ...(memberLines.length > 0 ? memberLines : ['- None found']), + ``, + `PROCESSES TOUCHING THIS CLUSTER:`, + ...(processLines.length > 0 ? processLines : ['- None found']), + ].join('\n'); + } + + if (resolvedType === 'symbol') { + const nodeId = getRowValue(symbolRow, 0, 'id'); + const name = getRowValue(symbolRow, 1, 'name'); + const filePath = getRowValue(symbolRow, 2, 'filePath'); + const nodeType = getRowValue(symbolRow, 3, 'nodeType'); + + const clusterQuery = ` + MATCH (n:${nodeType} {id: '${String(nodeId).replace(/'/g, "''")}'}) + MATCH (n)-[:CodeRelation {type: 'MEMBER_OF'}]->(c:Community) + RETURN c.label AS label, c.description AS description + LIMIT 1 + `; + const processQuery = ` + MATCH (n:${nodeType} {id: '${String(nodeId).replace(/'/g, "''")}'}) + MATCH (n)-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process) + RETURN p.label AS label, r.step AS step, p.stepCount AS stepCount + ORDER BY r.step + `; + const connectionsQuery = ` + MATCH (n:${nodeType} {id: '${String(nodeId).replace(/'/g, "''")}'}) + OPTIONAL MATCH (n)-[r1:CodeRelation]->(dst) + OPTIONAL MATCH (src)-[r2:CodeRelation]->(n) + RETURN + collect(DISTINCT {name: dst.name, type: r1.type, confidence: r1.confidence}) AS outgoing, + collect(DISTINCT {name: src.name, type: r2.type, confidence: r2.confidence}) AS incoming + LIMIT 1 + `; + + const [clusterRes, processRes, connRes] = await Promise.all([ + executeQuery(clusterQuery), + executeQuery(processQuery), + executeQuery(connectionsQuery), + ]); + + const clusterLabel = clusterRes.length > 0 ? getRowValue(clusterRes[0], 0, 'label') : 'Unclustered'; + const clusterDesc = clusterRes.length > 0 ? getRowValue(clusterRes[0], 1, 'description') : ''; + + const processLines = processRes.map((row: any) => { + const plabel = getRowValue(row, 0, 'label'); + const step = getRowValue(row, 1, 'step'); + const stepCount = getRowValue(row, 2, 'stepCount'); + return `- ${plabel} (step ${step}/${stepCount ?? '?'})`; + }); + + let connections = 'None'; + if (connRes.length > 0) { + const row = connRes[0]; + const rawOutgoing = Array.isArray(row) ? row[0] : (row.outgoing || []); + const rawIncoming = Array.isArray(row) ? row[1] : (row.incoming || []); + const outgoing = (rawOutgoing || []).filter((c: any) => c && c.name).slice(0, 5); + const incoming = (rawIncoming || []).filter((c: any) => c && c.name).slice(0, 5); + + const fmt = (c: any, dir: 'out' | 'in') => { + const conf = c.confidence ? Math.round(c.confidence * 100) : 100; + return dir === 'out' + ? `-[${c.type} ${conf}%]-> ${c.name}` + : `<-[${c.type} ${conf}%]- ${c.name}`; + }; + const outList = outgoing.map((c: any) => fmt(c, 'out')); + const inList = incoming.map((c: any) => fmt(c, 'in')); + if (outList.length || inList.length) { + connections = [...outList, ...inList].join(', '); + } + } + + return [ + `SYMBOL: ${nodeType} ${name}`, + `ID: ${nodeId}`, + `File: ${filePath || 'n/a'}`, + `Cluster: ${clusterLabel}${clusterDesc ? ` β€” ${clusterDesc}` : ''}`, + ``, + `PROCESSES:`, + ...(processLines.length > 0 ? processLines : ['- None found']), + ``, + `CONNECTIONS:`, + connections, + ].join('\n'); + } + + return `Unable to explore "${target}".`; + }, + { + name: 'explore', + description: 'Deep dive on a symbol, cluster, or process. Shows membership, participation, and connections.', schema: z.object({ - nodeIds: z.array(z.string()).describe('Node IDs to highlight (e.g., ["Function:src/utils.ts:calculate"])'), - description: z.string().optional().nullable().describe('What these nodes represent'), + target: z.string().describe('Name or ID of a symbol, cluster, or process'), + type: z.enum(['symbol', 'cluster', 'process']).optional().nullable().describe('Optional target type (auto-detected if omitted)'), }), } ); // ============================================================================ - // TOOL 6: BLAST RADIUS (Impact analysis) + // TOOL 7: IMPACT (Impact analysis) // ============================================================================ - const blastRadiusTool = tool( + const impactTool = tool( async ({ target, direction, maxDepth, relationTypes, includeTests, minConfidence }: { target: string; direction: 'upstream' | 'downstream'; @@ -505,7 +961,7 @@ MATCH (n:Function {id: emb.nodeId}) RETURN n`, r.type AS edgeType, r.confidence AS confidence, r.reason AS reason - LIMIT 100 + LIMIT 300 ` : ` MATCH (target {id: '${targetId.replace(/'/g, "''")}'}) @@ -522,7 +978,7 @@ MATCH (n:Function {id: emb.nodeId}) RETURN n`, r.type AS edgeType, r.confidence AS confidence, r.reason AS reason - LIMIT 100 + LIMIT 300 ` : isFileTarget ? ` @@ -541,7 +997,7 @@ MATCH (n:Function {id: emb.nodeId}) RETURN n`, r.type AS edgeType, r.confidence AS confidence, r.reason AS reason - LIMIT 100 + LIMIT 300 ` : ` MATCH (target {id: '${targetId.replace(/'/g, "''")}'}) @@ -558,10 +1014,10 @@ MATCH (n:Function {id: emb.nodeId}) RETURN n`, r.type AS edgeType, r.confidence AS confidence, r.reason AS reason - LIMIT 100 + LIMIT 300 `; depthQueries.push(executeQuery(d1Query).catch(err => { - if (import.meta.env.DEV) console.warn('Blast radius d=1 query failed:', err); + if (import.meta.env.DEV) console.warn('Impact d=1 query failed:', err); return []; })); @@ -586,7 +1042,7 @@ MATCH (n:Function {id: emb.nodeId}) RETURN n`, r2.type AS edgeType, r2.confidence AS confidence, r2.reason AS reason - LIMIT 100 + LIMIT 200 ` : ` MATCH (target {id: '${targetId.replace(/'/g, "''")}'}) @@ -606,10 +1062,10 @@ MATCH (n:Function {id: emb.nodeId}) RETURN n`, r2.type AS edgeType, r2.confidence AS confidence, r2.reason AS reason - LIMIT 100 + LIMIT 200 `; depthQueries.push(executeQuery(d2Query).catch(err => { - if (import.meta.env.DEV) console.warn('Blast radius d=2 query failed:', err); + if (import.meta.env.DEV) console.warn('Impact d=2 query failed:', err); return []; })); } @@ -637,7 +1093,7 @@ MATCH (n:Function {id: emb.nodeId}) RETURN n`, r3.type AS edgeType, r3.confidence AS confidence, r3.reason AS reason - LIMIT 50 + LIMIT 100 ` : ` MATCH (target {id: '${targetId.replace(/'/g, "''")}'}) @@ -659,10 +1115,10 @@ MATCH (n:Function {id: emb.nodeId}) RETURN n`, r3.type AS edgeType, r3.confidence AS confidence, r3.reason AS reason - LIMIT 50 + LIMIT 100 `; depthQueries.push(executeQuery(d3Query).catch(err => { - if (import.meta.env.DEV) console.warn('Blast radius d=3 query failed:', err); + if (import.meta.env.DEV) console.warn('Impact d=3 query failed:', err); return []; })); } @@ -721,9 +1177,108 @@ MATCH (n:Function {id: emb.nodeId}) RETURN n`, return `No ${direction} dependencies found for "${target}" (types: ${activeRelTypes.join(', ')}). This code appears to be ${direction === 'upstream' ? 'unused (not called by anything)' : 'self-contained (no outgoing dependencies)'}.`; } + const depth1 = byDepth.get(1) || []; + const depth2 = byDepth.get(2) || []; + const depth3 = byDepth.get(3) || []; + + // Confidence buckets + const confidenceBuckets = { high: 0, medium: 0, low: 0 }; + for (const nodes of byDepth.values()) { + for (const n of nodes) { + const conf = n.confidence ?? 1; + if (conf >= 0.9) confidenceBuckets.high += 1; + else if (conf >= 0.8) confidenceBuckets.medium += 1; + else confidenceBuckets.low += 1; + } + } + + // Affected processes and clusters + const maxIdsForContext = 500; + const trimmedIds = allNodeIds.slice(0, maxIdsForContext); + const idList = trimmedIds.map(id => `'${id.replace(/'/g, "''")}'`).join(', '); + let affectedProcesses: Array<{ label: string; hits: number; minStep: number | null; stepCount: number | null }> = []; + let affectedClusters: Array<{ label: string; hits: number; impact: string }> = []; + + if (trimmedIds.length > 0) { + const processQuery = ` + MATCH (s)-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process) + WHERE s.id IN [${idList}] + RETURN p.label AS label, COUNT(DISTINCT s.id) AS hits, MIN(r.step) AS minStep, p.stepCount AS stepCount + ORDER BY hits DESC + LIMIT 20 + `; + const clusterQuery = ` + MATCH (s)-[:CodeRelation {type: 'MEMBER_OF'}]->(c:Community) + WHERE s.id IN [${idList}] + RETURN c.label AS label, COUNT(DISTINCT s.id) AS hits + ORDER BY hits DESC + LIMIT 20 + `; + const directIdList = depth1.map(n => `'${n.id.replace(/'/g, "''")}'`).join(', '); + const directClusterQuery = depth1.length > 0 ? ` + MATCH (s)-[:CodeRelation {type: 'MEMBER_OF'}]->(c:Community) + WHERE s.id IN [${directIdList}] + RETURN DISTINCT c.label AS label + ` : ''; + + const [processRes, clusterRes, directClusterRes] = await Promise.all([ + executeQuery(processQuery), + executeQuery(clusterQuery), + directClusterQuery ? executeQuery(directClusterQuery) : Promise.resolve([]), + ]); + + const directClusterSet = new Set(); + directClusterRes.forEach((row: any) => { + const label = Array.isArray(row) ? row[0] : row.label; + if (label) directClusterSet.add(label); + }); + + affectedProcesses = processRes.map((row: any) => ({ + label: Array.isArray(row) ? row[0] : row.label, + hits: Array.isArray(row) ? row[1] : row.hits, + minStep: Array.isArray(row) ? row[2] : row.minStep, + stepCount: Array.isArray(row) ? row[3] : row.stepCount, + })); + + affectedClusters = clusterRes.map((row: any) => { + const label = Array.isArray(row) ? row[0] : row.label; + const hits = Array.isArray(row) ? row[1] : row.hits; + const impact = directClusterSet.has(label) ? 'direct' : 'indirect'; + return { label, hits, impact }; + }); + } + + const directCount = depth1.length; + const processCount = affectedProcesses.length; + const clusterCount = affectedClusters.length; + let risk = 'LOW'; + if (directCount >= 30 || processCount >= 5 || clusterCount >= 5 || totalAffected >= 200) { + risk = 'CRITICAL'; + } else if (directCount >= 15 || processCount >= 3 || clusterCount >= 3 || totalAffected >= 100) { + risk = 'HIGH'; + } else if (directCount >= 5 || totalAffected >= 30) { + risk = 'MEDIUM'; + } + // ===== COMPACT TABULAR OUTPUT ===== const lines: string[] = [ - `πŸ”΄ BLAST RADIUS: ${target} | ${direction} | ${totalAffected} affected`, + `πŸ”΄ IMPACT: ${target} | ${direction} | ${totalAffected} affected`, + `Confidence: High ${confidenceBuckets.high} | Medium ${confidenceBuckets.medium} | Low ${confidenceBuckets.low}`, + ``, + `AFFECTED PROCESSES:`, + ...(affectedProcesses.length > 0 + ? affectedProcesses.map(p => `- ${p.label} - BROKEN at step ${p.minStep ?? '?'} (${p.hits} symbols, ${p.stepCount ?? '?'} steps)`) + : ['- None found']), + ``, + `AFFECTED CLUSTERS:`, + ...(affectedClusters.length > 0 + ? affectedClusters.map(c => `- ${c.label} (${c.impact}, ${c.hits} symbols)`) + : ['- None found']), + ``, + `RISK: ${risk}`, + `- Direct callers: ${directCount}`, + `- Processes affected: ${processCount}`, + `- Clusters affected: ${clusterCount}`, ``, ]; @@ -767,7 +1322,6 @@ MATCH (n:Function {id: emb.nodeId}) RETURN n`, }; // Depth 1 - Critical (with call site snippets) - const depth1 = byDepth.get(1) || []; if (depth1.length > 0) { const header = direction === 'upstream' ? `d=1 (Directly DEPEND ON ${target}):` @@ -786,7 +1340,6 @@ MATCH (n:Function {id: emb.nodeId}) RETURN n`, } // Depth 2 - High impact - const depth2 = byDepth.get(2) || []; if (depth2.length > 0) { const header = direction === 'upstream' ? `d=2 (Indirectly DEPEND ON ${target}):` @@ -798,7 +1351,6 @@ MATCH (n:Function {id: emb.nodeId}) RETURN n`, } // Depth 3 - Transitive - const depth3 = byDepth.get(3) || []; if (depth3.length > 0) { lines.push(`d=3 (Deep impact/dependency):`); depth3.slice(0, 5).forEach(n => lines.push(formatNode(n))); @@ -811,15 +1363,11 @@ MATCH (n:Function {id: emb.nodeId}) RETURN n`, lines.push(`⚠️ Optional: grep("${target}") for dynamic patterns`); lines.push(``); - // Add the marker for UI highlighting - const marker = `[BLAST_RADIUS:${allNodeIds.join(',')}]`; - lines.push(marker); - return lines.join('\n'); }, { - name: 'blastRadius', - description: `Analyze the blast radius (impact) of changing a function, class, or file. + name: 'impact', + description: `Analyze the impact of changing a function, class, or file. Use when users ask: - "What would break if I changed X?" @@ -838,7 +1386,12 @@ Confidence: 100% = certain, <80% = fuzzy match (may be false positive) relationTypes filter (optional): - Default: CALLS, IMPORTS, EXTENDS, IMPLEMENTS (usage-based) -- Can add CONTAINS, DEFINES for structural analysis`, +- Can add CONTAINS, DEFINES for structural analysis + +Additional output sections: +- Affected processes (with step impact) +- Affected clusters (direct/indirect) +- Risk summary (based on direct callers, processes, clusters)`, schema: z.object({ target: z.string().describe('Name of the function, class, or file to analyze'), direction: z.enum(['upstream', 'downstream']).describe('upstream = what depends on this; downstream = what this depends on'), @@ -859,7 +1412,8 @@ relationTypes filter (optional): cypherTool, grepTool, readTool, - highlightTool, - blastRadiusTool, + overviewTool, + exploreTool, + impactTool, ]; }; diff --git a/gitnexus/src/core/llm/types.ts b/gitnexus/src/core/llm/types.ts index 8df3114b0..d9ca8b30b 100644 --- a/gitnexus/src/core/llm/types.ts +++ b/gitnexus/src/core/llm/types.ts @@ -177,15 +177,13 @@ export interface ToolCallInfo { * Now supports step-based streaming where each step is a distinct message */ export interface AgentStreamChunk { - type: 'reasoning' | 'tool_call' | 'tool_result' | 'content' | 'highlight' | 'error' | 'done'; + type: 'reasoning' | 'tool_call' | 'tool_result' | 'content' | 'error' | 'done'; /** LLM's reasoning/thinking text (shown as a step) */ reasoning?: string; /** Final answer content (streamed token by token) */ content?: string; /** Tool call information */ toolCall?: ToolCallInfo; - /** Node IDs to highlight in the graph */ - highlightNodeIds?: string[]; /** Error message */ error?: string; } diff --git a/gitnexus/src/hooks/useAppState.tsx b/gitnexus/src/hooks/useAppState.tsx index bce50e94e..21d6ef32b 100644 --- a/gitnexus/src/hooks/useAppState.tsx +++ b/gitnexus/src/hooks/useAppState.tsx @@ -1042,10 +1042,10 @@ export const AppStateProvider = ({ children }: { children: ReactNode }) => { } } - // Parse blast radius marker from tool results - const blastMatch = tc.result.match(/\[BLAST_RADIUS:([^\]]+)\]/); - if (blastMatch) { - const rawIds = blastMatch[1].split(',').map((id: string) => id.trim()).filter(Boolean); + // Parse impact marker from tool results + const impactMatch = tc.result.match(/\[IMPACT:([^\]]+)\]/); + if (impactMatch) { + const rawIds = impactMatch[1].split(',').map((id: string) => id.trim()).filter(Boolean); if (rawIds.length > 0 && graph) { const matchedIds = new Set(); const graphNodeIds = graph.nodes.map(n => n.id); diff --git a/gitnexus/src/lib/constants.ts b/gitnexus/src/lib/constants.ts index a5c231b32..efe94286d 100644 --- a/gitnexus/src/lib/constants.ts +++ b/gitnexus/src/lib/constants.ts @@ -18,6 +18,7 @@ export const NODE_COLORS: Record = { Type: '#a78bfa', // Violet light CodeElement: '#64748b', // Slate - muted Community: '#818cf8', // Indigo light - cluster indicator + Process: '#f43f5e', // Rose - execution flow indicator }; // Node sizes by type - clear visual hierarchy with dramatic size differences @@ -39,6 +40,7 @@ export const NODE_SIZES: Record = { Type: 3, // Type alias - small CodeElement: 2, // Generic small Community: 0, // Hidden by default - metadata node + Process: 0, // Hidden by default - metadata node }; // Community color palette for cluster-based coloring diff --git a/gitnexus/src/types/pipeline.ts b/gitnexus/src/types/pipeline.ts index 0c7fec293..123be720b 100644 --- a/gitnexus/src/types/pipeline.ts +++ b/gitnexus/src/types/pipeline.ts @@ -1,7 +1,8 @@ import { GraphNode, GraphRelationship, KnowledgeGraph } from '../core/graph/types'; import { CommunityDetectionResult } from '../core/ingestion/community-processor'; +import { ProcessDetectionResult } from '../core/ingestion/process-processor'; -export type PipelinePhase = 'idle' | 'extracting' | 'structure' | 'parsing' | 'imports' | 'calls' | 'heritage' | 'communities' | 'enriching' | 'complete' | 'error'; +export type PipelinePhase = 'idle' | 'extracting' | 'structure' | 'parsing' | 'imports' | 'calls' | 'heritage' | 'communities' | 'processes' | 'enriching' | 'complete' | 'error'; export interface PipelineProgress { phase: PipelinePhase; @@ -20,6 +21,7 @@ export interface PipelineResult { graph: KnowledgeGraph; fileContents: Map; communityResult?: CommunityDetectionResult; + processResult?: ProcessDetectionResult; } // Serializable version for Web Worker communication diff --git a/todo_live_check_feature.md b/todo_live_check_feature.md new file mode 100644 index 000000000..fceab3bd9 --- /dev/null +++ b/todo_live_check_feature.md @@ -0,0 +1,171 @@ +# Feature Specification: GitNexus "Guardian" (Live Impact Check) + +## 1. Overview +The "Guardian" is an active, background monitoring system that provides real-time feedback to developers as they modify code. It leverages the deterministic Knowledge Graph (KuzuDB) to perform instant "Impact Analysis" and "Architecture Linting" without incurring LLM token costs. + +**Goal:** Provide a "Safety Net" that catches breaking changes, side effects, and architectural violations *before* a commit is made. + +## 2. Core Capabilities + +### A. Live "Blast Radius" Detection +* **Trigger:** + * **Manual:** File Save / Debounced Keystroke. + * **AI-Aware Heuristic:** "Burst Write Cooldown" (See Section 3). +* **Logic:** + 1. Identify modified symbols (Functions, Classes) via incremental Tree-sitter parsing. + 2. Execute Graph Query (Cypher) to find dependents. + * `MATCH (modified)<-[:CALLS*1..5]-(affected) RETURN affected` + 3. Filter "affected" nodes that are outside the current file. +* **User Experience:** + * **Toast/Status Bar:** "⚠️ Modification affects 12 external files." + * **Panel:** List of affected files/functions (e.g., "Breaks `PaymentService.process()`"). + * **"Fix Prompt" Generator:** Button to copy a prompt for the AI agent (e.g., "Check `AuthService.ts` for regressions caused by my changes to `Login.tsx`"). +* **Cost:** **Zero Tokens.** (Pure Graph Traversal). + +### B. Architecture "Linting" +* **Trigger:** File Save / New Import Added. +* **Logic:** + 1. Detect new `IMPORTS` or `CALLS` edges in the graph. + 2. Check against defined "Layer Rules" (e.g., defined in `.gitnexus/rules.yaml`). + * *Rule Example:* `Frontend` cannot import `Database`. + 3. Execute Graph Query: + * `MATCH (source)-[:IMPORTS]->(target) WHERE source.layer = 'Frontend' AND target.layer = 'Database' RETURN source, target` +* **User Experience:** + * **Inline Warning:** "❌ Architectural Violation: UI component cannot directly access Database types." +* **Cost:** **Zero Tokens.** (Rule-based Graph Matching). + +### C. "Smart" Explanation (On Demand) +* **Trigger:** User clicks "Explain Risk" on a warning. +* **Logic:** + 1. Gather context: Source code of the change + Signatures of affected functions. + 2. Send structured prompt to LLM (Small model: GPT-4o-mini / Local Llama). + 3. *Prompt:* "The user modified `calculateTotal()`. This function is called by `InvoiceGenerator`. Explain potential risks." +* **User Experience:** + * Natural language summary: "Changing the return type of `calculateTotal` will cause a compilation error in `InvoiceGenerator` which expects a number." +* **Cost:** **Low.** (Only on user request, highly targeted context). + +## 3. Technical Architecture + +### The "Watcher" Loop +1. **File Watcher:** Listens for `change` events in the workspace. +2. **AI Detection Heuristic (Burst Mode):** + * *Logic:* Detect rapid file writes (3+ files in <2s) typical of AI Agents. + * *Action:* Suppress immediate checks (don't spam). Wait for "Cooldown" (e.g., 5s silence). + * *Trigger:* After cooldown, run "Full Impact Scan" and present summary. +3. **Incremental Parser:** Updates the KuzuDB graph for the specific file (milliseconds). + * *Note:* Uses Tree-sitter's incremental parsing to be extremely fast. +3. **Graph Engine:** Runs pre-compiled Cypher queries against the updated graph. +4. **Client Interface (MCP/Extension):** + * Push notifications to VS Code / Cursor. + * Update "Impact" panel in the browser UI. + +## 4. Agent Integration (Mandatory Workflow) + +To ensure AI agents (Cursor, Claude, etc.) utilize GitNexus, we leverage "Prompt Engineering via Context" and Tool Definitions. + +### A. The `.cursorrules` Protocol +* **Mechanism:** A `.cursorrules` file in the project root enforces agent behavior. +* **Content:** + ```markdown + # GitNexus Safety Protocol + You have access to `@gitnexus` tools. + + **MANDATORY WORKFLOW:** + 1. **PLANNING:** Before writing code, use `gitnexus_get_dependencies` to map the blast radius. + 2. **VERIFICATION:** After EVERY code modification, you MUST run `gitnexus_detect_impact`. + 3. **REPORTING:** If risks are found, fix them or report to user. + ``` + +### B. "Bossy" Tool Descriptions +* **Mechanism:** Update MCP Tool definitions to be imperative. +* **Description:** "CRITICAL SAFETY TOOL. MUST be used after modifying code. Scans graph for broken dependencies. Returns affected files." + +### C. "Fix Prompt" Generator +* **Mechanism:** If the Agent ignores the tools and the "Guardian" detects a break, the popup offers a "Copy Fix Prompt" button. +* **Prompt:** "Your changes to X broke Y. Use `gitnexus_detect_impact` to verify and fix." + +## 5. Distributed Knowledge Graph (Git-Native Architecture) + +To enable "B2B / Team" features without a central server, we use Git itself as the synchronization mechanism for the Knowledge Graph. This is the **"Git-Native Knowledge Graph"** architecture. + +### A. The Core Concept: "Graph Manifest" +You cannot commit the raw KuzuDB database files (binaries) to Git. Instead, we use a lightweight, diff-friendly **Graph Manifest**. + +* **File Path:** `.gitnexus/graph-state.jsonl.gz` +* **Content:** A compressed JSON Lines dump of the Nodes and Edges (semantic data only). +* **Purpose:** Acts as the "Transport Layer" for the graph between machines. + +### B. The Workflow + +#### Step 1: The "Write" Op (Local Dev) +1. **Code Change:** Developer modifies `User.ts`. +2. **Local Indexing:** GitNexus CLI updates local KuzuDB instantly (Incremental Update). +3. **Pre-Commit Hook:** + * Trigger: `git commit` + * Action: GitNexus dumps the current KuzuDB state to `.gitnexus/graph-state.jsonl.gz`. + * Optimization: Only dumps semantic data (e.g., "Func A calls Func B"), not the full AST, keeping it small. + +#### Step 2: The "Transport" (Git Sync) +* `git push` uploads the Code + Graph Manifest. +* **Crucial:** The graph version is now cryptographically tied to the commit hash. No "drift" between code and graph. + +#### Step 3: The "Read" Op (Teammate Pull) +1. **Git Pull:** Teammate receives new code + new manifest. +2. **Post-Merge Hook / Hydration:** + * Trigger: Git detects change in `.gitnexus`. + * Action: GitNexus CLI reads the manifest and **bulk-inserts** it into the local KuzuDB. + * **Result:** Teammate has a fully queried graph in seconds (vs. minutes of re-parsing). + +### C. Conflict Resolution: "Discard and Rebuild" +What happens if two devs change the graph simultaneously? + +* **Scenario:** Merge conflict in `.gitnexus/graph-state.jsonl.gz`. +* **Strategy:** + 1. GitNexus detects the conflict in the manifest file. + 2. It **discards** the conflicted manifest. + 3. It runs the **Parser** locally on the *merged* source code (Source of Truth). + 4. It generates a **fresh, correct** manifest. +* **Philosophy:** The Graph is a *derivative* of the Code. We never manually merge the graph; we regenerate it from the source. + +### D. Architecture Diagram (Mermaid) + +```mermaid +graph TD + subgraph "Developer A (Write)" + CodeA[User.ts] -->|Parser| DB_A[(Local KuzuDB)] + DB_A -->|Pre-Commit Export| ManifestA[.gitnexus/graph-state.gnx] + ManifestA -->|git push| GitHub + end + + subgraph "GitHub / Git Server" + GitHub -->|git pull| DevB_Repo + end + + subgraph "Developer B (Read)" + DevB_Repo[Code + Manifest] -->|Hydration Hook| DB_B[(Local KuzuDB)] + ManifestA -->|Bulk Insert| DB_B + DB_B -->|Instant Query| Cursor_B[Cursor / IDE] + end + + subgraph "Enterprise Hub (Monetization)" + GitHub -->|Webhook| HubServer[Node.js Hub] + HubServer -->|Download Manifest| CentralDB[(Neo4j / Postgres)] + CentralDB -->|Analytics API| Dashboard[CTO Dashboard] + end +``` + +### E. B2B / Enterprise "Hub" Integration +The "Hub" is a lightweight server that monetizes this architecture. +1. **Action:** Subscribes to the repo's webhooks. +2. **Ingestion:** Downloads *only* the manifest file (not the source code). +3. **Storage:** Loads it into a centralized DB (Neo4j/Postgres) for organization-wide queries. +4. **Security:** "We don't see your code, only your graph structure." + + +## 7. Implementation Roadmap +1. **Phase 1:** Implement `FileWatcher` in CLI + Incremental Graph Update. +2. **Phase 2:** Create `ImpactQuery` engine (Cypher queries for dependents). +3. **Phase 3:** Build MCP Tool `get_live_impact` for Editor integration and "Bossy" descriptions. +4. **Phase 4:** Implement Git-Native Graph Sync (Manifest generation + Pre-commit hook). +5. **Phase 5:** Add Architecture Rule definition schema (`.gitnexus/rules.yaml`). +