feat(skills): add ce-plan — GitNexus+PDG implementation-planning skill

Adds .claude/skills/ce-plan: a planning-only skill that builds
implementation-ready plans from GitNexus graph navigation (query/context/
impact/trace), bounded statement-level PDG slices (pdg_query, impact
mode:pdg, explain), and targeted source verification, with a context
ledger to prevent repeated reads and a machine-readable implementation
context pack (stable contract for a future ce-implement). Whitelisted in
.gitignore and registered in AGENTS.md and CLAUDE.md outside the
auto-managed gitnexus block.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Gergo Magyar 2026-07-11 06:53:29 +00:00
parent b249aa4c2d
commit aeddd88c25
9 changed files with 623 additions and 0 deletions

View file

@ -0,0 +1,70 @@
# ce-plan — Compound Engineering Plan
Generates deep, implementation-ready engineering plans by combining GitNexus
repository intelligence, statement-level Program Dependence Graph analysis,
and Claude Code's native targeted source verification.
## Invocation
```
/ce-plan Add retry support to the ingestion pipeline
/ce-plan Fix the stale warm-cache invalidation bug in exportedTypeMap
/ce-plan depth:deep impact_depth:3 Migrate the emit phase to streaming COPY
```
Output: `docs/plans/YYYY-MM-DD-ce-plan-<slug>.md` — a 13-section plan whose
section 11 is a machine-readable **implementation context pack** that a
follow-up agent can consume without re-investigating the repository.
## Architecture note: how GitNexus and Claude Code interact
Three layers, strictly ordered:
1. **GitNexus navigates** (`query` → `context` → `impact`/`trace` →
`cypher` last-resort). The graph answers *where to look* and *what is
connected*: execution flows, callers/callees, blast radius, related tests.
Every call must answer a named planning question.
2. **PDG constrains** (`pdg_query` controls/flows, `impact {mode:"pdg",
line}` statement slices, `explain` for taint). The statement-level layers
answer *what gates and feeds the behavior* inside the few functions the
change centers on. Results are filtered into a bounded slice
(`references/pdg-slice.md`), never dumped.
3. **Claude Code verifies** (targeted line-range reads). Current source is
authoritative; graph results are navigation hints until verified. On
disagreement: trust source, record the discrepancy, recommend re-indexing.
Token efficiency comes from the **context ledger**
(`references/context-ledger.md`): every query and read is recorded with the
question it answered, and nothing is re-fetched unless the source changed or a
contradiction surfaced. The ledger also enforces symbol budgets (5 primary /
20 related by default), and progressive disclosure keeps the big schemas out
of context until the phase that needs them.
## Files
| File | Purpose |
| --- | --- |
| `SKILL.md` | The skill: phases 0–5, hard rules, config, fallback |
| `references/pdg-slice.md` | PDG slice construction: tools, inclusion criteria, schema, security/performance modes |
| `references/context-ledger.md` | Ledger schema + anti-reread rules |
| `references/plan-template.md` | The 13-section plan document template |
| `references/context-pack.md` | Implementation context pack schema + stability contract |
## Requirements and graceful degradation
- Requires a GitNexus index (`node .gitnexus/run.cjs analyze`); statement-level
sections additionally require `analyze --pdg`.
- No PDG layer → the plan says so and skips statement-level claims (never
reconstructs fake edges).
- No GitNexus at all → fallback mode: targeted grep/read exploration, findings
labelled **source-derived**, with a recommendation to index.
## Limitations
- `pdg_query` is intra-procedural; cross-function flow comes from `explain`
(taint) or `impact {mode:"pdg"}` inter-procedural reach.
- If the `compound-engineering` plugin is installed alongside this repo skill,
both expose a skill named `ce-plan` (the plugin's under the
`compound-engineering:` namespace). Invoke this one as the bare `/ce-plan`;
disambiguate by full name if your harness prompts.
- The skill is read-only by contract; it will not fix what it finds.

View file

@ -0,0 +1,191 @@
---
name: ce-plan
description: "Use when you need a deep, implementation-ready engineering plan for a code change — built from GitNexus graph intelligence, statement-level PDG analysis, and targeted source verification, compact enough that an implementation agent can start without re-investigating. Examples: \"/ce-plan Add retry support to the ingestion pipeline\", \"/ce-plan Fix the stale warm-cache invalidation bug\", \"plan this change using the knowledge graph\"."
---
# ce-plan — Compound Engineering Plan
Produce an implementation-ready plan for an engineering task. GitNexus is the
navigation layer (where to look), statement-level PDG is the constraint layer
(what gates and feeds the behavior), Claude Code source reads are the
verification layer (what is actually true right now). The output is a plan
document plus a compact, machine-readable **implementation context pack** that
a follow-up agent (`ce-implement` or any executor) can consume without
repeating the investigation.
```
/ce-plan <task description>
/ce-plan impact_depth:3 depth:deep <task description> # knob overrides, see Configuration
```
**This skill plans. It never implements.** Do not modify production code,
tests, or configuration while running it. The only file it writes is the plan
document.
## Hard rules
- **Ledger first.** Before every GitNexus call and every file read, check the
context ledger (below). Never repeat a query or reread an unchanged range
that already answered the same question.
- **Every graph query answers a named planning question.** Record the question
and the conclusion in the ledger. No exploratory dredging.
- **Source beats graph.** The graph navigates; current source is authoritative.
Verify before asserting (see Phase 4).
- **No fabrication.** Never invent symbols, filenames, test names, tool
results, or PDG edges. Unknowns go to *Assumptions and Open Questions*.
- **Stop when you have enough.** Sufficient evidence ends exploration; plans
do not improve monotonically with tokens spent.
## Phase 0 — Parse and classify
Restate the task as an interpreted goal plus acceptance criteria (open the
ledger with these). Classify it:
| Category | Depth posture |
| --- | --- |
| Bug fix (local) | Narrow: 1–2 primary symbols, impact depth 1–2 |
| Feature | Default knobs |
| Refactor / shared API change | Impact analysis mandatory, impact depth 3 |
| Performance | Default + performance PDG mode (see `references/pdg-slice.md`) |
| Security | Default + security PDG mode + `explain` taint findings |
| Dependency upgrade / migration | Impact + compatibility focus; PDG usually unnecessary |
| Concurrency / transactional | Control-flow and state-mutation focus in the PDG slice |
| Test improvement / docs | Narrowest: usually no impact or PDG pass |
| Architecture change / spike | Widest: clusters + processes resources first |
Raise a bounded-depth knob only when the change clearly crosses architectural
boundaries; say so in the plan when you do.
## Phase 1 — Anchor and freshness
1. Resolve the target repo: `list_repos` if in doubt, else the indexed repo
covering the working directory. Pass `repo` explicitly on every call when
more than one repo is indexed.
2. Read `gitnexus://repo/{name}/context` — codebase overview + staleness check.
- Stale index → recommend `node .gitnexus/run.cjs analyze`, note the
staleness in the plan's Assumptions, and continue with source
verification weighted higher.
- GitNexus unavailable entirely → switch to **Fallback mode** (below).
3. For architecture-scale tasks only, also read
`gitnexus://repo/{name}/clusters` and `.../processes`.
## Phase 2 — Graph navigation ladder
Use the narrowest operation that answers the current ledger question, in this
order. Budgets: at most `max_primary_symbols` (5) primary symbols and
`max_related_symbols` (20) related symbols enter the ledger.
1. `query {search_query, task_context}` — locate concepts, execution flows,
modules, and related tests for the task.
2. `context {name}` — 360° view of each candidate primary symbol: callers,
callees, categorized refs, processes. Promote to primary or discard.
3. `impact {target, direction}` — upstream/downstream blast radius for shared
or high-connectivity symbols (`maxDepth` = `impact_depth`; `summaryOnly:
true` first for hub symbols). Record d=1 items — the plan must account for
every one of them.
4. `trace {from, to}` — when the task hinges on *how A reaches B*, one call
instead of chained context hops.
5. Statement-level PDG — Phase 3, for the functions the change centers on.
6. `cypher` — last resort, only for a precise graph question the tools above
cannot express. Read `gitnexus://repo/{name}/schema` first; anchor and
LIMIT every query.
7. `detect_changes {scope}` — only when planning against existing uncommitted
or branch work.
Do not run every tool by default. A local test fix may finish the ladder at
step 2.
## Phase 3 — Statement-level PDG slice
For the 1–3 functions most central to the change, build a **PDG context
slice**: read `references/pdg-slice.md` and follow it. In short:
- `pdg_query {mode: "controls", target}` — what gates the behavior (guards,
branch senses, early exits).
- `pdg_query {mode: "flows", target, variable?}` — def→use flow of the
variables the change touches.
- `impact {mode: "pdg", target, line}` — statement-anchored dependence slice
plus inter-procedural reach, when one statement is the seed of the change.
- Security tasks add `explain` (taint findings); performance tasks add the
loop/blocking-call checklist.
Filter hard: only statements meeting the slice inclusion criteria enter the
plan, bounded by `pdg_data_depth`/`pdg_control_depth` (2). Never paste a raw
PDG dump. No `--pdg` layer indexed → record "PDG unavailable" in the ledger,
skip to Phase 4, and say so in the plan; do not reconstruct fake edges.
## Phase 4 — Targeted source verification
GitNexus said where to look; now confirm what is there. Using ordinary file
reads (exact line ranges, not whole files unless genuinely required):
- Read every source range the plan will cite: signatures, branch conditions,
state mutations, error paths, nearby comments that change behavior.
- Read the tests GitNexus associated with the primary symbols; never claim a
test exists without having located it.
- Check repo conventions that constrain the change (AGENTS.md, GUARDRAILS.md,
lint/build config) — only the parts the change touches.
- Mark each ledger symbol `source_verified: true` as you go. **A symbol that
is named in Proposed Changes must be source-verified.**
- On graph/source disagreement: trust source, record the discrepancy in the
ledger and the plan, recommend re-indexing. Never present stale graph data
as fact.
Evidence hierarchy, strongest first: current source and config → current tests
and executable behavior → compiler/build/lint output → GitNexus graph and PDG
→ documentation and comments.
## Phase 5 — Compose the plan
1. Read `references/plan-template.md` and fill all 13 sections from the
ledger. Distinguish **confirmed facts / evidence-backed inferences /
assumptions / open questions** throughout.
2. Build the implementation context pack per `references/context-pack.md`
(this is section 11 of the plan).
3. Write the document to `docs/plans/YYYY-MM-DD-ce-plan-<slug>.md` (create the
directory if missing; kebab-case slug, 3–5 words). Repo-relative paths
everywhere.
4. Present in chat: objective, proposed-changes summary, implementation
sequence, top risks, open questions, and the plan file path. Do not paste
the whole document into chat.
## Context ledger
Maintain the ledger from Phase 0 onward — it is the skill's working memory and
its token budget enforcement. Schema and reread rules: `references/context-ledger.md`.
The final ledger feeds the plan; it is not itself published.
## Configuration
Defaults; override inline with `key:value` tokens before the task text (the
repo has no skill-config file mechanism — invocation args are the mechanism):
| Knob | Default | Meaning |
| --- | --- | --- |
| `depth` | by category | `narrow` / `default` / `deep` posture |
| `call_depth` | 2 | Caller/callee expansion in `context`/`trace` reasoning |
| `impact_depth` | 2 | `maxDepth` for `impact` |
| `pdg_data_depth` | 2 | Data-dependence hops in the PDG slice |
| `pdg_control_depth` | 2 | Control-dependence hops in the PDG slice |
| `max_primary_symbols` | 5 | Ledger budget |
| `max_related_symbols` | 20 | Ledger budget |
| `max_snippet_lines` | 30 | Longest source excerpt quoted in the plan |
## Fallback mode (GitNexus or PDG unavailable)
1. Say so, first thing, in chat and in the plan.
2. Use targeted repo exploration (grep/glob/reads) to approximate callers,
dependencies, execution flow, state changes, and related tests.
3. Label every such finding **source-derived** in the plan — never present it
as graph-derived, and never fabricate statement-level edges.
4. Recommend `node .gitnexus/run.cjs analyze` (add `--pdg` for the PDG layers)
when it would materially raise confidence.
## Never
- Implement the feature, edit production code, or run mutating commands.
- Dump unfiltered graph/PDG output or full files into the plan.
- Repeat a search or reread an unchanged range already in the ledger.
- Expand scope into unrelated refactoring.
- Treat comments as stronger evidence than executable code.
- Continue exploring after the ledger answers all open planning questions.

View file

@ -0,0 +1,66 @@
# Context ledger
The ledger is ce-plan's working memory. It exists to make repeated
investigation impossible-by-discipline: **before every GitNexus call and every
file read, check it.** Keep it as structured notes in your working context (or
a scratchpad file for very long sessions); it is never published verbatim —
the plan and context pack are distilled from it.
## Schema
```yaml
context_ledger:
task:
original_request: ""
interpreted_goal: ""
category: "" # Phase 0 classification
acceptance_criteria: []
established_facts: [] # each with its evidence source
symbols: # budgeted: max_primary_symbols / max_related_symbols
- name: ""
kind: ""
file: ""
relevance: "primary | related | discarded"
source_verified: false # flipped in Phase 4; required before naming in Proposed Changes
files_read:
- file: ""
ranges: [] # e.g. ["120-188"]
purpose: ""
content_hash_or_version: "" # e.g. git blob hash or "HEAD@<sha>"
gitnexus_queries:
- query: "" # tool + args
purpose: "" # the planning question it answers
conclusion: "" # one line; details stay in working memory
pdg_slices:
- symbol: ""
purpose: ""
conclusion: ""
unresolved_questions: []
assumptions: [] # explicit, carried into plan §12
decisions: [] # with rationale, carried into plan §6/§7
```
## Reread rules
Do **not** repeat a query or reread a source range unless one of:
- the previous result was incomplete for a *new* question;
- the source changed (hash/version mismatch, or an edit is known to have
happened);
- validation exposed a contradiction between graph and source.
When a repeat is justified, note in the ledger *why* the earlier entry was
insufficient. A ledger full of near-duplicate queries is the failure signal —
stop and plan with what is established.
## Discarding
Symbols and queries that turned out irrelevant stay in the ledger marked
`discarded` with a one-line reason. That is what prevents re-walking dead
ends later in the session.

View file

@ -0,0 +1,71 @@
# Implementation context pack
Section 11 of the plan. The stable, machine-readable contract a follow-up
implementation agent (`ce-implement` or any executor) consumes to start work
**without repeating the investigation**. Distilled from the ledger; every
entry traceable to verified evidence.
## Schema
```yaml
implementation_context:
task_summary: ""
acceptance_criteria: []
primary_symbols:
- symbol: ""
file: ""
lines: ""
role: ""
related_symbols:
- symbol: ""
relationship: "" # CALLS / IMPORTS / EXTENDS / test-of / ...
relevance: ""
execution_path: [] # ordered prose steps, from §2/§5
pdg_constraints: # from the PDG slice; empty + note if no layer
- description: ""
affected_statements: [] # "<file>:<line>" refs
implementation_consequence: ""
architectural_patterns:
- pattern: ""
example_location: "" # repo-relative file (+ symbol)
usage_guidance: ""
files_to_modify:
- file: ""
symbols: []
intended_change: ""
tests:
- file: "" # existing file to update, or new path to create
scenarios: [] # input → action → expected outcome
verification_commands: [] # real commands from repo scripts/CI, verified to exist
risks: []
assumptions: [] # verbatim from plan §12
avoid:
- "Do not repeat full repository discovery"
- "Do not replace established patterns without evidence"
# + task-specific prohibitions discovered during planning
```
## Must not contain
- full files;
- large raw GitNexus responses;
- unfiltered PDG dumps;
- duplicate code excerpts (cite `file:line`, don't re-quote);
- speculative implementation details presented as facts.
## Stability contract
Field names above are the interface for a future `ce-implement`. Add fields
freely; do not rename or repurpose existing ones. `assumptions` and `avoid`
are load-bearing: an executor treats `assumptions` as things to re-verify
cheaply before relying on them, and `avoid` as hard constraints.

View file

@ -0,0 +1,92 @@
# Building the PDG context slice
Statement-level evidence for the 1–3 functions most central to the change.
Goal: a compact slice the planning LLM can hold, never a graph dump.
## Tools (all verified against `gitnexus/src/mcp/tools.ts`)
| Question | Call |
| --- | --- |
| Under what condition does X run? Guards? | `pdg_query {mode: "controls", target}` |
| Where does variable Y flow inside the function? | `pdg_query {mode: "flows", target, variable}` |
| What depends on the statement at line N? | `impact {mode: "pdg", target, line: N}` |
| Source→sink taint paths (security mode) | `explain {target}` |
Contract caveats that shape interpretation:
- CDG branch sense is `'T'`/`'F'` in the edge's `reason`; a guard's sense
depends on its predicate (`if (!ok) return;` rides `'T'`) — never filter
guards by a fixed label. Early return/throw edges carry `guard: true`.
- `pdg_query` is intra-procedural and always anchored. Cross-function flow is
taint's domain (`explain`) or `impact {mode:"pdg"}`'s inter-procedural reach.
- Every `switch` case arm is `'T'` (per-case conditions not distinguished).
- No `--pdg` layer → the tools return a "no PDG layer" note, not an error.
Record it, skip the slice, recommend `analyze --pdg`. Do not reconstruct
edges from source by hand.
## Inclusion criteria
A statement enters the slice only if it is at least one of:
- directly matched to the task;
- a data-flow predecessor or successor of a relevant statement (within
`pdg_data_depth`, default 2);
- a control dependency of a relevant statement (within `pdg_control_depth`,
default 2);
- a state mutation affecting the requested behavior;
- an external call on the execution path;
- an error-handling or fallback branch;
- part of an affected return value;
- required to explain a test assertion.
Everything else is cut. If the slice exceeds ~15 statements per function,
tighten relevance rather than raising depth.
## Slice representation (goes into the ledger; summarized in plan §5)
```yaml
pdg_context:
entry_symbol: "processFileGroup"
source: { file: "gitnexus/src/core/ingestion/worker.ts", start_line: 120, end_line: 188 }
relevant_statements:
- id: "stmt-12" # stable id or "<file>:<line>"
lines: "128-130"
type: "condition | call | mutation | return | throw"
code: "if (request.retryable) {"
relevance: "Controls whether retry scheduling is entered"
defines: []
uses: ["request.retryable"]
control_dependencies: ["stmt-4"]
data_dependencies: []
execution_flow: # ordered, prose steps
- "Validate request"
- "Schedule retry"
critical_dependencies:
- { from: "stmt-7", to: "stmt-18", type: "data", explanation: "Validated request becomes scheduler input" }
behavioural_observations:
- "Persistence occurs before scheduler invocation"
planning_implications:
- "Changes to scheduling must account for partial failure"
```
Adapt field names to what the tools actually returned; keep it
machine-readable and short. `behavioural_observations` are confirmed facts;
`planning_implications` are inferences — keep the distinction.
## Security mode (task category: security)
Additionally identify and record: untrusted inputs, validation points,
sanitisation points, authn/authz checks, privilege boundaries, sensitive data,
persistence operations, network calls, dangerous sinks, and error paths that
bypass validation. Run `explain {target}` for persisted source→sink taint
paths and include the hop paths for findings relevant to the task. Absence of
a taint finding is **not** proof of safety (intra-procedural only) — say so
when it matters.
## Performance mode (task category: performance)
Additionally scan the slice for: loops, repeated calls, blocking operations,
network calls, database calls, allocation-heavy paths, caching boundaries,
concurrency, fan-out, repeated data transformations. State likely hot-path
implications as inferences; never claim measured improvements without
benchmark evidence.

View file

@ -0,0 +1,122 @@
# Plan document template
Fill every section. If a section is genuinely empty for this task (e.g. no PDG
layer indexed), keep the heading and state why in one line — never silently
drop it. Repo-relative paths everywhere. Distinguish confirmed facts,
evidence-backed inferences, assumptions, and open questions throughout.
```markdown
# Compound Engineering Plan
## 1. Objective
A concise description of the requested outcome.
## 2. Current Behaviour
Describe the current implementation and execution path.
Include the most relevant symbols, files, and statement-level observations.
## 3. Relevant Architecture
Explain the involved modules, boundaries, dependencies, and established patterns.
## 4. GitNexus Findings
Summarise:
- primary symbols;
- callers and callees;
- impact radius;
- related implementations;
- related tests;
- important cross-module relationships.
## 5. Statement-Level PDG Findings
For each critical symbol, explain:
- relevant statements;
- control dependencies;
- data dependencies;
- state mutations;
- error branches;
- side effects;
- ordering constraints;
- planning implications.
Do not paste an unfiltered graph dump.
## 6. Proposed Changes
For every proposed change include:
- file;
- symbol;
- exact responsibility;
- intended behavioural change;
- dependencies;
- constraints;
- implementation notes.
## 7. Implementation Sequence
Provide an ordered sequence of implementation steps.
Each step must be independently actionable.
## 8. Test Strategy
Describe:
- tests to add;
- tests to update;
- edge cases;
- failure paths;
- regression coverage;
- integration boundaries;
- relevant verification commands.
## 9. Risk and Impact Analysis
Include:
- high-risk symbols;
- downstream consumers;
- compatibility concerns;
- performance concerns;
- concurrency or transaction risks;
- migration risks;
- observability requirements.
## 10. Files Expected to Change
| File | Symbols | Reason |
|---|---|---|
## 11. Reusable Implementation Context
The machine-readable context pack — see `context-pack.md`.
## 12. Assumptions and Open Questions
Clearly separate assumptions from confirmed facts.
## 13. Definition of Done
Concrete, testable completion criteria.
```
Composition notes:
- §2/§5 quote source excerpts at most `max_snippet_lines` (30) lines each, and
only when the excerpt carries the argument.
- §4 findings each name the tool call they came from (traceable to the
ledger); stale-index or fallback-mode findings are labelled as such.
- §6 changes may only name symbols the ledger marks `source_verified`.
- §7 steps are ordered by dependency and independently actionable — an
executor can stop after any step with the tree still coherent.
- §8 names real, located test files for updates; new tests get concrete
scenario lists (input → action → expected outcome).
- §9 must account for every d=1 symbol the impact pass reported.

1
.gitignore vendored
View file

@ -98,6 +98,7 @@ gitnexus/vendor/**/node_modules/
.claude/skills/*
!.claude/skills/gitnexus/
!.claude/skills/gitnexus-pr-swarm-review/
!.claude/skills/ce-plan/
.history/

View file

@ -55,6 +55,15 @@ listed in [`pr-swarm-review/README.md`](pr-swarm-review/README.md); edit review
in the canonical files, never in the wrappers. The review is read-only — it never edits,
commits, or posts.
## Engineering planning (`/ce-plan`)
To produce a deep, implementation-ready plan for a code change, invoke the **`ce-plan`**
skill (`.claude/skills/ce-plan/SKILL.md`): GitNexus graph intelligence for navigation,
statement-level PDG slices for behavioral constraints, targeted source reads for
verification. Output lands in `docs/plans/` with a reusable implementation context pack
(section 11) that a follow-up implementation agent consumes without re-investigating.
The skill is planning-only — it never edits code.
## Changelog
| Date | Version | Change |

View file

@ -37,6 +37,7 @@ If always-on instructions grow, load deep conventions via conditional reads (e.g
- **This repository:** [AGENTS.md](AGENTS.md) (Cursor + monorepo notes), [ARCHITECTURE.md](ARCHITECTURE.md), [CONTRIBUTING.md](CONTRIBUTING.md), [GUARDRAILS.md](GUARDRAILS.md).
- **Call & inheritance resolution:** See ARCHITECTURE.md § Scope-Resolution Pipeline. Shared pipeline code in `gitnexus/src/core/ingestion/` must not name languages — use `LanguageProvider` / `ScopeResolver` hooks instead (see AGENTS.md). (The legacy call-resolution DAG was removed in #942.)
- **GitNexus:** `.claude/skills/gitnexus/`; MCP and indexed-repo rules live only in [AGENTS.md](AGENTS.md) (`gitnexus:start` … `gitnexus:end`). See **GitNexus rules** below.
- **Engineering plans:** `/ce-plan <task>` — implementation-ready plans via GitNexus + statement-level PDG + source verification; spec in `.claude/skills/ce-plan/SKILL.md` (see AGENTS.md § Engineering planning).
## Changelog