fix(skills): apply ce-plan validation findings (tool contract, consistency, conventions)

Tool contract: impact mode:'pdg' shape now includes the schema-required
direction param; CDG branch sense documented as the result 'label' field
(reason is cypher/raw-edge only); explain caveats corrected to its real
false-negative classes (cross-function TAINT_PATH is modeled).

Consistency: PDG slice homed in working memory (ledger keeps one-liners);
depth knob defined and category-overrides-baseline ordering stated;
call_depth (consumed by nothing) and content-hash bookkeeping dropped;
Never section folded into Hard rules; Phase 3 deduplicated to a pointer;
allowed-repeat escalations defined; budget/discard accounting clarified;
verification-commands gathering added to Phase 4; open_questions added to
the context pack.

From scenario runs: plans now pin the verified-at HEAD commit and index
freshness in a header, tag claims [verified]/[graph]/[inferred]/[assumed],
quote load-bearing tool output, prefer pre-hook-carrying npm scripts, and
support an out:<path> destination override; output path defined as the
Phase 1 target repo root.

Conventions: AGENTS.md 1.9.0 / CLAUDE.md 1.4.0 changelog rows + metadata
bumps; future ce-implement qualified as future.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Gergo Magyar 2026-07-11 07:18:04 +00:00
parent aeddd88c25
commit b34b171100
8 changed files with 133 additions and 98 deletions

View file

@ -25,7 +25,8 @@ Three layers, strictly ordered:
connected*: execution flows, callers/callees, blast radius, related tests.
Every call must answer a named planning question.
2. **PDG constrains** (`pdg_query` controls/flows, `impact {mode:"pdg",
line}` statement slices, `explain` for taint). The statement-level layers
direction, line}` statement slices, `explain` for taint). The
statement-level layers
answer *what gates and feeds the behavior* inside the few functions the
change centers on. Results are filtered into a bounded slice
(`references/pdg-slice.md`), never dumped.

View file

@ -9,9 +9,9 @@ Produce an implementation-ready plan for an engineering task. GitNexus is the
navigation layer (where to look), statement-level PDG is the constraint layer
(what gates and feeds the behavior), Claude Code source reads are the
verification layer (what is actually true right now). The output is a plan
document plus a compact, machine-readable **implementation context pack** that
a follow-up agent (`ce-implement` or any executor) can consume without
repeating the investigation.
document plus a compact, machine-readable **implementation context pack**
that a follow-up implementation agent (a future `ce-implement`, or any
executor) can consume without repeating the investigation.
```
/ce-plan <task description>
@ -19,33 +19,38 @@ repeating the investigation.
```
**This skill plans. It never implements.** Do not modify production code,
tests, or configuration while running it. The only file it writes is the plan
document.
tests, or configuration while running it. The only repository file it writes
is the plan document (a working ledger kept outside the repo is fine).
## Hard rules
- **Ledger first.** Before every GitNexus call and every file read, check the
context ledger (below). Never repeat a query or reread an unchanged range
that already answered the same question.
- **Ledger first.** Before every GitNexus call and every repo file read, check
the context ledger. Never repeat a query or reread an unchanged range that
already answered the same question (allowed repeats are defined in
`references/context-ledger.md`; this skill's own reference files are exempt
from ledger bookkeeping).
- **Every graph query answers a named planning question.** Record the question
and the conclusion in the ledger. No exploratory dredging.
- **Source beats graph.** The graph navigates; current source is authoritative.
Verify before asserting (see Phase 4).
Verify before asserting (see Phase 4). Comments are the weakest evidence —
never stronger than executable code.
- **No fabrication.** Never invent symbols, filenames, test names, tool
results, or PDG edges. Unknowns go to *Assumptions and Open Questions*.
- **No scope creep.** Adjacent refactors the task didn't ask for go to plan
§12 as suggestions, not into Proposed Changes.
- **Stop when you have enough.** Sufficient evidence ends exploration; plans
do not improve monotonically with tokens spent.
## Phase 0 — Parse and classify
Restate the task as an interpreted goal plus acceptance criteria (open the
ledger with these). Classify it:
Read `references/context-ledger.md` and open the ledger with the task:
original request, interpreted goal, acceptance criteria. Classify the task:
| Category | Depth posture |
| --- | --- |
| Bug fix (local) | Narrow: 1–2 primary symbols, impact depth 1–2 |
| Bug fix (local) | Narrow: 1–2 primary symbols, `impact_depth` 1 |
| Feature | Default knobs |
| Refactor / shared API change | Impact analysis mandatory, impact depth 3 |
| Refactor / shared API change | Impact analysis mandatory, `impact_depth` 3 |
| Performance | Default + performance PDG mode (see `references/pdg-slice.md`) |
| Security | Default + security PDG mode + `explain` taint findings |
| Dependency upgrade / migration | Impact + compatibility focus; PDG usually unnecessary |
@ -53,36 +58,44 @@ ledger with these). Classify it:
| Test improvement / docs | Narrowest: usually no impact or PDG pass |
| Architecture change / spike | Widest: clusters + processes resources first |
Raise a bounded-depth knob only when the change clearly crosses architectural
boundaries; say so in the plan when you do.
The category posture overrides the Configuration baseline; explicit `key:value`
invocation knobs override both. A task matching several rows combines them:
take the widest depth, union the focus areas.
## Phase 1 — Anchor and freshness
1. Resolve the target repo: `list_repos` if in doubt, else the indexed repo
covering the working directory. Pass `repo` explicitly on every call when
more than one repo is indexed.
2. Read `gitnexus://repo/{name}/context` — codebase overview + staleness check.
2. Record the repo's current HEAD commit in the ledger — every line-number
citation in the plan is pinned to it.
3. Read `gitnexus://repo/{name}/context` — codebase overview + staleness check.
- Stale index → recommend `node .gitnexus/run.cjs analyze`, note the
staleness in the plan's Assumptions, and continue with source
verification weighted higher.
- Resources unreadable but tools working → proceed on tools alone, treat
freshness as unknown (weight source higher), and note it in the plan.
- GitNexus unavailable entirely → switch to **Fallback mode** (below).
3. For architecture-scale tasks only, also read
4. For architecture-scale tasks only, also read
`gitnexus://repo/{name}/clusters` and `.../processes`.
## Phase 2 — Graph navigation ladder
Use the narrowest operation that answers the current ledger question, in this
order. Budgets: at most `max_primary_symbols` (5) primary symbols and
`max_related_symbols` (20) related symbols enter the ledger.
`max_related_symbols` (20) related symbols active in the ledger.
1. `query {search_query, task_context}` — locate concepts, execution flows,
modules, and related tests for the task.
2. `context {name}` — 360° view of each candidate primary symbol: callers,
callees, categorized refs, processes. Promote to primary or discard.
callees, categorized refs, processes. Promote to primary or discard. An
`ambiguous` result (ranked candidates) is answered by one retry narrowed
with `kind` / `file_path` / uid — that retry is an allowed repeat.
3. `impact {target, direction}` — upstream/downstream blast radius for shared
or high-connectivity symbols (`maxDepth` = `impact_depth`; `summaryOnly:
true` first for hub symbols). Record d=1 items — the plan must account for
every one of them.
true` first for hub symbols, then drill in — an allowed repeat). Record the
d=1 items — the **direct (depth-1) dependents** — the plan must account
for every one of them.
4. `trace {from, to}` — when the task hinges on *how A reaches B*, one call
instead of chained context hops.
5. Statement-level PDG — Phase 3, for the functions the change centers on.
@ -97,22 +110,10 @@ step 2.
## Phase 3 — Statement-level PDG slice
For the 1–3 functions most central to the change, build a **PDG context
slice**: read `references/pdg-slice.md` and follow it. In short:
- `pdg_query {mode: "controls", target}` — what gates the behavior (guards,
branch senses, early exits).
- `pdg_query {mode: "flows", target, variable?}` — def→use flow of the
variables the change touches.
- `impact {mode: "pdg", target, line}` — statement-anchored dependence slice
plus inter-procedural reach, when one statement is the seed of the change.
- Security tasks add `explain` (taint findings); performance tasks add the
loop/blocking-call checklist.
Filter hard: only statements meeting the slice inclusion criteria enter the
plan, bounded by `pdg_data_depth`/`pdg_control_depth` (2). Never paste a raw
PDG dump. No `--pdg` layer indexed → record "PDG unavailable" in the ledger,
skip to Phase 4, and say so in the plan; do not reconstruct fake edges.
For the 1–3 functions most central to the change, build a bounded **PDG
context slice**. Read `references/pdg-slice.md` and follow it — it owns the
tool calls, inclusion criteria, depth bounds, slice schema, the security and
performance modes, and the no-PDG-layer fallback.
## Phase 4 — Targeted source verification
@ -123,6 +124,10 @@ reads (exact line ranges, not whole files unless genuinely required):
state mutations, error paths, nearby comments that change behavior.
- Read the tests GitNexus associated with the primary symbols; never claim a
test exists without having located it.
- Verify the build/test commands the plan will name actually exist
(package.json scripts / CI workflows), and prefer the script form that
carries its prerequisites (pre-hooks) over invoking underlying binaries
directly.
- Check repo conventions that constrain the change (AGENTS.md, GUARDRAILS.md,
lint/build config) — only the parts the change touches.
- Mark each ledger symbol `source_verified: true` as you go. **A symbol that
@ -138,38 +143,35 @@ and executable behavior → compiler/build/lint output → GitNexus graph and PD
## Phase 5 — Compose the plan
1. Read `references/plan-template.md` and fill all 13 sections from the
ledger. Distinguish **confirmed facts / evidence-backed inferences /
assumptions / open questions** throughout.
ledger, using its claim-tagging convention to distinguish confirmed facts,
evidence-backed inferences, assumptions, and open questions.
2. Build the implementation context pack per `references/context-pack.md`
(this is section 11 of the plan).
3. Write the document to `docs/plans/YYYY-MM-DD-ce-plan-<slug>.md` (create the
directory if missing; kebab-case slug, 3–5 words). Repo-relative paths
everywhere.
3. Write the document to `docs/plans/YYYY-MM-DD-ce-plan-<slug>.md` under the
root of the repo being planned (the Phase 1 target repo, not necessarily
the cwd). Create the directory if missing; kebab-case slug, 3–5 words.
The `out:<path>` knob overrides the destination (use it for read-only
checkouts). Repo-relative paths inside the document.
4. Present in chat: objective, proposed-changes summary, implementation
sequence, top risks, open questions, and the plan file path. Do not paste
the whole document into chat.
## Context ledger
Maintain the ledger from Phase 0 onward — it is the skill's working memory and
its token budget enforcement. Schema and reread rules: `references/context-ledger.md`.
The final ledger feeds the plan; it is not itself published.
## Configuration
Defaults; override inline with `key:value` tokens before the task text (the
repo has no skill-config file mechanism — invocation args are the mechanism):
Baseline defaults — the Phase 0 category posture overrides them, and inline
`key:value` tokens before the task text override both (the repo has no
skill-config file mechanism; invocation args are the mechanism):
| Knob | Default | Meaning |
| --- | --- | --- |
| `depth` | by category | `narrow` / `default` / `deep` posture |
| `call_depth` | 2 | Caller/callee expansion in `context`/`trace` reasoning |
| `depth` | by category | `narrow` = `impact_depth` 1, PDG only if one function is clearly central; `default` = this table; `deep` = `impact_depth` 3 + clusters/processes read |
| `impact_depth` | 2 | `maxDepth` for `impact` |
| `pdg_data_depth` | 2 | Data-dependence hops in the PDG slice |
| `pdg_control_depth` | 2 | Control-dependence hops in the PDG slice |
| `max_primary_symbols` | 5 | Ledger budget |
| `max_related_symbols` | 20 | Ledger budget |
| `max_primary_symbols` | 5 | Ledger budget (active symbols; discards don't count) |
| `max_related_symbols` | 20 | Ledger budget (active symbols; discards don't count) |
| `max_snippet_lines` | 30 | Longest source excerpt quoted in the plan |
| `out` | `docs/plans/` in target repo | Plan document destination |
## Fallback mode (GitNexus or PDG unavailable)
@ -180,12 +182,3 @@ repo has no skill-config file mechanism — invocation args are the mechanism):
as graph-derived, and never fabricate statement-level edges.
4. Recommend `node .gitnexus/run.cjs analyze` (add `--pdg` for the PDG layers)
when it would materially raise confidence.
## Never
- Implement the feature, edit production code, or run mutating commands.
- Dump unfiltered graph/PDG output or full files into the plan.
- Repeat a search or reread an unchanged range already in the ledger.
- Expand scope into unrelated refactoring.
- Treat comments as stronger evidence than executable code.
- Continue exploring after the ledger answers all open planning questions.

View file

@ -1,10 +1,11 @@
# Context ledger
The ledger is ce-plan's working memory. It exists to make repeated
investigation impossible-by-discipline: **before every GitNexus call and every
file read, check it.** Keep it as structured notes in your working context (or
a scratchpad file for very long sessions); it is never published verbatim —
the plan and context pack are distilled from it.
investigation impossible-by-discipline: **before every GitNexus call and
every repo file read, check it.** Keep it as structured notes in your working
context (or a scratchpad file *outside the repo* for very long sessions); it
is never published verbatim — the plan and context pack are distilled from
it. This skill's own reference files are exempt from ledger bookkeeping.
## Schema
@ -16,9 +17,14 @@ context_ledger:
category: "" # Phase 0 classification
acceptance_criteria: []
verified_at_commit: "" # target repo HEAD, recorded once in Phase 1;
# every line citation in the plan pins to it
established_facts: [] # each with its evidence source
symbols: # budgeted: max_primary_symbols / max_related_symbols
symbols: # budgets count active (primary/related) only;
# discards are free — but on budget overflow,
# discard something before promoting
- name: ""
kind: ""
file: ""
@ -29,12 +35,12 @@ context_ledger:
- file: ""
ranges: [] # e.g. ["120-188"]
purpose: ""
content_hash_or_version: "" # e.g. git blob hash or "HEAD@<sha>"
gitnexus_queries:
- query: "" # tool + args
purpose: "" # the planning question it answers
conclusion: "" # one line; details stay in working memory
key_output: "" # one-line raw quote when the plan leans on this result
pdg_slices:
- symbol: ""
@ -50,11 +56,17 @@ context_ledger:
Do **not** repeat a query or reread a source range unless one of:
- the previous result was incomplete for a *new* question;
- the source changed (hash/version mismatch, or an edit is known to have
happened);
- the previous result was incomplete for the question at hand;
- the source is known to have changed (an edit happened);
- validation exposed a contradiction between graph and source.
**Allowed repeats** (deliberate escalations, not violations):
- `summaryOnly: true` → full drill-down on the same `impact` target;
- an `ambiguous` result retried once with `kind` / `file_path` / uid narrowing;
- the same tool re-run with a changed parameter that answers a *new* planning
question (e.g. `pdg_query` `controls` then `flows` on one function).
When a repeat is justified, note in the ledger *why* the earlier entry was
insufficient. A ledger full of near-duplicate queries is the failure signal —
stop and plan with what is established.

View file

@ -1,9 +1,9 @@
# Implementation context pack
Section 11 of the plan. The stable, machine-readable contract a follow-up
implementation agent (`ce-implement` or any executor) consumes to start work
**without repeating the investigation**. Distilled from the ledger; every
entry traceable to verified evidence.
implementation agent (a future `ce-implement`, or any executor) consumes to
start work **without repeating the investigation**. Distilled from the
ledger; every entry traceable to verified evidence.
## Schema
@ -44,10 +44,12 @@ implementation_context:
- file: "" # existing file to update, or new path to create
scenarios: [] # input → action → expected outcome
verification_commands: [] # real commands from repo scripts/CI, verified to exist
verification_commands: [] # real commands verified to exist AND be runnable —
# prefer npm/CI scripts that carry their pre-hooks
risks: []
assumptions: [] # verbatim from plan §12
assumptions: [] # faithful condensation of plan §12 assumptions
open_questions: [] # faithful condensation of plan §12 open questions
avoid:
- "Do not repeat full repository discovery"

View file

@ -9,18 +9,24 @@ Goal: a compact slice the planning LLM can hold, never a graph dump.
| --- | --- |
| Under what condition does X run? Guards? | `pdg_query {mode: "controls", target}` |
| Where does variable Y flow inside the function? | `pdg_query {mode: "flows", target, variable}` |
| What depends on the statement at line N? | `impact {mode: "pdg", target, line: N}` |
| What depends on the statement at line N? | `impact {mode: "pdg", target, direction: "upstream", line: N}` |
| Source→sink taint paths (security mode) | `explain {target}` |
Contract caveats that shape interpretation:
- CDG branch sense is `'T'`/`'F'` in the edge's `reason`; a guard's sense
depends on its predicate (`if (!ok) return;` rides `'T'`) — never filter
guards by a fixed label. Early return/throw edges carry `guard: true`.
- `impact` requires `direction` in every mode, `mode: "pdg"` included —
`"upstream"` for "what depends on this statement", `"downstream"` for what
it depends on. Omitting it fails schema validation.
- CDG branch sense is `'T'`/`'F'` in the result's `label` field; a guard's
sense depends on its predicate (`if (!ok) return;` rides `'T'`) — never
filter guards by a fixed label. Early return/throw edges carry `guard:
true`. (The raw edge stores the sense in `reason`, visible only via
`cypher`.)
- `pdg_query` is intra-procedural and always anchored. Cross-function flow is
taint's domain (`explain`) or `impact {mode:"pdg"}`'s inter-procedural reach.
- Every `switch` case arm is `'T'` (per-case conditions not distinguished).
- No `--pdg` layer → the tools return a "no PDG layer" note, not an error.
The note is repo-wide: one probe settles it — do not re-probe per function.
Record it, skip the slice, recommend `analyze --pdg`. Do not reconstruct
edges from source by hand.
@ -42,7 +48,11 @@ A statement enters the slice only if it is at least one of:
Everything else is cut. If the slice exceeds ~15 statements per function,
tighten relevance rather than raising depth.
## Slice representation (goes into the ledger; summarized in plan §5)
## Slice representation
Working-memory material: keep the full slice in working context while
planning, summarize it into the ledger's one-line `pdg_slices` entries, and
distill it into plan §5.
```yaml
pdg_context:
@ -79,9 +89,11 @@ Additionally identify and record: untrusted inputs, validation points,
sanitisation points, authn/authz checks, privilege boundaries, sensitive data,
persistence operations, network calls, dangerous sinks, and error paths that
bypass validation. Run `explain {target}` for persisted source→sink taint
paths and include the hop paths for findings relevant to the task. Absence of
a taint finding is **not** proof of safety (intra-procedural only) — say so
when it matters.
paths (intra-procedural TAINTED edges and cross-function TAINT_PATH flows)
and include the hop paths for findings relevant to the task. Absence of a
taint finding is **not** proof of safety — closure/callback flows,
property/field flows, and implicit flows are not modeled, and guard-style
sanitizers may be missed — say so when it matters.
## Performance mode (task category: performance)

View file

@ -2,12 +2,20 @@
Fill every section. If a section is genuinely empty for this task (e.g. no PDG
layer indexed), keep the heading and state why in one line — never silently
drop it. Repo-relative paths everywhere. Distinguish confirmed facts,
evidence-backed inferences, assumptions, and open questions throughout.
drop it. Repo-relative paths for all repo artifacts.
**Claim tagging.** Tag every load-bearing claim with its evidence class:
`[verified]` (source-read at the pinned commit), `[graph]` (GitNexus/PDG
output, not source-confirmed), `[inferred]` (evidence-backed reasoning),
`[assumed]` (unverified — must also appear in §12). Untagged prose is
narrative, not evidence.
```markdown
# Compound Engineering Plan
> Task: <one line>
> Evidence verified at commit <HEAD sha>; GitNexus index <fresh | N commits behind | not used>.
## 1. Objective
A concise description of the requested outcome.
@ -112,11 +120,16 @@ Composition notes:
- §2/§5 quote source excerpts at most `max_snippet_lines` (30) lines each, and
only when the excerpt carries the argument.
- §4 findings each name the tool call they came from (traceable to the
ledger); stale-index or fallback-mode findings are labelled as such.
- §4 findings each name the tool call they came from (tool + key args), plus a
one-line quote of the result when the plan leans on it — that is what makes
a tool claim auditable later. Stale-index or fallback-mode findings are
labelled as such.
- §6 changes may only name symbols the ledger marks `source_verified`.
- §7 steps are ordered by dependency and independently actionable — an
executor can stop after any step with the tree still coherent.
- §8 names real, located test files for updates; new tests get concrete
scenario lists (input → action → expected outcome).
- §9 must account for every d=1 symbol the impact pass reported.
scenario lists (input → action → expected outcome). Verification commands
must exist AND be runnable: prefer the npm/CI script form that carries its
prerequisites (pre-hooks, builds) over invoking underlying binaries directly.
- §9 must account for every direct (depth-1) dependent the impact pass
reported.

View file

@ -1,7 +1,7 @@
<!-- version: 1.7.0 -->
<!-- Last updated: 2026-04-23 -->
<!-- version: 1.9.0 -->
<!-- Last updated: 2026-07-11 -->
Last reviewed: 2026-04-23
Last reviewed: 2026-07-11
**Project:** GitNexus · **Environment:** dev · **Maintainer:** repository maintainers (see GitHub)
@ -68,6 +68,7 @@ The skill is planning-only — it never edits code.
| Date | Version | Change |
|------|---------|--------|
| 2026-07-11 | 1.9.0 | Added Engineering planning (`/ce-plan`) section; registered the `ce-plan` skill (`.claude/skills/ce-plan/`). |
| 2026-05-22 | 1.8.0 | Kotlin added to `MIGRATED_LANGUAGES` (registry-primary call resolution by default). Closes #1756 (companion-vs-instance dispatch) and #1757 (lambda scopes); refs #1746. RFC §6.4 corpus criterion waived (corpus-mode wiring is #927-scope); fixture criterion met. |
| 2026-04-23 | 1.7.0 | TypeScript added to `MIGRATED_LANGUAGES` (registry-primary call resolution by default). |
| 2026-04-20 | 1.6.0 | Added scope-resolution pipeline pointer (RFC #909 Ring 3); Python migrated to registry-primary. |

View file

@ -1,10 +1,10 @@
<!-- version: 1.3.0 -->
<!-- version: 1.4.0 -->
<!--
Metadata: version, last reviewed, scope, model policy, reference docs, changelog.
Last updated: 2026-03-22
Last updated: 2026-07-11
-->
Last reviewed: 2026-04-13
Last reviewed: 2026-07-11
**Project:** GitNexus · **Environment:** dev · **Maintainer:** repository maintainers (see GitHub)
@ -43,6 +43,7 @@ If always-on instructions grow, load deep conventions via conditional reads (e.g
| Date | Version | Change |
|------|---------|--------|
| 2026-07-11 | 1.4.0 | Added `/ce-plan` pointer to Reference Documentation. |
| 2026-04-13 | 1.3.0 | Updated GitNexus index stats after DAG refactor. |
| 2026-03-24 | 1.2.0 | Removed duplicated gitnexus:start block and scope table; replaced with pointers to AGENTS.md. |
| 2026-03-23 | 1.1.0 | Updated agent instructions to match AGENTS.md. |