Deep-audit both engineering folders (engineering/ + engineering-team/) against the June 2026 baseline and score every skill on a new 6-dimension agentic-readiness rubric (goal intake, decomposition, deterministic execution, verification, loop discipline, close-out). Combined: 26 HARNESS-READY, 39 LOOP-CAPABLE, 43 TOOL-ONLY, 7 PROSE-ONLY. Headline finding: loop discipline (AR5) is the repo-wide gap. Ship engineering/agent-harness — the thin unifying layer that turns any of the repo's 18 domains into a bounded, self-verifying agent loop: - harness_manifest_builder.py: scan a domain -> manifest.v1 (skills, tools, checks, signals) - goal_compiler.py: goal + manifest -> plan.v1; refuses vague goals (exit 3) / no-match (4) - loop_controller.py: init/next/record/verify/close state machine; runs checks itself via subprocess (no verification theater), caps attempts+iterations with escalation, refuses to close while any task is unverified; atomic state writes - 18 committed per-domain manifests, JSON schema, harness-runner agent, /cs:harness command, 3 references citing the 2024-2026 harness canon - reuses agenthub / autoresearch locked-evaluator / tc-tracker / loop-library primitives Audit record under audit/engineering-agentic-2026-07/ (master + 2 domain reports + improvement-fields rollup + research digest + rubric). Counters: 82->83 plugins, 354->355 skills, 593->596 tools, 722->725 refs (derive_counters --check passes). All CI gates green: plugin.json, smoke --help/--sample, JSON output, path linter, dual-publish, counters. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L4JerbGv6vqitUMhqHPA9g
2.5 KiB
Agentic-Readiness Rubric (AR v1)
Audit date: 2026-07-03 · Branch: claude/engineering-audit-agentic-loops-hv9x9m
The June 2026 audit (../newgen-2026-06/RUBRIC.md) asked "does this skill earn its context window?" This follow-up asks the next question: can an agent pick this skill up with a goal and drive it to a verified close? — the gather-context → take-action → verify-work → repeat loop the 2024–2026 harness canon converged on (see research-digest.md).
The six dimensions (0–2 each, total 0–12)
| # | Dimension | 0 | 1 | 2 |
|---|---|---|---|---|
| AR1 | Goal intake | Accepts any input silently | Asks context questions | Forcing questions / intake tool / refuses vague input (exit-code gate) |
| AR2 | Task decomposition | No plan step | Prose phases | Explicit planning step or tool whose output the workflow consumes |
| AR3 | Deterministic execution | No wired tools | Tools named, CLIs incomplete | Exact runnable CLIs; output consumed by a named next step |
| AR4 | Verification | None | Checklist prose | Machine-checkable gate (exit codes, JSON assertions) the workflow REQUIRES before proceeding |
| AR5 | Loop discipline | No retry/stop rules | "Re-run until clean" without a cap | Iteration caps, stop conditions, escalation thresholds |
| AR6 | Close-out | Work just ends | Informal done statement | Definition of done + state persistence or handoff artifact |
Classes
| Class | Criteria | Meaning |
|---|---|---|
| HARNESS-READY | total ≥ 9 AND AR4 ≥ 1 AND AR5 ≥ 1 | An agent can run this skill inside a bounded loop today |
| LOOP-CAPABLE | total 6–8 (or ≥9 failing an AR4/AR5 gate) | One or two targeted additions from harness-ready |
| TOOL-ONLY | total 3–5 | Good tools, no loop spine |
| PROSE-ONLY | total 0–2 | Knowledge dump; needs structural rebuild |
The AR4/AR5 gate is deliberate: a skill with perfect intake and tools but no verification gate or stop condition is more dangerous in an autonomous loop, not less — it runs confidently and forever.
Executable enforcement
The rubric is now mechanized: engineering/agent-harness/skills/agent-harness/scripts/harness_manifest_builder.py
records per-skill agentic_signals (static evidence for AR1/AR4/AR5/AR6) in every domain
manifest, and loop_controller.py enforces AR4/AR5/AR6 at run time regardless of the
skill's own discipline. Improvement PRs should move skills up this ladder; the manifests
make regressions diffable.