claude-skills/audit/engineering-agentic-2026-07/RUBRIC.md
Claude 0a5d18ceba
feat(engineering): agent-harness skill + agentic-readiness audit of both engineering domains
Deep-audit both engineering folders (engineering/ + engineering-team/) against the
June 2026 baseline and score every skill on a new 6-dimension agentic-readiness rubric
(goal intake, decomposition, deterministic execution, verification, loop discipline,
close-out). Combined: 26 HARNESS-READY, 39 LOOP-CAPABLE, 43 TOOL-ONLY, 7 PROSE-ONLY.
Headline finding: loop discipline (AR5) is the repo-wide gap.

Ship engineering/agent-harness — the thin unifying layer that turns any of the repo's
18 domains into a bounded, self-verifying agent loop:
- harness_manifest_builder.py: scan a domain -> manifest.v1 (skills, tools, checks, signals)
- goal_compiler.py: goal + manifest -> plan.v1; refuses vague goals (exit 3) / no-match (4)
- loop_controller.py: init/next/record/verify/close state machine; runs checks itself via
  subprocess (no verification theater), caps attempts+iterations with escalation, refuses
  to close while any task is unverified; atomic state writes
- 18 committed per-domain manifests, JSON schema, harness-runner agent, /cs:harness command,
  3 references citing the 2024-2026 harness canon
- reuses agenthub / autoresearch locked-evaluator / tc-tracker / loop-library primitives

Audit record under audit/engineering-agentic-2026-07/ (master + 2 domain reports +
improvement-fields rollup + research digest + rubric).

Counters: 82->83 plugins, 354->355 skills, 593->596 tools, 722->725 refs (derive_counters
--check passes). All CI gates green: plugin.json, smoke --help/--sample, JSON output,
path linter, dual-publish, counters.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L4JerbGv6vqitUMhqHPA9g
2026-07-03 06:01:43 +00:00

2.5 KiB
Raw Permalink Blame History

Agentic-Readiness Rubric (AR v1)

Audit date: 2026-07-03 · Branch: claude/engineering-audit-agentic-loops-hv9x9m

The June 2026 audit (../newgen-2026-06/RUBRIC.md) asked "does this skill earn its context window?" This follow-up asks the next question: can an agent pick this skill up with a goal and drive it to a verified close? — the gather-context → take-action → verify-work → repeat loop the 2024–2026 harness canon converged on (see research-digest.md).

The six dimensions (0–2 each, total 0–12)

# Dimension 0 1 2
AR1 Goal intake Accepts any input silently Asks context questions Forcing questions / intake tool / refuses vague input (exit-code gate)
AR2 Task decomposition No plan step Prose phases Explicit planning step or tool whose output the workflow consumes
AR3 Deterministic execution No wired tools Tools named, CLIs incomplete Exact runnable CLIs; output consumed by a named next step
AR4 Verification None Checklist prose Machine-checkable gate (exit codes, JSON assertions) the workflow REQUIRES before proceeding
AR5 Loop discipline No retry/stop rules "Re-run until clean" without a cap Iteration caps, stop conditions, escalation thresholds
AR6 Close-out Work just ends Informal done statement Definition of done + state persistence or handoff artifact

Classes

Class Criteria Meaning
HARNESS-READY total ≥ 9 AND AR4 ≥ 1 AND AR5 ≥ 1 An agent can run this skill inside a bounded loop today
LOOP-CAPABLE total 6–8 (or ≥9 failing an AR4/AR5 gate) One or two targeted additions from harness-ready
TOOL-ONLY total 3–5 Good tools, no loop spine
PROSE-ONLY total 0–2 Knowledge dump; needs structural rebuild

The AR4/AR5 gate is deliberate: a skill with perfect intake and tools but no verification gate or stop condition is more dangerous in an autonomous loop, not less — it runs confidently and forever.

Executable enforcement

The rubric is now mechanized: engineering/agent-harness/skills/agent-harness/scripts/harness_manifest_builder.py records per-skill agentic_signals (static evidence for AR1/AR4/AR5/AR6) in every domain manifest, and loop_controller.py enforces AR4/AR5/AR6 at run time regardless of the skill's own discipline. Improvement PRs should move skills up this ladder; the manifests make regressions diffable.