mirror of
https://github.com/alirezarezvani/claude-skills.git
synced 2026-10-09 03:17:54 +00:00
Deep-audit both engineering folders (engineering/ + engineering-team/) against the June 2026 baseline and score every skill on a new 6-dimension agentic-readiness rubric (goal intake, decomposition, deterministic execution, verification, loop discipline, close-out). Combined: 26 HARNESS-READY, 39 LOOP-CAPABLE, 43 TOOL-ONLY, 7 PROSE-ONLY. Headline finding: loop discipline (AR5) is the repo-wide gap. Ship engineering/agent-harness — the thin unifying layer that turns any of the repo's 18 domains into a bounded, self-verifying agent loop: - harness_manifest_builder.py: scan a domain -> manifest.v1 (skills, tools, checks, signals) - goal_compiler.py: goal + manifest -> plan.v1; refuses vague goals (exit 3) / no-match (4) - loop_controller.py: init/next/record/verify/close state machine; runs checks itself via subprocess (no verification theater), caps attempts+iterations with escalation, refuses to close while any task is unverified; atomic state writes - 18 committed per-domain manifests, JSON schema, harness-runner agent, /cs:harness command, 3 references citing the 2024-2026 harness canon - reuses agenthub / autoresearch locked-evaluator / tc-tracker / loop-library primitives Audit record under audit/engineering-agentic-2026-07/ (master + 2 domain reports + improvement-fields rollup + research digest + rubric). Counters: 82->83 plugins, 354->355 skills, 593->596 tools, 722->725 refs (derive_counters --check passes). All CI gates green: plugin.json, smoke --help/--sample, JSON output, path linter, dual-publish, counters. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L4JerbGv6vqitUMhqHPA9g
33 lines
1.7 KiB
Markdown
33 lines
1.7 KiB
Markdown
# agent-harness
|
|
|
|
Turn any domain folder of this repo into a **bounded agentic loop**: pick up a goal,
|
|
compile it into tasks with machine-run verification, execute, verify, retry with caps,
|
|
escalate to a human when budgets exhaust, and close only when everything is verified.
|
|
|
|
```
|
|
GOAL → goal_compiler → PLAN → loop_controller: [execute → verify]* → CLOSE
|
|
↑ retry ≤ caps, changed approach
|
|
└ ESCALATE — never fake success
|
|
```
|
|
|
|
## What ships
|
|
|
|
| Piece | Purpose |
|
|
|---|---|
|
|
| `scripts/harness_manifest_builder.py` | Scan a domain folder → `manifest.v1` JSON (skills, tools, checks, agentic signals) |
|
|
| `scripts/goal_compiler.py` | Goal + manifest → `plan.v1` task plan; refuses vague goals (exit 3, forcing questions) |
|
|
| `scripts/loop_controller.py` | `init/next/record/verify/close/status` state machine; controller runs checks itself |
|
|
| `assets/harnesses/*.json` | 18 committed per-domain manifests (regenerable, diff-stable) |
|
|
| `assets/harness_manifest.schema.json` | Manifest schema |
|
|
| `references/` | Agentic-loop canon, verification discipline, domain-harness design (cited) |
|
|
| `agents/harness-runner.md` | Stateless one-task-per-invocation executor |
|
|
| `commands/cs-harness.md` | `/cs:harness <domain> <goal>` end-to-end driver |
|
|
|
|
All tools are stdlib-only, pass `--help` and `--sample`, and emit JSON.
|
|
|
|
## Design lineage
|
|
|
|
Anthropic's long-running-agents harness (feature-list + stateless shifts), verifier's law,
|
|
SWE-agent's environment-feedback lesson, Ralph-loop fresh-context iteration, Cognition's
|
|
serialize-writers rule, and this repo's own tc-tracker / autoresearch locked-evaluator /
|
|
loop-library stop-state primitives. See `skills/agent-harness/references/`.
|