mirror of
https://github.com/alirezarezvani/claude-skills.git
synced 2026-10-09 03:17:54 +00:00
- CHANGELOG.md gains the [2.12.0] entry (first tagged release since v2.9.0):
consolidates the previously documented but untagged v2.10.0-v2.11.2 work,
all post-2.11.2 merges, and the full 17-issue triage sweep; the ten stacked
[Unreleased] sections are demoted into the 2.12.0 body so the Release
workflow tags and publishes the whole span. Verified parseable with
scripts/extract_release_notes.py (version 2.12.0, 554-line body).
- Version markers bumped to 2.12.0: marketplace.json metadata,
CLAUDE.md current-version header + footer.
- Counters trued to the derived values (380 skills / 96 plugins / 20 domains /
706 tools / 823 refs / 114 agents / 138 commands) in README badges + prose,
CLAUDE.md, marketplace.json, and the long-stale mkdocs.yml/docs/index.md
site description (was still claiming 345/78/17).
- Docs site regenerated via scripts/generate-docs.py (568 generated pages;
new pages for the recently merged plugins); codex/gemini mirrors resynced;
mkdocs build verified locally with the same plugin set static.yml uses
(670 HTML pages, no errors).
- Fix: the three hivemind worker personas (assets/agents/{coder,scout,tester}.md,
merged via #979 while Actions was not triggering) lacked the frontmatter
`name:` field and hard-failed the blocking G10 gate — named
hive-coder/hive-scout/hive-tester; 645 files now scan with 0 errors.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qgc6RYXWJPr5oW9DHU7zR4
4.1 KiB
4.1 KiB
| title | description |
|---|---|
| Phase 3 — Grade → Iterate (the bounded loop) — Agent Skill for Claude Managed Agents | Phase 3 of building a Claude Managed Agent — the bounded grade→iterate loop. Define a CMA outcome (a required markdown rubric graded by an isolated. Agent skill for Claude Code, Codex CLI, Gemini CLI, OpenClaw. |
Phase 3 — Grade → Iterate (the bounded loop)
:material-rocket-launch-outline: Agent Launcher
:material-identifier: `grade-iterate`
:material-github: Source
Install:
claude /plugin install agent-launcher-skills
This is the plugin's loop: CMA's outcome primitive self-grades the agent's
work in an isolated context and feeds failing verdicts back for the next attempt.
It is always bounded by max_iterations (1..20) — never "improve forever".
See [references/loops-and-workflows.md](https://github.com/alirezarezvani/claude-skills/tree/main/agent-launcher/references/loops-and-workflows.md)
and the outcome section of
[references/cma-primitives.md](https://github.com/alirezarezvani/claude-skills/tree/main/agent-launcher/references/cma-primitives.md).
Workflow
- Define the outcome.
The rubric is required;python3 scripts/outcome_builder.py \ --sheet ./my-agent/build-sheet.json --max-iterations 5 \ --out ./my-agent/payloads/outcome.jsonmax_iterationsis clamped to 1..20. Send the payload as auser.define_outcomeevent (append to the running session). - Read every verdict first.
Tables the rubric outcome and recommends: SHIP (python3 scripts/verdict_reader.py --result ./my-agent/last-verdict.jsonsatisfied), SHARPEN then re-run (needs_revision), ESCALATE (max_iterations_reached/failed), RESUME (interrupted). With ≤1 iteration left it flips to "make the single highest-value fix or escalate now". - Loop invariant. Each iteration must move ≥1 rubric line fail→pass, or the run halts at the cap and escalates. Don't burn the budget on cosmetic edits.
- Once a version passes, run held-back eval.
Held-back cases (never seen during iteration) run in parallel, capped at the 25-thread CMA ceiling, each graded against the same rubric.python3 scripts/eval_scaffold.py \ --sheet ./my-agent/build-sheet.json --out ./my-agent/eval.json --concurrency 5 - Decide. SHIP as v0, or promote to a scheduled deployment (Phase 4). Record
the verdict on the goal:
goal_state.py set --phase run-without-you.
Hard rules
- Bounded, always. No outcome without a
max_iterationscap. - Read the verdict before acting. The grader's explanation drives the next move.
- Held-back cases are held back. Never grade generalization on cases the agent already iterated against.
Forcing-question library (recommend + cite)
- "What are the 3–5 rubric lines?" Recommend: grounded, checkable criteria. Cite: cma-primitives.md (rubric required).
- "How many iterations before you'd rather look yourself?" Recommend: 3–5. Cite: loops-and-workflows.md (bounded loop).
- "On a fail, sharpen the prompt or the tools?" Recommend: whichever rubric line failed points to. Cite: verdict_reader next-move table.
- "Which cases did the agent NOT see?" Recommend: hold back ≥3 for generalization. Cite: this SKILL (held-back eval).
Tools
scripts/outcome_builder.py— user.define_outcome payload (rubric required, cap 1..20).scripts/verdict_reader.py— grader result → next move.scripts/eval_scaffold.py— held-back cases + parallel run plan (≤25 threads).