Two reviewer nits batched: derive_counters.py's module docstring and --check
help still described the pre-#989 three-source coverage (flagged on #989);
check_model_freshness.py's EXCLUDED_DIRS did double duty as a directory AND
filename exclusion set, which the name hid (flagged on #985 and #988's
reviews) — renamed with a comment stating both roles. No behavior change;
both gates re-verified passing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qgc6RYXWJPr5oW9DHU7zR4
Review follow-up on #989: shortDescription/longDescription carry their own
counts and were just trued — include them in the gated source text so
standardized-phrasing claims in them are checked (non-matching prose is
simply not read). Verified: planting 997 in shortDescription fails the
gate; restored passes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qgc6RYXWJPr5oW9DHU7zR4
Adopts the two verified findings from PR #940 (credit: @benrfairless):
- .codex-plugin/plugin.json still said v2.2.0 / 223 skills / 23 agents /
298 tools / 9 domains — roughly nine releases behind, and it is the
manifest Codex users see. Version, description, and the interface
short/long descriptions are trued to the v2.12.0 counters (380 skills /
20 domains / 706 tools / 823 refs / 114 agents / 138 commands / 96
plugins), with the top-level description written in the standardized
claim phrasing so the gate can read it.
- mkdocs.yml's site_description was content-correct after v2.12.0 but
ungated and phrased invisibly to extract_claims ('agent skills',
'installable plugins') — reworded to the standardized phrasing.
- derive_counters.py run_check() now reads both as claim sources
(mkdocs.yml restricted to the site_description line since its !!python
tags reject safe_load). Verified: planting 999/998 in the two sites
fails the gate naming both; restored values pass.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qgc6RYXWJPr5oW9DHU7zR4
Caught by review on the v2.12.0 promotion PR #985: README's Skills Overview
heading still said 370 and CLAUDE.md's footer Status line said 379 while the
banner/badges/scope line say the derived 380. Both wordings ('370 skills
across', '379 skills deployed across') were invisible to derive_counters.py's
claim patterns, which is why they could drift — reworded both into the
standardized '<N> production-ready skills across <D> domains' phrasing, made
extract_claims() validate every occurrence of a claim pattern instead of only
the first, and run_check() now reads CLAUDE.md's Status footer line alongside
Current Scope. Verified: planting 999/998 in the two lines fails the gate
naming both; restored values pass.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qgc6RYXWJPr5oW9DHU7zR4
PR #984's counter true-up bumped README badges and CHANGELOG to the derived
114 agents / 138 commands but left the CLAUDE.md Current Scope line and
marketplace.json metadata.description at the stale 111/131 (caught by review
on #984). derive_counters.py --check passed because CLAIM_PATTERNS had no
agents/commands patterns — added both (agents anchored on the "(cs-" suffix
so prose like "9 more coding agents" can't false-match), verified the new
gate fails on the pre-fix docs and passes post-fix.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qgc6RYXWJPr5oW9DHU7zR4
Per PR #891 review: check_readme_badges silently skipped a badge whose regex
found no match, so a renamed or removed shield would quietly drop out of the
gate — the same silent-drift class this gate exists to catch. Now a missing
badge appends a mismatch (mirrors run_check's "no recognizable counter claims
found" precedent), so it fails loudly. Verified: renaming a badge trips exit 1;
the current README still passes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L4JerbGv6vqitUMhqHPA9g
Per PR #891 review: `mismatches: list = []` was an inconsistent drive-by
annotation vs the un-annotated locals elsewhere in the file. Revert to keep
the diff minimal and the style consistent. No behavior change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L4JerbGv6vqitUMhqHPA9g
Follow-up to the merged agent-harness PR (#890), applying the automated review nits:
- loop_controller.py: drop unused `import shlex`; simplify cmd_record's exit-code
expression to the clearer form already used in cmd_verify (behavior-equivalent)
- SKILL.md + references/verification_discipline.md: document that plan/state files are
a trust boundary (verify shell-executes their cmd strings) — run the harness only on
files produced by goal_compiler, never untrusted input
- README.md: bump Agents 96->97 and Commands 102->103 badges (drift the previous PR
missed because derive_counters didn't validate these badges)
- scripts/derive_counters.py: add check_readme_badges — validates the Skills/Agents/
Commands shields against derived counts, closing the CI blind spot that let the badge
drift ship. Verified it fails (exit 1) on drift and passes when correct.
All gates green: plugin.json (83 OK), smoke --help/--sample (600 pass), JSON output
(0 fail), path linter (0 findings), derive_counters --check (pass).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L4JerbGv6vqitUMhqHPA9g
Two housekeeping items surfaced during the contributor-PR hardening session.
1. CHANGELOG backfill — add [Unreleased] entries for five skills that merged
without their own changelog blocks (roast #865, named-persona-adversarial-review
#867, agent-decision-receipts #868/#869, zero-hallucination-coder #870,
deep-research #872). Earlier merges updated headline counters but not this log.
2. Per-domain counter validation — scripts/derive_counters.py --check now also
validates the README "Skills Overview" per-domain table: each domain row's
count must equal the SKILL.md count in its linked folder, and every on-disk
domain must have a row. Previously --check only validated headline aggregates,
so per-domain rows drifted silently. Verified: passes on the fixed state, fails
on a wrong count, fails on a missing row, and parses exactly the 18 real domain
rows (bold-first-cell install/skills-vs-agents tables are not false-flagged).
Trued up the README table to make the new check pass: fixed six stale row
counts (engineering-team 51->52, engineering 78->80, marketing 47->48,
productivity 6->7, ra-qm-team 18->19, c-level 66->68), added the missing
markdown-html row (5), and named the newly-merged skills in their domain
descriptions. Per-domain rows now sum to the 354 headline.
Headline aggregates unchanged (354 skills / 722 refs / 82 plugins / 18 domains).
- derive_counters.py: python_tools condition simplified to the equivalent
parts[0] != 'scripts' (reviewer M1); dead root_scripts variable removed;
--check still passes with identical values
- fda-consultant-specialist quick-start: 820.30 example annotated as a legacy
checklist key mapping to ISO 13485 §7.3 (reviewer m3 — note: switching the
example to '--section 7.3' as suggested would break; the checker's CLI keys
are intentionally the legacy 820.x checklist indices, documented in --help)
- CLAUDE.md: audit/ directory documented as an intentional public audit
record, distinct from the gitignored AUDIT_REPORT.md (reviewer m2)
Reviewer m1 (agents/CLAUDE.md 'engineering-team/' link) is a false positive:
agents/engineering-team/ exists as an agents subfolder containing exactly the
two linked files; check_paths.py confirms 0 unresolvable references.
https://claude.ai/code/session_019AJddAL1NADWMXsy1qNPQF