Round-2 sweep after re-auditing all 15 reported issues against the merged dev:
- #885 generalized: the original fix only renamed self-improving-agent's
status/review, but three more plugins shipped skills whose bare names
shadow Claude Code built-ins. Renamed with the same convention:
playwright-pro init/review -> pw-init/pw-review, agenthub init/status ->
hub-init/hub-status, autoresearch-agent status/resume -> ar-status/
ar-resume. All command references (/pw: /hub: /ar:), docs, audit records,
harness manifests, and mirror trees/indexes updated; the flat mirror
namespace no longer collides on 'status'. New scripts/check_skill_names.py
gate (wired into ci-quality-gate.yml as blocking) fails CI on any future
bare reserved name; rule added to SKILL-AUTHORING-STANDARD.md.
- #969 follow-through: five more scripts print box-drawing characters that
cannot exist in cp1252 (api_scorecard, api_linter,
breaking_change_detector, humanizer_scorer, content_scorer) — same
guarded UTF-8 reconfigure applied; all smoke-tested under a forced
legacy encoding.
Verified: check_skill_names (incl. negative test), check_plugin_json,
check_paths, derive_counters, check_dual_publish, smoke_scripts (634/634),
0 broken mirror symlinks.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qgc6RYXWJPr5oW9DHU7zR4
Clears every reference the new G7 lint flags, then makes it blocking so the
class cannot drift back. audit/engineering-agentic-2026-07 marked the
senior-ml-engineer half of this STILL-OPEN.
Deleted rather than updated:
- agent-designer/agent_evaluator.py's _define_cost_benchmarks() held
per-token prices for gpt-4, gpt-3.5-turbo and claude-3 at 2024 rates. The
result was assigned to self.cost_benchmarks and never read by anything, so
the method is gone. Cost analysis uses the cost_usd the caller supplies per
execution log, which is the only figure that can be accurate
Made model-agnostic, following the precedent already set by
senior-prompt-engineer/scripts/prompt_optimizer.py's --price-per-mtok:
- senior-ml-engineer SKILL.md and llm_integration_guide.md drop both 2024
price tables and the context-window table (which claimed GPT-4 = 8,192).
calculate_cost() takes rates as parameters; count_tokens() takes an
encoding name, since encodings outlive model IDs and
encoding_for_model() raises KeyError on anything unmapped
- OpenAIProvider loses its default model, so the caller must pass one
- llm-cost-optimizer's routing table names tiers, not models
Pinned to current IDs where an example genuinely needs one: SKILL_PIPELINE.md
(claude-opus-4-6 -> claude-opus-5), prompt-governance (claude-sonnet-4-5 ->
claude-sonnet-5), agent-designer README. Both dual-publish copies of the CAIO
pricing move together, so G4 stays green.
TEAM_STRUCTURE_GUIDE.md documented `prompt_optimizer.py --model gpt-4 --task
classification`. That contract no longer exists: there is no --task flag and
`prompt` is a required positional. Replaced with a runnable invocation.
Four references stay, with reasons in the allowlist: two litreview examples
where the retired model is the subject of the literature being reviewed, one
dated Computer Use citation, and the embedding benchmark already labelled a
2024 snapshot.
Assisted-by: Claude Code:claude-opus-5
audit/newgen-2026-06/00-MASTER.md proposed a "model-name freshness ... regex
deny-list for retired model identifiers" gate. It was never built, which is
why retired IDs and 2024 price tables survived both the June and July 2026
audits and are still in the tree today.
check_model_freshness.py flags references that mislead or break on execution:
script defaults, config values, cost tables keyed on a retired model, and
copy-pasteable CLI examples pinning a retired versioned ID. It distinguishes
these from legitimate dated citations, which stay silent when the line carries
a year, an arXiv ID, or wording like "model card" / "as of" / "historical" —
unless the line also looks like a live default, since
`model: str = "claude-3-opus" # 2024 default` still breaks.
Haiku 4.5 is excluded from the Claude 4 sweep in the patterns rather than
per-file, because claude-haiku-4-5-20251001 is current.
Advisory (continue-on-error) for now: it reports 34 references, 13 of them in
executable positions, and the content fixes land in the next change. Flip to
blocking there. --executable-only prints just the 13 that matter first.
Assisted-by: Claude Code:claude-opus-5
Every existing gate reads frontmatter with a regex or a line scan
(generate-docs.py, sync-codex-skills.py, check_paths.py), so a block that is
not valid YAML passed CI while Claude Code loaded the skill with no metadata.
The 14 files fixed in the previous commit had drifted that way unnoticed.
check_frontmatter.py parses each block with yaml.safe_load and enforces what
Claude Code actually reads:
errors - unparseable YAML, non-mapping frontmatter, missing description,
missing agent name, an agent name containing ':' (refused since
CC 2.1.218), or a missing frontmatter block
warnings - keys outside the current skill/agent frontmatter spec, and a
combined description + when_to_use over the 1536-char cap that
the skill listing truncates at
Warnings are non-blocking so this lands without requiring the wider metadata
cleanup; --strict flips them fatal. The run also tallies the off-spec keys no
runtime reads (license 172, metadata 125, domain 76, compatible_tools 37,
triggers 14), which gives that cleanup a worklist regenerated on every run.
Clean on the current tree: 593 files, 0 errors, 17 warnings.
Assisted-by: Claude Code:claude-opus-5
Implements issue #654 Option A (embedded-sample convention) plus the
verification harness the issue asked for:
- scripts/smoke_json_output.py — new advisory gate (G9) that discovers
every tool whose --help advertises JSON output, runs <tool> --sample
<json-flag>, and asserts the stdout parses as JSON. Tools advertising
JSON without --sample are reported as 'uncovered' (a backlog, not a
failure) so the gate can be adopted incrementally; --strict flips that
to a hard failure once coverage is high. Wired into ci-quality-gate.yml
alongside G8.
- Added --sample embedded fixtures to the 5 tools named in #654:
error_budget_calculator, slo_review, blast_radius_calculator,
audit_log_analyzer, api_linter. Their required args are now optional
when --sample is passed; missing-arg behavior is unchanged otherwise.
- Fixed 4 tools the new gate surfaced (prompt_rater, coach_tip_classifier,
cheat_code_filter, redaction_linter): their --sample path printed human
text and ignored --json; it now honors the JSON flag.
- Synced the 3 dual-published standalone copies (slo-architect x2,
chaos-engineering) so the drift guard stays green.
Gate now reports 16 tools covered, 16 verified, 0 failures.
https://claude.ai/code/session_01CUWsrUNZP9jpxvAwq67UiT
- enforce-pr-target.yml: drop the no-op split/trim/join on the comment body
(array join already produces the final text)
- ci-quality-gate.yml: safety findings now emit a workflow warning instead
of being silently absorbed by '|| true'
- check_paths.py: fnmatch import hoisted to module level
- smoke_scripts.py: stale exception entries now fail the gate (exit 3) so
scripts/smoke_exceptions.txt stays tidy
https://claude.ai/code/session_019AJddAL1NADWMXsy1qNPQF
Every advisory run was green through PR #835, so the burn-in SLA
(2026-07-01 or 10 green runs) is satisfied early. Edge cases route to the
in-repo allowlists (check_paths_allowlist.txt, smoke_exceptions.txt)
instead of continue-on-error.
https://claude.ai/code/session_019AJddAL1NADWMXsy1qNPQF
- enforce-pr-target.yml: the maintainer exemption let any maintainer PR
target main — replaced with the branch-based hard rule from CLAUDE.md:
only dev->main promotion PRs are allowed, regardless of author.
Maintainer PRs now fail the check with retarget instructions (not
auto-closed); non-maintainer PRs are commented and closed as before.
Re-checks on edited/ready_for_review so retargeting clears it.
- ci-quality-gate.yml: flip-to-blocking SLA documented for the 4 advisory
gates (2026-07-01 or 10 consecutive green runs on dev)
- CHANGELOG.md: Deprecated/Removed Skills section with migration paths for
command-guide, ai-seo (-> aeo), release-manager (-> changelog-generator)
Addresses automated review feedback on PR #835 (items 1, 3, 6).
https://claude.ai/code/session_019AJddAL1NADWMXsy1qNPQF
Move 54 internal-only files out of the public tree via git rm --cached
+ .gitignore. Files remain on the maintainer's local disk; future fresh
clones see only production skill packages and user-facing docs.
Hidden:
- documentation/ (sprint plans, strategy, roadmaps) — 16 files
- eval-workspace/ (Tessl eval outputs) — 6 files
- megaprompts/ (Path-B draft specs) — 16 files
- tests/ (pytest suite — run locally, not in CI) — 14 files
- .autoresearch/ (autoresearch workspace) — 1 file
- AUDIT_REPORT.md (stale tracked despite prior ignore)
CI: remove the pytest step from ci-quality-gate.yml since tests/ is
no longer tracked. Other quality gates (yamllint, plugin.json schema,
compile-all, safety, link-check) remain in place.
CLAUDE.md: add a Maintainer-Local Folders section near the top so
future readers understand why these paths are referenced but absent
from GitHub.
No history rewrite — going-forward only. Past commits still contain
these files; only HEAD is cleaned.
Issue #686 was the second round of the same Claude Code path-validator
tightening: v2.1.107 rejected bare "./" (fixed in #539 by moving to
"./skills"), then v2.1.133 also rejected "./skills". The validator that
codified the #539 fix was still recommending "./skills" verbatim — so a
future round 3 would have hit the same trap.
This commit makes the validator catch the regression and runs it in CI:
- scripts/check_plugin_json.py
- Reject any "skills" string starting with "./" (catches both
"./skills" and "./skills/sub" patterns)
- Update docstring + error message to point at the layout-correct
forms instead of the now-broken "./skills"
- Recognize "source" and "attribution" as approved extension fields
(already documented in CLAUDE.md but not in the validator), so the
21 pre-existing false-positives go away and CI can run blocking
- Drop the "./" rejection inside arrays — CLAUDE.md says ["./"] is
the correct single-skill-at-root form
- .github/workflows/ci-quality-gate.yml
- Add blocking "Validate plugin.json manifests" step that runs the
validator on every PR
- CLAUDE.md
- Add an Enforcement note pointing at the validator and the lockstep
rule: when CC tightens its path validator again, update validator
rules and CLAUDE.md together
Verified: 69/69 manifests pass; 6-case smoke test confirms validator
rejects all three known-broken forms ("./skills", "./", "./skills/sub")
and accepts all three documented-valid forms ("skills", ["./"],
explicit array).
Fork-based PRs (like PR #498) caused all CI checks to fail due to:
- ci-quality-gate: checkout failed because fork branch names don't exist
in the base repo. Now uses commit SHA for PR events.
- skill-security-audit: comment posting failed with read-only GITHUB_TOKEN.
Now continues on error and writes results to job summary as fallback.
- claude-code-review: fallback comment step failed silently. Now continues
on error and writes status to job summary.
https://claude.ai/code/session_01X1RKFAkEwxgg6gQvJG1KCa