- Remove focused-fix from Integration Points (was a substitution for
systematic-debugging which does not exist in this repo)
- Add scripts/ship_gate_scanner.py: stdlib-only pre-production audit CLI
covering all 8 categories (SEC, DB, CODE, DEP, AI, DEPLOY, FE, OBS)
with JSON output, ANSI color, interactive manual checks, and exit codes
Implements Karpathy's 4 coding principles (Think Before Coding, Simplicity
First, Surgical Changes, Goal-Driven Execution) as an active enforcement
plugin, not just passive guidelines. Derived from Karpathy's X post on LLM
coding pitfalls but goes far beyond the source material with automated
detection tools, a review agent, and CI integration patterns.
Differentiator vs forrestchang/andrej-karpathy-skills (prompt-only, single
SKILL.md): this version ships real tooling that DETECTS violations instead
of just documenting principles.
Plugin contents (engineering/karpathy-coder/):
- SKILL.md with `context: fork` for skill chaining
- 4 Python tools (stdlib only):
- complexity_checker.py — cyclomatic complexity, class density, nesting
depth, function length, premature abstractions (Principle #2)
- diff_surgeon.py — diff noise ratio: comment-only changes, whitespace,
style drift, drive-by refactors, quote-style swaps (Principle #3)
- assumption_linter.py — detects "just", "obviously", "should work",
vague actions, unscoped users, missing format specs (Principle #1)
- goal_verifier.py — scores plan steps 0-3 for verification quality,
flags vague criteria, checks for final verification (Principle #4)
- 1 sub-agent: karpathy-reviewer (runs all 4 principles against a diff)
- 1 slash command: /karpathy-check (dispatches the reviewer)
- 1 pre-commit hook: karpathy-gate.sh (non-blocking, warns on violations)
- 3 reference docs: karpathy-principles.md (full context + when to relax),
anti-patterns.md (10+ before/after examples), enforcement-patterns.md
(Husky, pre-commit framework, GitHub Actions CI integration)
- .claude-plugin/plugin.json manifest (v2.3.0)
- Cross-tool compatible: works with any AGENTS.md-based CLI
All 4 scripts verified: --help passes, smoke tests run correctly.
complexity_checker catches its own nesting depth. assumption_linter
correctly flags "just", "obviously", "should work". goal_verifier
correctly scores plans with/without verification steps.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Issue #506 reported that `hooks/hooks.json` used `./hooks/error-capture.sh`
which fails for any session started outside the plugin dir. That specific
file was already fixed in commit 217b199 (which closed#392) — both
`hooks/hooks.json` and `settings.json` already use `${CLAUDE_PLUGIN_ROOT}`.
However, two stale example paths were still surfacing the bug in
documentation:
1. `engineering-team/self-improving-agent/CLAUDE.md` line 74 — "To enable"
example with `./skills/self-improving-agent/hooks/error-capture.sh`
2. `engineering-team/self-improving-agent/hooks/error-capture.sh` header
comment — install example with the same broken path
Both examples would teach users to copy the broken pattern into their own
settings.json, reproducing the exact bug #506 describes.
Fix: rewrite both examples to use `${CLAUDE_PLUGIN_ROOT}/hooks/error-capture.sh`
and add explicit "do not use relative paths" warnings. Also clarify in
CLAUDE.md that manual hook wiring is NOT needed when installing via
`/plugin install` — the hook is registered automatically from the plugin's
hooks.json.
Fixes#506
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Bug: after `/plugin install self-improving-agent@claude-code-skills`, only
1 skill appeared and `/si:review`, `/si:promote`, `/si:extract`, `/si:status`,
`/si:remember` were all unknown commands. The 5 sub-skills were silently
registered under the wrong namespace.
Root cause: two issues in the plugin manifest layer.
1. **Slash-command namespace is derived from `.claude-plugin/plugin.json`
`name`**, not from the marketplace entry name, the settings.json name, or
frontmatter. Previous `name: "self-improving-agent"` caused sub-skills to
register as `/self-improving-agent:review` etc — never matching the
documented `/si:*` commands.
2. **`command: /si:<op>` frontmatter in sub-skill SKILL.md files is a
non-standard field** not in the Claude Code Skills spec. Claude Code
silently ignores it. It created the illusion that the commands were being
registered when they were not.
Fix:
- Change `engineering-team/self-improving-agent/.claude-plugin/plugin.json`
`name` from "self-improving-agent" → "si". This is the namespace root; it
does NOT affect the marketplace install identifier (which stays
`self-improving-agent` via the marketplace.json `name` field). After the
fix, skills register as `/si:review`, `/si:promote`, `/si:extract`,
`/si:status`, `/si:remember` — matching the README and CLAUDE.md docs.
- Remove the non-standard `command: /si:<op>` frontmatter line from all 5
sub-skill SKILL.md files (review, promote, extract, status, remember).
Frontmatter now contains only `name` and `description` per the Claude Code
Skills spec.
- Bump plugin.json version 2.1.2 → 2.3.0 to match repo release.
- Update marketplace.json entry: version 2.2.0 → 2.3.0, expand description
to list all 5 slash commands and 2 sub-agents.
OpenClaw compat: the legacy `settings.json` inside the skill directory still
uses `"name": "self-improving-agent"` for OpenClaw's install path. Claude Code
ignores settings.json entirely, so this is safe to leave as-is.
Install flow (unchanged, verified correct after fix):
/plugin marketplace add alirezarezvani/claude-skills
/plugin install self-improving-agent@claude-code-skills
# → 5 skills register as /si:review, /si:promote, /si:extract,
# /si:status, /si:remember
Known related issue (not fixed in this PR to keep scope tight): `agenthub`
has the identical bug. Its plugin.json `name` is "agenthub", so sub-skills
register as `/agenthub:init` rather than the documented `/hub:init`. Same
fix applies: rename plugin.json `name` to "hub". Will file as a follow-up.
Fixes#505
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Implements Karpathy's LLM Wiki pattern as a production-grade plugin. The LLM
incrementally ingests sources into a persistent, interlinked Obsidian vault —
updating entity/concept/source pages, flagging contradictions, maintaining an
index and append-only log. Knowledge compounds instead of being re-derived by
RAG on every query.
Plugin contents (engineering/llm-wiki/):
- SKILL.md with `context: fork` frontmatter for skill chaining
- 3 sub-agents: wiki-ingestor, wiki-librarian, wiki-linter
- 5 slash commands: /wiki-init, /wiki-ingest, /wiki-query, /wiki-lint, /wiki-log
- 8 Python tools (stdlib only): init_vault, ingest_source, update_index,
append_log, wiki_search (BM25), lint_wiki, graph_analyzer, export_marp
- 8 reference docs: schema, page-formats, ingest/query/lint workflows,
obsidian-setup, cross-tool-setup, memex-principles
- Vault templates: CLAUDE.md, AGENTS.md, .cursorrules, index.md, log.md,
5 page templates (entity, concept, source, comparison, synthesis)
- Worked example vault on "LLM interpretability"
- .claude-plugin/plugin.json manifest
Cross-tool compatibility: the scripts are pure Python stdlib. Only the schema
loader changes per tool (CLAUDE.md for Claude Code, AGENTS.md for Codex CLI /
Cursor / Antigravity / OpenCode / Gemini CLI, .cursorrules for legacy Cursor).
init_vault.py --tool all installs all three.
Repo-level registration:
- Commands mirrored to top-level commands/ for repo-wide discovery
- Agents mirrored to agents/engineering/ as cs-wiki-{ingestor,librarian,linter}
- .claude-plugin/marketplace.json: new llm-wiki entry + version bump to v2.3.0
- CLAUDE.md updated: 234 skills, 313 Python tools, 432 refs, 28 agents, 27 commands
Also saved (deferred): craighewitt-mattpocock reimplementation plan at
documentation/implementation/ — 4-pod proposal for building better versions
of selected skills from thecraighewitt-skills and mattpocock-skills
collections. Not executed; awaiting user confirmation on scope.
End-to-end smoke test passed: init_vault → ingest → update_index → append_log
→ wiki_search → lint → graph_analyzer → export_marp all run against a fresh
vault with real pages, wikilinks, and frontmatter.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Fork-based PRs (like PR #498) caused all CI checks to fail due to:
- ci-quality-gate: checkout failed because fork branch names don't exist
in the base repo. Now uses commit SHA for PR events.
- skill-security-audit: comment posting failed with read-only GITHUB_TOKEN.
Now continues on error and writes results to job summary as fallback.
- claude-code-review: fallback comment step failed silently. Now continues
on error and writes status to job summary.
https://claude.ai/code/session_01X1RKFAkEwxgg6gQvJG1KCa
Self-contained skill for tracking technical changes with structured JSON
records, an enforced state machine, and a session handoff format that lets
a new AI session resume work cleanly when a previous one expires.
Includes:
- 5 stdlib-only Python scripts (init, create, update, status, validator)
all supporting --help and --json
- 3 reference docs (lifecycle state machine, JSON schema, handoff format)
- /tc dispatcher in commands/tc.md
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Move from data-analysis/ to engineering/
- Fix 5 cross-references to use correct domain paths
- Fix Python 3.9 compat in sample_size_calculator.py
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Adds statistical-analyst skill — fills a gap in the repo (no hypothesis
testing or experiment analysis tooling exists; only ab-test-setup for
instrumentation, but zero analysis capability).
Three stdlib-only Python scripts:
- hypothesis_tester.py: Z-test (proportions), Welch's t-test (means),
Chi-square (categorical) with p-value, CI, Cohen's d/h, Cramér's V
- sample_size_calculator.py: required n per variant for proportion and
mean tests, with power/MDE tradeoff table and duration estimates
- confidence_interval.py: Wilson score interval (proportions) and
z-based interval (means) with margin of error and precision notes
Validator: 86.4/100 (GOOD). Security audit: PASS (0 critical/high).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>