Runs engineering/book-to-skill end to end on its first real source: OpenAI's Spinning Up in Deep RL (MIT, (c) 2018 OpenAI; primarily developed by Joshua Achiam). Cloned openai/spinningup and compiled its docs/ reStructuredText tree (38 files, ~37k words, ~49K tokens) through the full pipeline -- extract --mode technical, analysis, 20 chapter files, glossary/patterns/cheatsheet, master SKILL.md, validator, plugin emitter. The compiled skill passes book_skill_validator.py in --strict mode with every file inside budget: a 2,101-token resident core (cap 4,000) plus 20 on-demand chapters averaging ~1,256 tokens each. Chapter structure follows the source's own toctree rather than a heading scan: user documentation (ch01-06), Introduction to RL Parts 1-3 (ch07-09), the researcher essay / key papers / exercises / benchmarks (ch10-13), one chapter per algorithm in lineage order (ch14-19: VPG to TRPO to PPO, DDPG to TD3 and SAC), and the logger/MPI/ExperimentGrid utilities (ch20). Rights basis is open-license, not fair use -- the emitter's Step-11 gate refuses a shareable package without one. Upstream's MIT notice is reproduced in full in the plugin's LICENSE beside this package's own, and README.md names the source, the author and the source's frozen version; a sidecar JSON is not a license notice. Also fixes a defect the emitter only reveals at its final step: skill_plugin_emitter.py wrote its whole `source` provenance block into plugin.json, on a stale inline claim that `source`/`attribution` were approved extension fields. Claude Code rejects an entire manifest on any unrecognized key (issue #954) and scripts/check_plugin_json.py hard-fails such a manifest, so every package the emitter produced failed the blocking CI gate on commit. _plugin_manifest() now emits spec fields only and a new _authoring_notes() writes .claude-plugin/authoring-notes.json. Recorded as deviation 26 in engineering/book-to-skill/README.md; the printed marketplace.json snippet is unchanged, since `source` is a valid key there. Counters: skills 386 -> 387, agents 116 -> 117, commands 146 -> 147, plugins 97 -> 98. Tools and references unchanged -- a compiled knowledge base ships notes, not scripts. All blocking CI gates verified locally: compileall, check_plugin_json --all, check_skill_names, check_paths, check_frontmatter, check_dual_publish, check_model_freshness, smoke_scripts (692/692), derive_counters --check. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UySnyf5upm4y8xhYA3w6yw
2.4 KiB
Spinning Up in Deep RL
Knowledge-base plugin compiled from Spinning Up in Deep RL by Joshua Achiam (OpenAI) by
engineering/book-to-skill. 20 chapters indexed.
What is in here
| File | Contents |
|---|---|
skills/spinning-up-deep-rl/SKILL.md |
Core frameworks, chapter index, topic index (resident, under 4k tokens) |
skills/spinning-up-deep-rl/chapters/ |
One summary per chapter — loaded on demand, never all at once |
skills/spinning-up-deep-rl/glossary.md |
Every significant term, alphabetized, with its chapter |
skills/spinning-up-deep-rl/patterns.md |
Techniques and design patterns with trade-offs |
skills/spinning-up-deep-rl/cheatsheet.md |
Decision rules, thresholds and trade-off matrices |
Use
/cs:spinning-up-deep-rl # core frameworks + chapter index
/cs:spinning-up-deep-rl <topic> # resolve via topic index, read one chapter
/cs:spinning-up-deep-rl ch05 # read one chapter summary
Or invoke the cs-spinning-up-deep-rl agent for a working session anchored to this source.
Provenance and limits
Source: OpenAI's Spinning Up in Deep RL
(openai/spinningup), primarily developed by
Joshua Achiam. Compiled from the docs/ reStructuredText tree at the January 2020
PyTorch update.
Rights basis: open-license. The source is MIT, Copyright (c) 2018 OpenAI, which
permits derivative distribution. The full upstream notice is reproduced in
LICENSE alongside this package's own; the top-level license field in
plugin.json covers the scaffolding only.
Generated, not hand-authored: every claim traces to the source document. It carries that source's blind spots, and it is a set of structured notes — not a copy of the work and not a substitute for reading it.
What it does not cover: DQN and the discrete-action value-learning family, recurrent or
convolutional architectures, partially-observed settings, model-based implementations, and any
deep RL work after early 2020. The six implementations documented are educational; ch13 records
which are research-grade (DDPG, TD3, SAC) and which are not (VPG, TRPO, PPO).
Distribution: shareable. Regenerate or extend with
python3 engineering/book-to-skill/skills/book-to-skill/scripts/extract_document.py, then re-run
book_skill_validator.py before loading the result.