Phase 3 of the multi-skill build effort. Same 14-step pipeline. Composes
explicitly with feature-flags-architect (kill switches as abort triggers)
and kubernetes-operator (operators are common chaos targets).
## What landed
### New skill: engineering/chaos-engineering
End-to-end chaos engineering discipline. Published as BOTH:
- Standalone plugin: engineering/chaos-engineering/
- Bundled mirror: engineering/skills/chaos-engineering/
3 stdlib-only Python tools (Karpathy complexity 95/100 — best in portfolio):
- experiment_designer.py — generates structured plans with hypothesis,
steady-state, blast radius, abort criteria,
rollback. Refuses to render plans without
abort criteria (exit code 1).
- blast_radius_calculator.py — computes affected users + error budget
consumption + GREEN/YELLOW/RED risk score.
Validates inputs (0 ≤ traffic-share ≤ 1).
- experiment_postmortem.py — blameless postmortems from plan + result log;
detects blame-laden language ("fault of",
"should have known", "stupid", etc.) and
warns at write time.
4 reference docs:
- chaos_principles.md — 4 founding principles + 5th abort principle,
maturity model, history, when-to-start checklist
- experiment_design.md — 7-section plan structure, pre-flight checklist,
time-boxing, escalation
- attack_taxonomy.md — 7 attack types (latency / error / resource /
network-partition / dependency-failure / time-skew
/ infrastructure) with magnitudes and tooling
- tooling_landscape.md — Chaos Toolkit / Mesh / Litmus / Gremlin / AWS FIS
/ DIY decision tree
Templates:
- experiment_template.md — fill-in plan with all 7 sections
- postmortem_template.md — blameless postmortem structure
Plus: SKILL.md (213 lines), README.md, /chaos-experiment slash command.
### Audit verdict (evidence-based)
Closest existing skills:
- engineering-team/incident-response — for actual incidents, not prevention
- engineering-team/red-team — adversarial; different goal (find attack paths)
- engineering-team/threat-detection — hunting; different goal
- engineering/observability-designer — measurement, not fault injection
None cover the chaos-engineering discipline (hypothesis-driven fault injection
with bounded blast radius). Verdict: BUILD. Gap is real and tooling-shaped.
### Composition story (Phase 1+2+3 form a stack)
```
feature-flags-architect.kill_switch_audit.py
↓ defines kill switches that ↓
chaos-engineering.experiment_designer.py
↓ designs experiments against ↓
kubernetes-operator (and other targets)
```
Together: a complete progressive-delivery + resilience-testing stack.
### Marketplace / registry
- marketplace.json: chaos-engineering registered as standalone plugin
- engineering-advanced-skills bundle: 47 → 48 skills, version → 2.4.2
- engineering/.claude-plugin/plugin.json: version + skill list updated
- mkdocs.yml: nav entry under "Engineering - POWERFUL"
- docs/skills/engineering/chaos-engineering.md: docs page (manual,
pending generate-docs.py classification fix)
- docs/commands/chaos-experiment.md: auto-generated
- .codex/, .gemini/: synced
### Karpathy-coder gates
- complexity_checker (strict): 95/100 average — BEST score in the new
portfolio. Only 1 WARN (depth 5 in blast_radius_calculator.py validation
branches; the other 2 scripts hit no findings whatsoever).
- All 1666 tests pass (was 1648; added 18 for the new skill).
- mkdocs build --strict: succeeded in 13.33s.
### Verifiable success criteria (all green)
✓ scripts/*.py --help → exit 0 for all 3 scripts
✓ SKILL.md frontmatter → name + description + tags + compatible_tools
✓ plugin.json schema → 8 fields exact (verified by check_plugin_json.py)
✓ sync_skill_bundles → standalone ↔ bundled mirror in sync
✓ marketplace.json → standalone entry + bundle counts updated
✓ generate-docs.py → command page generated (skill page manual)
✓ mkdocs build --strict → succeeded
✓ cross-tool sync → codex + gemini synced
✓ pytest tests/ → 1666 passed, 0 failed
✓ CHANGELOG.md → [Unreleased] entry expanded for Phase 3
✓ Self-test (RED case) → 50% blast radius on 99.9% baseline correctly
classifies as RED (17.33% of monthly budget) and
returns ABORT recommendation
✓ Composition test → references named skills explicitly compose
## Phase 1+2+3 cumulative
- 3 new skills: feature-flags-architect, kubernetes-operator, chaos-engineering
- 9 new Python tools (all stdlib, all <200 LOC, average complexity 90/100)
- 12 new reference docs (~250-500 lines each)
- 3 new slash commands (/flag-cleanup, /operator-audit, /chaos-experiment)
- 2 repo-infrastructure scripts (sync_skill_bundles, check_plugin_json)
- 1 pre-existing test fix (full-page-screenshot CI red)
## Files
- engineering/chaos-engineering/ (new standalone plugin)
- engineering/skills/chaos-engineering/ (new bundled mirror)
- commands/chaos-experiment.md (new slash command)
- docs/skills/engineering/chaos-engineering.md (new docs page)
- docs/commands/chaos-experiment.md (auto-generated)
- mkdocs.yml (nav entries)
- .claude-plugin/marketplace.json (registered)
- engineering/.claude-plugin/plugin.json (bundle bumped)
- CHANGELOG.md ([Unreleased] expanded)
- .codex/, .gemini/ (cross-tool sync)
https://claude.ai/code/session_01Dq12xJakFRxwaoU8Pqejdm
5.7 KiB
Attack taxonomy
7 categories of fault injection. Each tests a different system property. Pick the one whose failure mode matches your hypothesis.
1. Latency
What it tests: timeouts, retries, circuit breakers, fallback paths.
Inject: add N ms of delay to network responses to a target.
When to use:
- "What if dependency X is slow?"
- "Are timeouts configured correctly upstream?"
- "Does the retry budget kick in?"
Tools:
- Linux
tc(traffic control) — direct kernel-level shaping - Chaos Mesh
NetworkChaos(delay) - Toxiproxy — proxy-based, language-agnostic
- AWS FIS —
aws:network:traffic-controlaction
Example magnitude: +200ms (90% of typical timeouts), +2000ms (test backoff), +30s (test giving-up logic).
2. Error injection
What it tests: error handling paths, fallback behavior, retry policies.
Inject: return errors (5xx, exceptions) for a fraction of requests.
When to use:
- "What happens when X starts failing?"
- "Does the fallback path actually work in prod?"
- "Are we logging errors correctly?"
Tools:
- Chaos Mesh
HTTPChaos - Service mesh (Istio, Linkerd) fault injection
- Toxiproxy with error toxic
- Application-level feature flag for synthetic errors
Example magnitude: 1% errors (test handler), 50% errors (test retry), 100% errors (test fallback path).
3. Resource exhaustion
What it tests: saturation handling, autoscaling, OOM behavior, disk-full handling.
Inject: consume CPU, memory, or disk on the target.
When to use:
- "What if memory leaks?"
- "Does the autoscaler kick in?"
- "What happens when disk fills?"
Sub-types:
- CPU pressure — peg cores at N% usage
- Memory pressure — allocate large blocks
- Disk fill — write large files until partition fills
- I/O saturation — high random read/write
Tools:
stress-ng— CPU/memory/IO/disk- Chaos Mesh
StressChaosandIOChaos - AWS FIS
aws:ssm:send-commandwith stress-ng
Example magnitude: 80% CPU sustained, 90% memory, fill /var to 95%.
4. Network partition
What it tests: consensus protocols, leader election, split-brain prevention, region failover.
Inject: drop all packets between a set of hosts.
When to use:
- "What if AZ-A loses connectivity to AZ-B?"
- "Does the database elect a new primary?"
- "Does the cluster avoid split-brain?"
Tools:
- Chaos Mesh
NetworkChaos(partition mode) tcwith iptables drop rules- AWS FIS
aws:network:disrupt-connectivity
Example magnitude: drop 100% to peer X (full partition), drop 50% (degraded link).
5. Dependency failure
What it tests: graceful degradation, fallback to cache, fallback to default values.
Inject: make a downstream dependency unavailable (timeout, refuse connections).
When to use:
- "What if the rec engine goes down?"
- "Does Search degrade gracefully when ML models are unreachable?"
- "Is cache the fallback for the user-pref service?"
Tools:
- Service mesh fault injection (most flexible)
- Toxiproxy
- iptables rules to refuse connections
- Chaos Mesh
NetworkChaoswithcorruptordrop
Example magnitude: 100% requests to dep X timeout (full outage), 25% timeout (intermittent), 0% available for 5 min (sustained outage).
6. Time skew
What it tests: time-sensitive logic — token expiry, cron schedules, TTLs, retry backoff.
Inject: alter the wall clock seen by a process.
When to use:
- "What if NTP fails?"
- "What if a process clock drifts +5 minutes?"
- "Do tokens correctly fail validation when expired?"
- "Does cron skip or double-fire?"
Tools:
libfaketime— preload library- Chaos Mesh
TimeChaos - Custom: change container's
/etc/localtime
Example magnitude: +1 minute (subtle), +5 minutes (TLS / token failures), +1 day (catastrophic for some logic).
Caution: time skew can cause cluster-wide consensus failures. Test in isolation first.
7. Infrastructure (kill instance / pod / container)
What it tests: auto-recovery, failover, replica count maintenance.
Inject: terminate an instance, pod, or container.
When to use:
- "Does Kubernetes restart the pod?"
- "Does the load balancer remove the instance from rotation?"
- "Is the replication factor maintained?"
Tools:
- Chaos Monkey (the original)
- Chaos Mesh
PodChaos(kill, fail) - AWS FIS
aws:ec2:terminate-instances kubectl delete pod(manual, simplest)
Example magnitude: kill 1 of N pods (Chaos Monkey level), kill all pods of a deployment (test recreation), kill 1 of 3 replica DB nodes (test failover).
Choosing an attack
| Hypothesis pattern | Attack type |
|---|---|
| "What if X is slow?" | Latency |
| "What if X is failing?" | Error |
| "What if we run hot?" | Resource |
| "What if regions partition?" | Network partition |
| "What if dep X is down?" | Dependency failure |
| "What if clocks drift?" | Time skew |
| "What if a node dies?" | Infrastructure |
Combining attacks
Real outages often combine attacks (e.g., latency + saturation). Once basic experiments are stable, run combinations:
- Latency on dependency + CPU pressure on app → tests timeout + retry budget interaction
- Pod kill + network partition → tests recovery during a partition
- Disk fill + dependency failure → tests fallback path while disk is constrained
Combinations have higher risk; reduce blast radius accordingly.
Severity ladder
S1 — Latency (small) ← start here
S2 — Error injection (low %)
S3 — Resource pressure (CPU/mem)
S4 — Latency (large) / errors (high %)
S5 — Single instance kill
S6 — Network partition (single peer)
S7 — Multiple instance kill
S8 — Region partition / time skew
S9 — Combinations of S5-S8 ← here be dragons
Don't skip levels. Earn confidence at S1-S3 before attempting S5+.