claude-skills/engineering/skills/chaos-engineering/references/attack_taxonomy.md
Claude 23eefc2e9a
feat(skills): ship chaos-engineering (Phase 3 — resilience testing discipline)
Phase 3 of the multi-skill build effort. Same 14-step pipeline. Composes
explicitly with feature-flags-architect (kill switches as abort triggers)
and kubernetes-operator (operators are common chaos targets).

## What landed

### New skill: engineering/chaos-engineering

End-to-end chaos engineering discipline. Published as BOTH:
- Standalone plugin: engineering/chaos-engineering/
- Bundled mirror:    engineering/skills/chaos-engineering/

3 stdlib-only Python tools (Karpathy complexity 95/100 — best in portfolio):
- experiment_designer.py        — generates structured plans with hypothesis,
                                   steady-state, blast radius, abort criteria,
                                   rollback. Refuses to render plans without
                                   abort criteria (exit code 1).
- blast_radius_calculator.py    — computes affected users + error budget
                                   consumption + GREEN/YELLOW/RED risk score.
                                   Validates inputs (0 ≤ traffic-share ≤ 1).
- experiment_postmortem.py      — blameless postmortems from plan + result log;
                                   detects blame-laden language ("fault of",
                                   "should have known", "stupid", etc.) and
                                   warns at write time.

4 reference docs:
- chaos_principles.md      — 4 founding principles + 5th abort principle,
                              maturity model, history, when-to-start checklist
- experiment_design.md      — 7-section plan structure, pre-flight checklist,
                              time-boxing, escalation
- attack_taxonomy.md        — 7 attack types (latency / error / resource /
                              network-partition / dependency-failure / time-skew
                              / infrastructure) with magnitudes and tooling
- tooling_landscape.md      — Chaos Toolkit / Mesh / Litmus / Gremlin / AWS FIS
                              / DIY decision tree

Templates:
- experiment_template.md    — fill-in plan with all 7 sections
- postmortem_template.md    — blameless postmortem structure

Plus: SKILL.md (213 lines), README.md, /chaos-experiment slash command.

### Audit verdict (evidence-based)

Closest existing skills:
- engineering-team/incident-response — for actual incidents, not prevention
- engineering-team/red-team — adversarial; different goal (find attack paths)
- engineering-team/threat-detection — hunting; different goal
- engineering/observability-designer — measurement, not fault injection
None cover the chaos-engineering discipline (hypothesis-driven fault injection
with bounded blast radius). Verdict: BUILD. Gap is real and tooling-shaped.

### Composition story (Phase 1+2+3 form a stack)

```
feature-flags-architect.kill_switch_audit.py
  ↓ defines kill switches that ↓
chaos-engineering.experiment_designer.py
  ↓ designs experiments against ↓
kubernetes-operator (and other targets)
```

Together: a complete progressive-delivery + resilience-testing stack.

### Marketplace / registry

- marketplace.json: chaos-engineering registered as standalone plugin
- engineering-advanced-skills bundle: 47 → 48 skills, version → 2.4.2
- engineering/.claude-plugin/plugin.json: version + skill list updated
- mkdocs.yml: nav entry under "Engineering - POWERFUL"
- docs/skills/engineering/chaos-engineering.md: docs page (manual,
  pending generate-docs.py classification fix)
- docs/commands/chaos-experiment.md: auto-generated
- .codex/, .gemini/: synced

### Karpathy-coder gates

- complexity_checker (strict): 95/100 average — BEST score in the new
  portfolio. Only 1 WARN (depth 5 in blast_radius_calculator.py validation
  branches; the other 2 scripts hit no findings whatsoever).
- All 1666 tests pass (was 1648; added 18 for the new skill).
- mkdocs build --strict: succeeded in 13.33s.

### Verifiable success criteria (all green)

✓  scripts/*.py --help     → exit 0 for all 3 scripts
✓  SKILL.md frontmatter    → name + description + tags + compatible_tools
✓  plugin.json schema      → 8 fields exact (verified by check_plugin_json.py)
✓  sync_skill_bundles      → standalone ↔ bundled mirror in sync
✓  marketplace.json        → standalone entry + bundle counts updated
✓  generate-docs.py        → command page generated (skill page manual)
✓  mkdocs build --strict   → succeeded
✓  cross-tool sync         → codex + gemini synced
✓  pytest tests/           → 1666 passed, 0 failed
✓  CHANGELOG.md            → [Unreleased] entry expanded for Phase 3
✓  Self-test (RED case)    → 50% blast radius on 99.9% baseline correctly
                             classifies as RED (17.33% of monthly budget) and
                             returns ABORT recommendation
✓  Composition test        → references named skills explicitly compose

## Phase 1+2+3 cumulative

- 3 new skills: feature-flags-architect, kubernetes-operator, chaos-engineering
- 9 new Python tools (all stdlib, all <200 LOC, average complexity 90/100)
- 12 new reference docs (~250-500 lines each)
- 3 new slash commands (/flag-cleanup, /operator-audit, /chaos-experiment)
- 2 repo-infrastructure scripts (sync_skill_bundles, check_plugin_json)
- 1 pre-existing test fix (full-page-screenshot CI red)

## Files

- engineering/chaos-engineering/                                (new standalone plugin)
- engineering/skills/chaos-engineering/                         (new bundled mirror)
- commands/chaos-experiment.md                                  (new slash command)
- docs/skills/engineering/chaos-engineering.md                  (new docs page)
- docs/commands/chaos-experiment.md                             (auto-generated)
- mkdocs.yml                                                    (nav entries)
- .claude-plugin/marketplace.json                               (registered)
- engineering/.claude-plugin/plugin.json                        (bundle bumped)
- CHANGELOG.md                                                  ([Unreleased] expanded)
- .codex/, .gemini/                                             (cross-tool sync)

https://claude.ai/code/session_01Dq12xJakFRxwaoU8Pqejdm
2026-05-09 21:24:16 +00:00

5.7 KiB

Attack taxonomy

7 categories of fault injection. Each tests a different system property. Pick the one whose failure mode matches your hypothesis.

1. Latency

What it tests: timeouts, retries, circuit breakers, fallback paths.

Inject: add N ms of delay to network responses to a target.

When to use:

  • "What if dependency X is slow?"
  • "Are timeouts configured correctly upstream?"
  • "Does the retry budget kick in?"

Tools:

  • Linux tc (traffic control) — direct kernel-level shaping
  • Chaos Mesh NetworkChaos (delay)
  • Toxiproxy — proxy-based, language-agnostic
  • AWS FIS — aws:network:traffic-control action

Example magnitude: +200ms (90% of typical timeouts), +2000ms (test backoff), +30s (test giving-up logic).

2. Error injection

What it tests: error handling paths, fallback behavior, retry policies.

Inject: return errors (5xx, exceptions) for a fraction of requests.

When to use:

  • "What happens when X starts failing?"
  • "Does the fallback path actually work in prod?"
  • "Are we logging errors correctly?"

Tools:

  • Chaos Mesh HTTPChaos
  • Service mesh (Istio, Linkerd) fault injection
  • Toxiproxy with error toxic
  • Application-level feature flag for synthetic errors

Example magnitude: 1% errors (test handler), 50% errors (test retry), 100% errors (test fallback path).

3. Resource exhaustion

What it tests: saturation handling, autoscaling, OOM behavior, disk-full handling.

Inject: consume CPU, memory, or disk on the target.

When to use:

  • "What if memory leaks?"
  • "Does the autoscaler kick in?"
  • "What happens when disk fills?"

Sub-types:

  • CPU pressure — peg cores at N% usage
  • Memory pressure — allocate large blocks
  • Disk fill — write large files until partition fills
  • I/O saturation — high random read/write

Tools:

  • stress-ng — CPU/memory/IO/disk
  • Chaos Mesh StressChaos and IOChaos
  • AWS FIS aws:ssm:send-command with stress-ng

Example magnitude: 80% CPU sustained, 90% memory, fill /var to 95%.

4. Network partition

What it tests: consensus protocols, leader election, split-brain prevention, region failover.

Inject: drop all packets between a set of hosts.

When to use:

  • "What if AZ-A loses connectivity to AZ-B?"
  • "Does the database elect a new primary?"
  • "Does the cluster avoid split-brain?"

Tools:

  • Chaos Mesh NetworkChaos (partition mode)
  • tc with iptables drop rules
  • AWS FIS aws:network:disrupt-connectivity

Example magnitude: drop 100% to peer X (full partition), drop 50% (degraded link).

5. Dependency failure

What it tests: graceful degradation, fallback to cache, fallback to default values.

Inject: make a downstream dependency unavailable (timeout, refuse connections).

When to use:

  • "What if the rec engine goes down?"
  • "Does Search degrade gracefully when ML models are unreachable?"
  • "Is cache the fallback for the user-pref service?"

Tools:

  • Service mesh fault injection (most flexible)
  • Toxiproxy
  • iptables rules to refuse connections
  • Chaos Mesh NetworkChaos with corrupt or drop

Example magnitude: 100% requests to dep X timeout (full outage), 25% timeout (intermittent), 0% available for 5 min (sustained outage).

6. Time skew

What it tests: time-sensitive logic — token expiry, cron schedules, TTLs, retry backoff.

Inject: alter the wall clock seen by a process.

When to use:

  • "What if NTP fails?"
  • "What if a process clock drifts +5 minutes?"
  • "Do tokens correctly fail validation when expired?"
  • "Does cron skip or double-fire?"

Tools:

  • libfaketime — preload library
  • Chaos Mesh TimeChaos
  • Custom: change container's /etc/localtime

Example magnitude: +1 minute (subtle), +5 minutes (TLS / token failures), +1 day (catastrophic for some logic).

Caution: time skew can cause cluster-wide consensus failures. Test in isolation first.

7. Infrastructure (kill instance / pod / container)

What it tests: auto-recovery, failover, replica count maintenance.

Inject: terminate an instance, pod, or container.

When to use:

  • "Does Kubernetes restart the pod?"
  • "Does the load balancer remove the instance from rotation?"
  • "Is the replication factor maintained?"

Tools:

  • Chaos Monkey (the original)
  • Chaos Mesh PodChaos (kill, fail)
  • AWS FIS aws:ec2:terminate-instances
  • kubectl delete pod (manual, simplest)

Example magnitude: kill 1 of N pods (Chaos Monkey level), kill all pods of a deployment (test recreation), kill 1 of 3 replica DB nodes (test failover).

Choosing an attack

Hypothesis pattern Attack type
"What if X is slow?" Latency
"What if X is failing?" Error
"What if we run hot?" Resource
"What if regions partition?" Network partition
"What if dep X is down?" Dependency failure
"What if clocks drift?" Time skew
"What if a node dies?" Infrastructure

Combining attacks

Real outages often combine attacks (e.g., latency + saturation). Once basic experiments are stable, run combinations:

  • Latency on dependency + CPU pressure on app → tests timeout + retry budget interaction
  • Pod kill + network partition → tests recovery during a partition
  • Disk fill + dependency failure → tests fallback path while disk is constrained

Combinations have higher risk; reduce blast radius accordingly.

Severity ladder

S1 — Latency (small)              ← start here
S2 — Error injection (low %)
S3 — Resource pressure (CPU/mem)
S4 — Latency (large) / errors (high %)
S5 — Single instance kill
S6 — Network partition (single peer)
S7 — Multiple instance kill
S8 — Region partition / time skew
S9 — Combinations of S5-S8        ← here be dragons

Don't skip levels. Earn confidence at S1-S3 before attempting S5+.