Phase 4 of the multi-skill build effort. Same 14-step pipeline.
## What landed
### New skill: engineering/slo-architect
End-to-end SLO discipline per Google SRE Workbook. Published as BOTH:
- Standalone plugin: engineering/slo-architect/
- Bundled mirror: engineering/skills/slo-architect/
3 stdlib-only Python tools (Karpathy complexity 95/100):
- slo_designer.py — generates SLO definitions; refuses to render
if required fields missing (owner, policy doc,
SLI numerator/denominator). Supports 5 SLI
types: request-success-rate, request-latency,
availability-time, data-freshness, correctness.
- error_budget_calculator.py — computes error budget AND the canonical
multi-window burn-rate alert thresholds:
fast (1h/5m, page), slow (6h/30m, page),
ticket (3d/6h). Output is PromQL-shaped,
ready to paste into Prometheus rules.
- slo_review.py — audits SLO docs for 7 common bugs:
target ≥99.99, target ≤99, window <7d,
window >90d, no SLI definition, no error
budget policy, CPU-as-SLI.
4 reference docs:
- slo_principles.md — SLI vs SLO vs SLA, Google SRE Workbook canon
- sli_design.md — 5 SLI types with examples and anti-patterns
- error_budget.md — error budget math, burn-rate alerts, budget policy
- composition.md — how SLOs feed feature-flags, chaos, kubernetes-operator
Asset templates:
- slo_template.yaml — fillable SLO YAML with all required fields
- error_budget_policy.md — fillable 4-state policy (HEALTHY / CAUTION /
CRITICAL / VIOLATED)
Plus: SKILL.md, README.md, /slo-design slash command.
## Composition with prior phases
Explicit wire-up to the rest of the portfolio:
- feature-flags-architect.kill_switch_audit references SLO burn-rate
- chaos-engineering.blast_radius_calculator takes SLO error budget as input
- kubernetes-operator capability level L4 requires SLOs + Prometheus rules
The SLO is the unifying number: rollout abort, chaos blast radius, and
operator capability all reference it. references/composition.md walks
through end-to-end use.
## Audit verdict (evidence-based)
Closest existing skill: engineering/observability-designer covers SLI/SLO as
ONE topic among many (metrics, logs, traces, dashboards, alerting). It has
no dedicated tools and is breadth-not-depth. slo-architect is the focused
SLO discipline with deterministic Python tools — same gap pattern as
kubernetes-operator vs senior-devops.
## Marketplace / registry
- marketplace.json: slo-architect registered as standalone plugin
- engineering-advanced-skills bundle: 49 → 50 skills, version → 2.4.4
- engineering/.claude-plugin/plugin.json: version + skill list updated
- mkdocs.yml: nav entry under "Engineering - POWERFUL"
- docs/skills/engineering/slo-architect.md: docs page (manual)
- docs/commands/slo-design.md: auto-generated
- .codex/, .gemini/: synced
## Karpathy-coder gates
- complexity_checker (strict): 95/100 average — same top score as
chaos-engineering. 1 WARN (depth 7 in slo_review.py from generator
expressions). Verdict: WARN, not FAIL.
- All 1689 tests pass (was 1671; +18 for the new skill).
- mkdocs build --strict: succeeded in 12.47s.
## Verifiable success criteria (all green)
✓ scripts/*.py --help → exit 0 for all 3 scripts
✓ SKILL.md frontmatter → name + description + tags + compatible_tools
✓ plugin.json schema → 8 fields exact (verified)
✓ sync_skill_bundles → standalone ↔ bundled mirror in sync
✓ marketplace.json → standalone entry + bundle counts updated
✓ generate-docs.py → command page generated (skill page manual)
✓ mkdocs build --strict → succeeded
✓ cross-tool sync → codex + gemini synced
✓ pytest tests/ → 1689 passed, 0 failed
✓ CHANGELOG.md → [Unreleased] entry expanded for Phase 4
✓ Self-test → error_budget_calculator on 99.9% / 28d emits
correct burn-rate (14.4 fast, 6 slow, 1 ticket)
✓ Composition → references named skills explicitly compose
## Phase 1+2+3+4 cumulative
- 4 new skills: feature-flags-architect, kubernetes-operator,
chaos-engineering, slo-architect
- 12 new Python tools (all stdlib, all <250 LOC, average complexity 92/100)
- 16 new reference docs
- 4 new slash commands (/flag-cleanup, /operator-audit, /chaos-experiment,
/slo-design)
https://claude.ai/code/session_01Dq12xJakFRxwaoU8Pqejdm
Co-authored-by: Claude <noreply@anthropic.com>
5.1 KiB
Composition with the rest of the portfolio
slo-architect is the keystone. Three other skills in this library already lean on the SLO + error budget concept. This page shows how to wire them together for a coherent reliability stack.
The unified concept: error budget
┌────────────────────────────────────────────────────────────┐
│ slo-architect │
│ defines SLO, error budget, burn rate │
└──────────┬─────────────────┬────────────────┬─────────────┘
│ │ │
▼ ▼ ▼
feature-flags- chaos-engineering kubernetes-
architect (blast-radius operator
(rollout abort) bound by EB) (cap level L4)
With feature-flags-architect
feature-flags-architect defines kill switches. Their abort triggers should reference SLO burn-rate, not arbitrary thresholds.
Before:
abort_if: "p99 > 1000ms OR error_rate > 1%"
After (SLO-driven):
abort_if: "burn_rate.fast > 14.4 over 1h (per SLO checkout-success)"
Wire-up:
- Define SLO via
slo_designer.py - Run
error_budget_calculator.pyto get the burn-rate threshold - Use that threshold in the flag's abort criteria
- The kill_switch_audit.py from feature-flags-architect now has a real signal to verify against
With chaos-engineering
chaos-engineering's blast_radius_calculator.py already takes monthly error budget as input — but the budget should come from the SLO, not be made up.
# 1. Get the budget from the SLO definition
python slo_architect/scripts/error_budget_calculator.py \
--target 99.9 --window-days 30 --format json \
| jq .budget_minutes
# 2. Pass it to the chaos blast-radius calculator
python chaos_engineering/scripts/blast_radius_calculator.py \
--traffic-share 0.05 \
--user-pop 1000000 \
--duration-min 15 \
--monthly-budget-min 43.2 # ← from step 1
Now blast radius is bounded by REAL error budget, not a number someone typed in.
With kubernetes-operator
OperatorHub Capability Level 4 ("Deep Insights") requires:
/metricsendpoint- Prometheus alert rules
- SLOs documented for the operator's managed resources
slo-architect provides the SLO definitions; error_budget_calculator.py provides the alert rules. Drop them in the operator's Helm chart or OperatorHub bundle.
End-to-end example
Goal: ship a new checkout flow.
-
Define the SLO (slo-architect):
slo_designer.py --service checkout-svc --sli-type request-success-rate \ --target 99.9 --window-days 28 --owner team-checkout -
Compute burn-rate alerts (slo-architect):
error_budget_calculator.py --target 99.9 --window-days 28 # → fast_burn threshold = 14.4 -
Define rollout (feature-flags-architect):
rollout_planner.py --population 100000 --target-percent 100 \ --duration-days 14 --strategy ring # 1% → 5% → 25% → 50% → 100% -
Wire the abort (feature-flags-architect):
abort_if: "burn_rate.fast > 14.4 (per SLO slo-checkout-svc-...)" -
Validate via chaos before going wide (chaos-engineering):
blast_radius_calculator.py --traffic-share 0.05 --user-pop 100000 \ --duration-min 15 --monthly-budget-min 40.32 # → GREEN if <1% of monthly budget -
Audit the operator if the service is operator-managed (kubernetes-operator):
operator_capability_audit.py --operator-dir ./checkout-operator # → confirm L4 includes the new SLO
Each step uses the previous step's output as input. The SLO is the unifying number.
What slo-architect does NOT replace
- observability-designer — broader observability strategy (metrics, logs, traces, dashboards beyond SLO)
- incident-response — SLO violation may trigger an incident, but incident response is a separate discipline
- performance-profiler — capacity planning needs different metrics than SLO does
Use slo-architect for SLO+error-budget; use the others for their specific scopes.
Anti-pattern: SLO without composition
A team defines SLOs in a spreadsheet. Nobody references them in:
- Feature flag rollouts
- Chaos experiment design
- Operator capability audits
- Incident postmortems
The SLOs become a reporting artifact, not an operating tool. The composition story is what makes SLOs change behavior.
Operational checklist
For any service with a new SLO, verify:
- SLO defined via
slo_designer.py(slo_review.pypasses) - Burn-rate alerts deployed via
error_budget_calculator.pyoutput - If using feature flags: rollout abort references the SLO burn-rate threshold
- If running chaos: blast radius bounded by SLO error budget
- If operator-managed: operator audit confirms L4 includes the SLO
- Postmortem template (when SLO violated) includes "SLO revision needed?" question