claude-skills/engineering/skills/slo-architect/references/composition.md
Alireza Rezvani 9dd6fd184c
feat(slo-architect): Phase 4 — SLO/SLI/error-budget discipline (#605)
Phase 4 of the multi-skill build effort. Same 14-step pipeline.

## What landed

### New skill: engineering/slo-architect

End-to-end SLO discipline per Google SRE Workbook. Published as BOTH:
- Standalone plugin: engineering/slo-architect/
- Bundled mirror:    engineering/skills/slo-architect/

3 stdlib-only Python tools (Karpathy complexity 95/100):
- slo_designer.py             — generates SLO definitions; refuses to render
                                 if required fields missing (owner, policy doc,
                                 SLI numerator/denominator). Supports 5 SLI
                                 types: request-success-rate, request-latency,
                                 availability-time, data-freshness, correctness.
- error_budget_calculator.py  — computes error budget AND the canonical
                                 multi-window burn-rate alert thresholds:
                                 fast (1h/5m, page), slow (6h/30m, page),
                                 ticket (3d/6h). Output is PromQL-shaped,
                                 ready to paste into Prometheus rules.
- slo_review.py               — audits SLO docs for 7 common bugs:
                                 target ≥99.99, target ≤99, window <7d,
                                 window >90d, no SLI definition, no error
                                 budget policy, CPU-as-SLI.

4 reference docs:
- slo_principles.md   — SLI vs SLO vs SLA, Google SRE Workbook canon
- sli_design.md       — 5 SLI types with examples and anti-patterns
- error_budget.md     — error budget math, burn-rate alerts, budget policy
- composition.md      — how SLOs feed feature-flags, chaos, kubernetes-operator

Asset templates:
- slo_template.yaml          — fillable SLO YAML with all required fields
- error_budget_policy.md     — fillable 4-state policy (HEALTHY / CAUTION /
                                CRITICAL / VIOLATED)

Plus: SKILL.md, README.md, /slo-design slash command.

## Composition with prior phases

Explicit wire-up to the rest of the portfolio:
- feature-flags-architect.kill_switch_audit references SLO burn-rate
- chaos-engineering.blast_radius_calculator takes SLO error budget as input
- kubernetes-operator capability level L4 requires SLOs + Prometheus rules

The SLO is the unifying number: rollout abort, chaos blast radius, and
operator capability all reference it. references/composition.md walks
through end-to-end use.

## Audit verdict (evidence-based)

Closest existing skill: engineering/observability-designer covers SLI/SLO as
ONE topic among many (metrics, logs, traces, dashboards, alerting). It has
no dedicated tools and is breadth-not-depth. slo-architect is the focused
SLO discipline with deterministic Python tools — same gap pattern as
kubernetes-operator vs senior-devops.

## Marketplace / registry

- marketplace.json: slo-architect registered as standalone plugin
- engineering-advanced-skills bundle: 49 → 50 skills, version → 2.4.4
- engineering/.claude-plugin/plugin.json: version + skill list updated
- mkdocs.yml: nav entry under "Engineering - POWERFUL"
- docs/skills/engineering/slo-architect.md: docs page (manual)
- docs/commands/slo-design.md: auto-generated
- .codex/, .gemini/: synced

## Karpathy-coder gates

- complexity_checker (strict): 95/100 average — same top score as
  chaos-engineering. 1 WARN (depth 7 in slo_review.py from generator
  expressions). Verdict: WARN, not FAIL.
- All 1689 tests pass (was 1671; +18 for the new skill).
- mkdocs build --strict: succeeded in 12.47s.

## Verifiable success criteria (all green)

✓  scripts/*.py --help     → exit 0 for all 3 scripts
✓  SKILL.md frontmatter    → name + description + tags + compatible_tools
✓  plugin.json schema      → 8 fields exact (verified)
✓  sync_skill_bundles      → standalone ↔ bundled mirror in sync
✓  marketplace.json        → standalone entry + bundle counts updated
✓  generate-docs.py        → command page generated (skill page manual)
✓  mkdocs build --strict   → succeeded
✓  cross-tool sync         → codex + gemini synced
✓  pytest tests/           → 1689 passed, 0 failed
✓  CHANGELOG.md            → [Unreleased] entry expanded for Phase 4
✓  Self-test               → error_budget_calculator on 99.9% / 28d emits
                             correct burn-rate (14.4 fast, 6 slow, 1 ticket)
✓  Composition             → references named skills explicitly compose

## Phase 1+2+3+4 cumulative

- 4 new skills: feature-flags-architect, kubernetes-operator,
                chaos-engineering, slo-architect
- 12 new Python tools (all stdlib, all <250 LOC, average complexity 92/100)
- 16 new reference docs
- 4 new slash commands (/flag-cleanup, /operator-audit, /chaos-experiment,
                        /slo-design)

https://claude.ai/code/session_01Dq12xJakFRxwaoU8Pqejdm

Co-authored-by: Claude <noreply@anthropic.com>
2026-05-10 07:39:05 +02:00

5.1 KiB

Composition with the rest of the portfolio

slo-architect is the keystone. Three other skills in this library already lean on the SLO + error budget concept. This page shows how to wire them together for a coherent reliability stack.

The unified concept: error budget

┌────────────────────────────────────────────────────────────┐
│                      slo-architect                         │
│             defines SLO, error budget, burn rate           │
└──────────┬─────────────────┬────────────────┬─────────────┘
           │                 │                │
           ▼                 ▼                ▼
   feature-flags-      chaos-engineering   kubernetes-
   architect           (blast-radius        operator
   (rollout abort)     bound by EB)         (cap level L4)

With feature-flags-architect

feature-flags-architect defines kill switches. Their abort triggers should reference SLO burn-rate, not arbitrary thresholds.

Before:

abort_if: "p99 > 1000ms OR error_rate > 1%"

After (SLO-driven):

abort_if: "burn_rate.fast > 14.4 over 1h (per SLO checkout-success)"

Wire-up:

  1. Define SLO via slo_designer.py
  2. Run error_budget_calculator.py to get the burn-rate threshold
  3. Use that threshold in the flag's abort criteria
  4. The kill_switch_audit.py from feature-flags-architect now has a real signal to verify against

With chaos-engineering

chaos-engineering's blast_radius_calculator.py already takes monthly error budget as input — but the budget should come from the SLO, not be made up.

# 1. Get the budget from the SLO definition
python slo_architect/scripts/error_budget_calculator.py \
  --target 99.9 --window-days 30 --format json \
  | jq .budget_minutes

# 2. Pass it to the chaos blast-radius calculator
python chaos_engineering/scripts/blast_radius_calculator.py \
  --traffic-share 0.05 \
  --user-pop 1000000 \
  --duration-min 15 \
  --monthly-budget-min 43.2  # ← from step 1

Now blast radius is bounded by REAL error budget, not a number someone typed in.

With kubernetes-operator

OperatorHub Capability Level 4 ("Deep Insights") requires:

  • /metrics endpoint
  • Prometheus alert rules
  • SLOs documented for the operator's managed resources

slo-architect provides the SLO definitions; error_budget_calculator.py provides the alert rules. Drop them in the operator's Helm chart or OperatorHub bundle.

End-to-end example

Goal: ship a new checkout flow.

  1. Define the SLO (slo-architect):

    slo_designer.py --service checkout-svc --sli-type request-success-rate \
      --target 99.9 --window-days 28 --owner team-checkout
    
  2. Compute burn-rate alerts (slo-architect):

    error_budget_calculator.py --target 99.9 --window-days 28
    # → fast_burn threshold = 14.4
    
  3. Define rollout (feature-flags-architect):

    rollout_planner.py --population 100000 --target-percent 100 \
      --duration-days 14 --strategy ring
    # 1% → 5% → 25% → 50% → 100%
    
  4. Wire the abort (feature-flags-architect):

    abort_if: "burn_rate.fast > 14.4 (per SLO slo-checkout-svc-...)"
    
  5. Validate via chaos before going wide (chaos-engineering):

    blast_radius_calculator.py --traffic-share 0.05 --user-pop 100000 \
      --duration-min 15 --monthly-budget-min 40.32
    # → GREEN if <1% of monthly budget
    
  6. Audit the operator if the service is operator-managed (kubernetes-operator):

    operator_capability_audit.py --operator-dir ./checkout-operator
    # → confirm L4 includes the new SLO
    

Each step uses the previous step's output as input. The SLO is the unifying number.

What slo-architect does NOT replace

  • observability-designer — broader observability strategy (metrics, logs, traces, dashboards beyond SLO)
  • incident-response — SLO violation may trigger an incident, but incident response is a separate discipline
  • performance-profiler — capacity planning needs different metrics than SLO does

Use slo-architect for SLO+error-budget; use the others for their specific scopes.

Anti-pattern: SLO without composition

A team defines SLOs in a spreadsheet. Nobody references them in:

  • Feature flag rollouts
  • Chaos experiment design
  • Operator capability audits
  • Incident postmortems

The SLOs become a reporting artifact, not an operating tool. The composition story is what makes SLOs change behavior.

Operational checklist

For any service with a new SLO, verify:

  • SLO defined via slo_designer.py (slo_review.py passes)
  • Burn-rate alerts deployed via error_budget_calculator.py output
  • If using feature flags: rollout abort references the SLO burn-rate threshold
  • If running chaos: blast radius bounded by SLO error budget
  • If operator-managed: operator audit confirms L4 includes the SLO
  • Postmortem template (when SLO violated) includes "SLO revision needed?" question