claude-skills/markdown-html/skills/md-review/references/severity_coding.md
Claude 8f6734a205
feat(markdown-html): add md-review v2.10.2 — code-review markdown→2-col HTML
Adds the fourth skill to the markdown-html/ domain. The Tier-2 use case
from Shihipar's essay ("Code Review and PR Writeups"): a markdown PR
writeup with ```diff blocks and severity callouts becomes a single-file
2-column HTML review with top jump-nav, diff on the left, severity-tagged
annotation cards on the right, and a mandatory named reviewer footer.

Three stdlib tools pipeline together:

1. diff_parser.py — scans markdown for ```diff fenced blocks, parses each
   as a unified diff (--- a/file, +++ b/file, @@ -10,7 +10,8 @@,
   space/+/- body lines), assigns per-line numbers on both old (lo) and
   new (ln) sides, preserves the per-hunk @@ header context. Supports
   --infer-diff for unfenced blocks. Stdlib regex + state machine.

2. annotation_extractor.py — extracts severity callouts (GFM
   > [!BLOCKER] style) and inline markers (nit:, blocker:, etc.).
   Default convention BLOCKER/MAJOR/MINOR/NIT per Google's Code Review
   Developer Guide; overridable via --severity-convention. Attaches each
   annotation to the nearest preceding diff block by source-line index;
   unanchored annotations go to a "general comments" section. Also
   captures LGTM/approve markers separately as approvals.

3. review_html_renderer.py — emits single-file 2-col HTML. Top jump-nav
   lists every annotation with severity badge + 80-char preview + jump
   link + per-tier counts in heading. Each hunk-row is a CSS grid with
   diff on the left (per-line numbers, +/- marks, addition/deletion bg
   tints from --md-success/--md-warn via color-mix) and annotation cards
   on the right. WCAG-1.4.1-compliant severity badges (color + icon +
   aria-label + text — color is NEVER the sole signal); BLOCKER danger
   color computed by hue-rotating the design-system accent 120° toward
   red so it stays brand-coherent. Approval bar when LGTM markers
   present and no findings. Collapses to stacked on viewports < 900px.
   Mandatory --reviewer (refuses exit 3 otherwise — research-ops
   named-owner discipline). Refuses exit 4 if no hunks present (wrong
   skill → route to md-document). No Prism CDN (diff coloring conflicts
   with syntax highlighting).

Plus 3 references each citing 5-7 sources:
  - diff_rendering_canon.md — POSIX diff format + GitHub/GitLab UI +
    difftastic + SWE at Google ch. 9
  - severity_coding.md — WCAG 1.4.1 + Google review taxonomy + Don
    Norman Design of Everyday Things + NN/g color UX
  - pr_annotation_ux.md — convergent 2-col UX from GitHub / GitLab /
    Reviewable / CodeStream + SWE at Google + NN/g F-shape
1 template asset documenting the canonical 2-col review HTML shape.
/cs:md-review slash command with 4 pre-flight gates
(under-100-lines, no-onboarding, missing-reviewer, no-hunks) + pipeline
+ output digest.

Repo-level updates:
- markdown-html/.claude-plugin/plugin.json: skills array adds
  ./skills/md-review; version 2.10.1 → 2.10.2; description updated.
- .claude-plugin/marketplace.json: markdown-html-skills entry version
  and description; top-level counters 341 → 342 skills, 542 → 545 Python
  tools, 685 → 688 references, 88 → 89 slash commands; metadata.version
  2.10.1 → 2.10.2.
- Root CLAUDE.md: v2.10.2 release-notes block above v2.10.1.

Validation:
- check_plugin_json.py → OK
- sync-codex-skills.py --dry-run → 1 new symlink, documentation: 4 skills,
  total 343
- skill_description_validator.py → PASS (all 5 checks: present, length,
  third-person, trigger, action verb)
- skill_review_checklist_runner.py → 5/6 PASS (97 lines passes
  under-100-lines check; minor WARN on "user" vs "developer" terminology
  which are contextually distinct — converter operator vs review subject)
- All 3 tools pass --help and --sample
- Hard rules verified end-to-end:
    no --reviewer → exit 3 with refusal message
    no hunks → exit 4 with refusal message + md-document routing hint
    custom severity convention "critical,important,suggestion,nit" works
- Full pipeline on sample PR (2 diff blocks, 2 callouts) produces 11.3 KB
  single-file HTML with all 14 expected components (reviewer footer, PR
  title, aria-labels per WCAG 1.4.1, --md-danger computed color, modern
  color-mix tints, 2-col grid, 900px responsive collapse, both file
  paths, annotation cards, jump-nav, addition+deletion classes, Findings
  heading).

Coming in v2.10.3: md-slides — slide splitter + presenter-notes parser
+ arrow-key/space-bar nav + @media print for PDF export. Reuses
md-document's renderer scaffolding + design-system/scripts/config_loader.py.

https://claude.ai/code/session_01BK2KoQot1U7J5oSosrCQdc
2026-06-03 05:43:45 +00:00

4.9 KiB
Raw Permalink Blame History

Severity Coding

Why this exists: Code-review annotations need a severity dimension or every comment carries the same weight. This skill ships a default 4-tier convention (BLOCKER / MAJOR / MINOR / NIT), accepts custom conventions via --severity-convention, and enforces WCAG 1.4.1 (color is not the sole signal — every badge has color + icon + aria-label).

The default convention

Tier Meaning Source Visual
BLOCKER Must fix before merge. Author cannot proceed without addressing. Found in many engineering teams' written conventions; Google calls this "must-fix" Filled square ■ + derived danger color (accent rotated 120° toward red)
MAJOR Strongly recommended. Author should address unless they have a good counter-argument. Common in Google's Code Review Developer Guide as the "non-nit substantive comment" Filled triangle ▲ + --md-warn
MINOR Worth fixing. Reasonable to address now; reasonable to defer. Common middle tier Filled circle ● + --md-link
NIT Cosmetic / style preference. Author may address or ignore. Google's nit: prefix is the canonical example Open circle ◦ + --md-text-muted

Position 0 (BLOCKER) is most severe; position 3 (NIT) is least. Position determines the severity_rank field in the JSON output, which the renderer uses for sort + display order.

Custom conventions

--severity-convention "critical,important,suggestion,nit" swaps the tier list. Same rank ordering applies (position 0 = most severe). Renderer uses the same icon/color mapping when the tier name matches a default; otherwise falls back to the _text-muted token + open-circle icon for the unknown tier.

WCAG 1.4.1 — color is not the sole signal

Every severity badge ships:

  1. Color — derived from the design-system palette, validated for AA contrast against bg.
  2. Icon — ■ / ▲ / ● / ◦ — a glyph that's distinct shape-wise. (Deuteranopic readers see the shape difference even if the color difference is muted.)
  3. aria-label — full text spelled out: "Blocker — must fix before merge". Screen readers announce this; sighted readers don't see it but get the icon + color + text.
  4. Text label — the severity name itself ("BLOCKER") is in the badge. Triple-redundancy: color, icon, text.

A reader who is completely color-blind, or viewing on a grayscale monitor, or screen-reading the page, still has full access to severity.

What we deliberately don't do

  • Emoji icons — render inconsistently across OS / font (Windows vs macOS vs Linux vs iOS), break in monochrome printing. We use Unicode geometric shapes (■▲●◦) that ship with every system font.
  • Color-only badges — the WCAG floor.
  • Severity arithmetic — no "this PR has score X = blockers × 10 + majors × 3 + …" auto-rollup. Review quality is qualitative; numeric rollups create gaming incentives.
  • Auto-block merge based on severity — that's a CI concern, not a renderer concern.

Sources

1. WCAG 2.2 §1.4.1 Use of Color (w3.org/WAI/WCAG22)

The hard rule: color must not be the only visual means of conveying information. Our color + icon + label + aria-label is the canonical implementation.

2. Google Code Review Developer Guide (google.github.io/eng-practices/review/reviewer)

Source of the nit: prefix convention and the "blocker / major / minor / nit" taxonomy. We adopt it as the default.

3. Software Engineering at Google (Manshreck & Wright, O'Reilly 2020), Ch. 9

The taxonomy is documented here: blockers are tickets that fail review; nits are "I'd prefer X but it's not a blocker"; majors are substantive comments.

4. Phabricator / Sourcegraph / Reviewable.io

Each tool ships its own severity vocabulary; ours is compatible by adopting the most-common 4-tier convention.

5. Don Norman — The Design of Everyday Things (2013 ed., Basic Books)

The "signifier" concept: visual elements must communicate function. A red badge is a signifier; a red badge that says "BLOCKER" with a filled square is a stronger signifier; a red badge with aria-label="Blocker — must fix before merge" adds the non-visual channel.

6. NN/g — Color in UX (Therese Fessenden, 2024)

Empirical: relying on color alone fails for 8% of male readers (red-green color blindness). The triple-redundancy approach is standard.

7. Anil Dash — The Web We Lost (dashes.com, 2012)

Indirectly: portable web artifacts. A single-file review must render correctly even when CSS variables fail to load — which is why the badge text label (in addition to color + icon) is essential. If var(--md-warn) doesn't resolve, the user still sees "MAJOR" with a triangle.

Applied to md-review

The _render_severity_badge helper in review_html_renderer.py emits the color + icon + label + aria-label combo for every severity, derived from the design-system palette. The convention is overridable via --severity-convention. Refuses to render a badge with color alone.