claude-skills/markdown-html/skills/md-review/scripts/annotation_extractor.py
Claude 8f6734a205
feat(markdown-html): add md-review v2.10.2 — code-review markdown→2-col HTML
Adds the fourth skill to the markdown-html/ domain. The Tier-2 use case
from Shihipar's essay ("Code Review and PR Writeups"): a markdown PR
writeup with ```diff blocks and severity callouts becomes a single-file
2-column HTML review with top jump-nav, diff on the left, severity-tagged
annotation cards on the right, and a mandatory named reviewer footer.

Three stdlib tools pipeline together:

1. diff_parser.py — scans markdown for ```diff fenced blocks, parses each
   as a unified diff (--- a/file, +++ b/file, @@ -10,7 +10,8 @@,
   space/+/- body lines), assigns per-line numbers on both old (lo) and
   new (ln) sides, preserves the per-hunk @@ header context. Supports
   --infer-diff for unfenced blocks. Stdlib regex + state machine.

2. annotation_extractor.py — extracts severity callouts (GFM
   > [!BLOCKER] style) and inline markers (nit:, blocker:, etc.).
   Default convention BLOCKER/MAJOR/MINOR/NIT per Google's Code Review
   Developer Guide; overridable via --severity-convention. Attaches each
   annotation to the nearest preceding diff block by source-line index;
   unanchored annotations go to a "general comments" section. Also
   captures LGTM/approve markers separately as approvals.

3. review_html_renderer.py — emits single-file 2-col HTML. Top jump-nav
   lists every annotation with severity badge + 80-char preview + jump
   link + per-tier counts in heading. Each hunk-row is a CSS grid with
   diff on the left (per-line numbers, +/- marks, addition/deletion bg
   tints from --md-success/--md-warn via color-mix) and annotation cards
   on the right. WCAG-1.4.1-compliant severity badges (color + icon +
   aria-label + text — color is NEVER the sole signal); BLOCKER danger
   color computed by hue-rotating the design-system accent 120° toward
   red so it stays brand-coherent. Approval bar when LGTM markers
   present and no findings. Collapses to stacked on viewports < 900px.
   Mandatory --reviewer (refuses exit 3 otherwise — research-ops
   named-owner discipline). Refuses exit 4 if no hunks present (wrong
   skill → route to md-document). No Prism CDN (diff coloring conflicts
   with syntax highlighting).

Plus 3 references each citing 5-7 sources:
  - diff_rendering_canon.md — POSIX diff format + GitHub/GitLab UI +
    difftastic + SWE at Google ch. 9
  - severity_coding.md — WCAG 1.4.1 + Google review taxonomy + Don
    Norman Design of Everyday Things + NN/g color UX
  - pr_annotation_ux.md — convergent 2-col UX from GitHub / GitLab /
    Reviewable / CodeStream + SWE at Google + NN/g F-shape
1 template asset documenting the canonical 2-col review HTML shape.
/cs:md-review slash command with 4 pre-flight gates
(under-100-lines, no-onboarding, missing-reviewer, no-hunks) + pipeline
+ output digest.

Repo-level updates:
- markdown-html/.claude-plugin/plugin.json: skills array adds
  ./skills/md-review; version 2.10.1 → 2.10.2; description updated.
- .claude-plugin/marketplace.json: markdown-html-skills entry version
  and description; top-level counters 341 → 342 skills, 542 → 545 Python
  tools, 685 → 688 references, 88 → 89 slash commands; metadata.version
  2.10.1 → 2.10.2.
- Root CLAUDE.md: v2.10.2 release-notes block above v2.10.1.

Validation:
- check_plugin_json.py → OK
- sync-codex-skills.py --dry-run → 1 new symlink, documentation: 4 skills,
  total 343
- skill_description_validator.py → PASS (all 5 checks: present, length,
  third-person, trigger, action verb)
- skill_review_checklist_runner.py → 5/6 PASS (97 lines passes
  under-100-lines check; minor WARN on "user" vs "developer" terminology
  which are contextually distinct — converter operator vs review subject)
- All 3 tools pass --help and --sample
- Hard rules verified end-to-end:
    no --reviewer → exit 3 with refusal message
    no hunks → exit 4 with refusal message + md-document routing hint
    custom severity convention "critical,important,suggestion,nit" works
- Full pipeline on sample PR (2 diff blocks, 2 callouts) produces 11.3 KB
  single-file HTML with all 14 expected components (reviewer footer, PR
  title, aria-labels per WCAG 1.4.1, --md-danger computed color, modern
  color-mix tints, 2-col grid, 900px responsive collapse, both file
  paths, annotation cards, jump-nav, addition+deletion classes, Findings
  heading).

Coming in v2.10.3: md-slides — slide splitter + presenter-notes parser
+ arrow-key/space-bar nav + @media print for PDF export. Reuses
md-document's renderer scaffolding + design-system/scripts/config_loader.py.

https://claude.ai/code/session_01BK2KoQot1U7J5oSosrCQdc
2026-06-03 05:43:45 +00:00

208 lines
7.6 KiB
Python

#!/usr/bin/env python3
"""annotation_extractor.py - Extract severity-tagged review annotations from markdown.
Stdlib-only. Scans the same markdown source the diff_parser walks, finds
severity callouts and inline review markers, and attaches each annotation to
the nearest preceding diff block. The result is what the renderer puts in
the right margin of the 2-column layout.
Severity conventions accepted:
GFM callouts (preferred):
> [!BLOCKER] must-fix before merge
> [!MAJOR] strongly recommend addressing
> [!MINOR] worth fixing
> [!NIT] cosmetic / style preference
Inline markers (legacy, less structured):
blocker: <prose>
major: <prose>
minor: <prose>
nit: <prose>
LGTM (treated as APPROVAL marker; not severity-coded)
The --severity-convention flag accepts a custom 4-tier ordering, e.g.
"critical,important,suggestion,nit" — but defaults to the BLOCKER/MAJOR/
MINOR/NIT convention. The order matters: position 0 = most severe.
Attachment heuristic: each annotation attaches to the most recent diff block
that appeared above it in the markdown source (by source line number). If
no diff appears above, the annotation is "unanchored" and the renderer
shows it in a "general comments" section.
NO LLM CALLS. Pure regex + line-index attachment.
Usage:
python annotation_extractor.py --input review.md --diff-blocks hunks.json
python annotation_extractor.py --sample
"""
from __future__ import annotations
import argparse
import json
import re
import sys
from pathlib import Path
from typing import Any
DEFAULT_SEVERITY_CONVENTION = ["BLOCKER", "MAJOR", "MINOR", "NIT"]
CALLOUT_OPEN_RE = re.compile(r"^>\s*\[!([A-Z]+)\]\s*$")
BLOCKQUOTE_RE = re.compile(r"^>\s?(.*)$")
INLINE_MARKER_RE = re.compile(r"^(?P<sev>[A-Za-z]+):\s+(?P<body>.+)$")
APPROVAL_RE = re.compile(r"^(LGTM|👍|approved|approve)\s*$", re.IGNORECASE)
def extract_annotations(
text: str,
diff_blocks: dict[str, Any] | None,
severity_convention: list[str],
) -> dict[str, Any]:
"""Returns {annotations: [...], approvals: [...], summary: {...}}."""
lines = text.splitlines()
upper_severities = {s.upper() for s in severity_convention}
# Index of diff-block source lines, for attachment lookup
diff_source_lines: list[tuple[int, int]] = [] # (source_line, block_index)
if diff_blocks:
for b in diff_blocks.get("blocks", []):
diff_source_lines.append((b["source_line"], b["block_index"]))
def nearest_preceding_block(line_index: int) -> int | None:
best: int | None = None
for src_line, block_idx in diff_source_lines:
if src_line <= line_index:
best = block_idx
else:
break
return best
annotations: list[dict[str, Any]] = []
approvals: list[dict[str, Any]] = []
i = 0
while i < len(lines):
ln = lines[i]
# GFM callout (multi-line)
m_open = CALLOUT_OPEN_RE.match(ln)
if m_open:
severity_raw = m_open.group(1).upper()
body_lines: list[str] = []
j = i + 1
while j < len(lines):
bq = BLOCKQUOTE_RE.match(lines[j])
if not bq:
break
body_lines.append(bq.group(1))
j += 1
if severity_raw in upper_severities:
annotations.append({
"kind": "callout",
"severity": severity_raw,
"severity_rank": severity_convention.index(
next(s for s in severity_convention if s.upper() == severity_raw)
),
"body": " ".join(b.strip() for b in body_lines).strip(),
"source_line": i,
"attached_block": nearest_preceding_block(i),
})
i = j
continue
# Approval marker (LGTM, etc.)
if APPROVAL_RE.match(ln.strip()):
approvals.append({
"kind": "approval",
"marker": ln.strip(),
"source_line": i,
"attached_block": nearest_preceding_block(i),
})
i += 1
continue
# Inline marker (single-line, prose-leading)
m_inline = INLINE_MARKER_RE.match(ln.strip())
if m_inline:
sev_raw = m_inline.group("sev").upper()
if sev_raw in upper_severities:
annotations.append({
"kind": "inline",
"severity": sev_raw,
"severity_rank": severity_convention.index(
next(s for s in severity_convention if s.upper() == sev_raw)
),
"body": m_inline.group("body").strip(),
"source_line": i,
"attached_block": nearest_preceding_block(i),
})
i += 1
# Sort annotations primarily by source_line (preserves narrative order
# in the renderer's jump-nav), then group by severity in summary.
annotations.sort(key=lambda a: a["source_line"])
counts_by_severity: dict[str, int] = {}
for a in annotations:
counts_by_severity[a["severity"]] = counts_by_severity.get(a["severity"], 0) + 1
return {
"annotations": annotations,
"approvals": approvals,
"summary": {
"total_annotations": len(annotations),
"total_approvals": len(approvals),
"counts_by_severity": counts_by_severity,
"severity_convention": severity_convention,
},
}
def main(argv: list[str]) -> int:
p = argparse.ArgumentParser(description=__doc__.split("\n")[0])
p.add_argument("--input", help="Path to markdown file, or '-' for stdin")
p.add_argument("--diff-blocks", help="Path to diff_parser JSON output (for attachment)")
p.add_argument("--severity-convention",
default=",".join(DEFAULT_SEVERITY_CONVENTION),
help="Comma-separated severity tier list, most-to-least severe. "
"Default: BLOCKER,MAJOR,MINOR,NIT")
p.add_argument("--output", help="Path to write JSON output (else stdout)")
p.add_argument("--sample", action="store_true",
help="Run on a built-in sample PR review")
args = p.parse_args(argv)
severity_convention = [s.strip().upper() for s in args.severity_convention.split(",")]
if len(severity_convention) < 2:
print("error: --severity-convention needs at least 2 tiers", file=sys.stderr)
return 2
if args.sample:
sys.path.insert(0, str(Path(__file__).resolve().parent))
import diff_parser
text = diff_parser.SAMPLE_MARKDOWN
diff_blocks = diff_parser.parse_markdown_for_diffs(text)
elif args.input:
text = sys.stdin.read() if args.input == "-" else Path(args.input).read_text(encoding="utf-8")
diff_blocks = (
json.loads(Path(args.diff_blocks).read_text(encoding="utf-8"))
if args.diff_blocks else None
)
else:
p.print_help()
return 0
result = extract_annotations(text, diff_blocks, severity_convention)
payload = json.dumps(result, indent=2)
if args.output:
Path(args.output).write_text(payload, encoding="utf-8")
print(f"wrote {args.output}: {result['summary']['total_annotations']} annotations, "
f"{result['summary']['total_approvals']} approvals, "
f"counts={result['summary']['counts_by_severity']}")
else:
print(payload)
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))