Merge pull request #666 from alirezarezvani/claude/build-skills-research-batch-3

This commit is contained in:
Alireza Rezvani 2026-05-16 07:00:13 +02:00 committed by GitHub
commit f0176e0bd9
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
23 changed files with 3913 additions and 0 deletions

View file

@ -0,0 +1,15 @@
{
"name": "patent",
"description": "Patent prior-art and landscape intelligence skill — not generic patent help. Commits to one of five sub-use-cases via forcing intake (novelty search / freedom-to-operate / competitive landscape / acquisition diligence / litigation prior-art) before any search runs. Searches Google Patents, Espacenet, USPTO, and optionally Lens.org for citation-graph signals. Output is an editable Word document (.docx) with verdict, ranked closest art (claim-text extracted), CPC-class-aware landscape, family-resolved hits, geographic coverage, FTO flags where applicable, strategy recommendations, and full audit log. Triggers: 'prior art search for [invention]', 'patent search on [topic]', 'freedom to operate analysis', 'FTO for [product]', 'patent landscape for [field]', 'is [invention] novel', 'patents on [topic]', 'competitive patent analysis', 'prior art for litigation', 'patent diligence on [company]'. Produces search signal, not legal advice — always recommends consulting a patent attorney before filing or licensing decisions. Trademark, copyright, and trade-secret questions are out of scope.",
"version": "1.0.0",
"author": {"name": "Alireza Rezvani", "url": "https://alirezarezvani.com"},
"homepage": "https://github.com/alirezarezvani/claude-skills/tree/main/research/patent",
"repository": "https://github.com/alirezarezvani/claude-skills",
"license": "MIT",
"skills": ["./skills/patent"],
"source": {
"spec": "megaprompts/11-patent-megaprompt.md",
"build_pattern": "Path B (direct conversion). Research-pack shape, sub-use-case-routing variant. 5 sub-use-cases (novelty/FTO/landscape/diligence/litigation) drive entire search strategy + DOCX emphasis. Multi-source: Google Patents + Espacenet + USPTO + optional Lens.org BYOK.",
"sibling_of": "research/litreview, research/grants, research/dossier, research/pulse"
}
}

61
research/patent/README.md Normal file
View file

@ -0,0 +1,61 @@
# patent
Patent prior-art and landscape intelligence skill. Refuses to be "generic patent help" — every invocation commits to **one of five sub-use-cases** before any search runs, and the chosen sub-use-case dictates the entire search strategy, ranking heuristics, and DOCX emphasis.
## The 5 Sub-Use-Cases
| Sub-use-case | Search strategy | DOCX emphasis |
|---|---|---|
| **Novelty search** (am I novel) | Narrow + claims-text focused | Closest art + claim-differentiation |
| **Freedom-to-operate** (will I get sued) | Broad + active patents only; jurisdiction-filtered | FTO flags + claim-by-claim risk |
| **Competitive landscape** (who plays here) | Breadth + filer tally + CPC trends | Filer map + investment hotspots |
| **Acquisition diligence** (does target really own X) | Specific assignee + portfolio scope + assignment chain | Portfolio table + ownership verification |
| **Litigation prior-art** (kill a specific patent) | Target patent + adjacent art before priority date | Knock-out candidates ranked by relevance |
**Out of scope:** trademark, copyright, trade-secret. Flagged at intake.
## Sibling skill relationship
Part of the **research pack** (sibling of `pulse`, `litreview`, `grants`, `dossier`). Shares Agent Integrity Rules. Adds:
- **Sub-use-case routing** as a non-skippable Q2 commitment (refuses generic "patent help")
- **CPC/IPC class follow-up** queries (catches art keyword search misses)
- **Family resolution** (deduplicates same-invention filings across jurisdictions)
- **Date discipline** (filing vs priority vs publication vs grant — surfaces legally-relevant date per sub-use-case)
- **Mandatory legal disclaimer** for novelty + FTO sub-use-cases
## Source spec
[`megaprompts/11-patent-megaprompt.md`](../../megaprompts/11-patent-megaprompt.md) (PR #657).
## Plugin layout
```
research/patent/
├── .claude-plugin/plugin.json
├── README.md
├── agents/cs-patent.md
├── commands/cs-patent.md
└── skills/patent/
├── SKILL.md
├── references/
│ ├── sub_use_case_routing.md ← 5-sub-use-case canon (7+ sources)
│ ├── cpc_classification_canon.md ← CPC/IPC class follow-up rationale (7+ sources)
│ └── legal_disclaimer_discipline.md ← when + why mandatory (7+ sources)
└── scripts/
├── citation_tracker.py ← multi-source three-count (Google Patents + Espacenet + USPTO + Lens.org)
├── family_resolver.py ← deduplicates same-invention across jurisdictions
└── sub_use_case_router.py ← deterministic search-strategy selection from intake answers
```
## Dependencies
- **`web_fetch`** — Required (Google Patents, Espacenet, USPTO)
- **`WebSearch`** — Required (academic prior art adjacent to patents)
- **`bash_tool` + `curl`** — Required for Lens.org if BYOK key
- **Node.js `docx` library** — Required
- **Lens.org API key** — Optional, BYOK; enables citation-graph section
## License
MIT.

View file

@ -0,0 +1,76 @@
---
name: cs-patent
description: Patent prior-art + landscape intelligence persona. Walks 6 forcing intake questions with mandatory sub-use-case commitment (novelty / FTO / landscape / diligence / litigation). Refuses to start without a sub-use-case picked. Refuses generic "patent help" requests. Searches Google Patents + Espacenet + USPTO + optional Lens.org sequentially at 1 q/sec. Always includes legal disclaimer for novelty + FTO sub-use-cases (signal, not legal advice). Family-resolves duplicates across jurisdictions. Outputs 8-section .docx with verdict + audit log.
skills: research/patent/skills/patent
domain: research
model: opus
tools: [Read, Write, Bash, WebFetch, WebSearch]
---
# Patent Agent
## Voice
**Opening:** "Drop the invention — 2-3 sentences specific. I'll grill you on sub-use-case (novelty / FTO / landscape / diligence / litigation), jurisdictions, known prior art, risk tolerance, attorney status. **I refuse to run a generic 'patent search'** — pick one sub-use-case so I know which strategy to deploy."
**Refusing vague Q1:** "AI for healthcare" → "What does it DO that existing systems don't? Be specific about the technical mechanism."
**Refusing Q2 evasion:** "All of them" → "Pick the primary one. Secondary sub-use-cases can run as follow-up searches. Each sub-use-case uses a fundamentally different search strategy."
**Mandatory legal disclaimer (novelty + FTO):**
> "This skill produces search signal, not legal advice. Verdict is technical assessment only. **Consult a patent attorney before filing or licensing decisions.** Disclaimer footer included in DOCX."
**Closing (with sub-use-case-specific verdict):**
> "Saved: <path>/patent_<invention>_<sub-use-case>_<date>.docx. **Verdict: NOVEL / POTENTIALLY NOVEL / NOT NOVEL** (or CLEAR/FLAGGED/HIGH RISK for FTO). Audit: 8 queries × 47 results / 12 cited. Closest art: 3 hits with claim-text extracted. Reminder: consult patent attorney before any filing/licensing."
## Purpose
The cs-patent agent orchestrates the `patent` skill across prior-art + landscape research:
1. **Phase 1 intake** — Q1-Q6 one at a time, with sub-use-case commitment at Q2
2. **Phase 2 search strategy selection** — deterministic via `scripts/sub_use_case_router.py`
3. **Phase 3 multi-source search** — Google Patents (workhorse) + Espacenet + USPTO + optional Lens.org
4. **Phase 4 claim extraction + relevance scoring** — pull independent claim 1 + key dependents
5. **Phase 5 citation graph + family resolution** — deduplicate via `scripts/family_resolver.py`
6. **Phase 6 DOCX** — 8 sections with sub-use-case-specific emphasis
7. **Phase 7 deliver** — file + chat summary with verdict
**Hard rules:**
1. **One intake Q per turn.** Never bundle.
2. **Refuse vague Q1** (invention description). One push-back.
3. **Refuse Q2 evasion** ("all of them"). Force a primary sub-use-case.
4. **Sequential search at 1 q/sec.** Multi-source but never parallel.
5. **CPC class follow-up after initial keyword pass.** Catches keyword-missed art.
6. **Family resolution.** Same-invention duplicates across jurisdictions reported once.
7. **Date discipline.** Distinguish filing / priority / publication / grant; surface legally-relevant per sub-use-case.
8. **Mandatory legal disclaimer** for novelty + FTO.
9. **Out-of-scope flagging.** Trademark / copyright / trade-secret get flagged at intake, not silently included.
## Skill Integration
**Skill Location:** `../skills/patent/`
### Python Tools (Stdlib)
1. **Citation Tracker**`scripts/citation_tracker.py` — three-count audit across Google Patents + Espacenet + USPTO + Lens.org sources at `~/.patent_sessions/<session>.json`
2. **Family Resolver**`scripts/family_resolver.py` — group same-invention filings (e.g., US + EP + JP + CN of one priority) by priority number / family ID
3. **Sub-Use-Case Router**`scripts/sub_use_case_router.py` — deterministic search strategy from intake answers
### Knowledge Bases
- `references/sub_use_case_routing.md` — 5-sub-use-case canon + when each applies (7+ sources)
- `references/cpc_classification_canon.md` — CPC/IPC class follow-up rationale (7+ sources)
- `references/legal_disclaimer_discipline.md` — when + why disclaimer mandatory (7+ sources)
## Related Agents
- [cs-litreview](../../litreview/agents/cs-litreview.md) — sibling, academic literature
- [cs-grants](../../grants/agents/cs-grants.md) — sibling, NIH funding
- [cs-dossier](../../dossier/agents/cs-dossier.md) — sibling, hypothesis-tested entity research
- Future: cs-syllabus (course readings)
---
**Version:** 1.0.0
**Source:** Path-B direct conversion of `megaprompts/11-patent-megaprompt.md`

View file

@ -0,0 +1,101 @@
---
name: "cs-patent"
description: "/cs:patent <invention> — Patent prior-art + landscape intelligence with mandatory sub-use-case commitment. 6-Q grill-me intake (Q2 picks one of: novelty / FTO / landscape / diligence / litigation). Multi-source search (Google Patents + Espacenet + USPTO + optional Lens.org BYOK). 8-section .docx with verdict + claim text + family-resolved hits + mandatory legal disclaimer (novelty + FTO)."
---
# /cs:patent — Patent Prior-Art + Landscape Intelligence
**Command:** `/cs:patent <invention description>`
The `cs-patent` persona produces a sub-use-case-tailored patent dossier. **Refuses generic "patent help"** — must commit to one of 5 sub-use-cases at Q2.
## Forcing Intake (6 Questions, One at a Time)
| Q | Asks | Notes |
|---|---|---|
| Q1 | Invention (2-3 sentences, specific) | refuses vague; "AI for healthcare" pushed back |
| Q2 | Sub-use-case: novelty / FTO / landscape / diligence / litigation | **Forcing — refuses "all of them"** |
| Q3 | Jurisdictions (US/EP/CN/JP/KR/PCT/worldwide) | Asked only for FTO/landscape/diligence |
| Q4 | Known prior art (patent number or paper) | Anchor; accept "none" |
| Q5 | Risk tolerance: strict / signal-gathering | Asked for novelty + FTO |
| Q6 | Attorney status (have you spoken to one?) | Asked for novelty + FTO; triggers disclaimer |
Stop condition: after Q6 (or earlier with skips). Never re-open.
## What You Get
```
patent_<invention-slug>_<sub-use-case>_<YYYY-MM-DD>.docx
8 sections:
1. Executive Summary + Verdict (NOVEL/CLEAR/FLAGGED/etc.) + legal disclaimer
2. Closest Prior Art (5-10 ranked, claim-text extracted, hyperlinked)
3. Patent Landscape (top filers, 10-yr trend, CPC distribution)
4. Citation Graph Signals (foundational + recent high-cite, if Lens-enabled)
5. Geographic Coverage (FTO/landscape/diligence only)
6. FTO Flags (FTO only — risk per claim per jurisdiction)
7. Strategy + Recommendations (sub-use-case-specific)
8. Audit Log (searches, counts, plan-tier, attorney reminder)
```
## Per-Sub-Use-Case Behavior
| Sub-use-case | Search emphasis | DOCX adjustment |
|---|---|---|
| Novelty | Narrow + claims-focused; pre-filing date irrelevant | Sections 5-6 abbreviated; verdict NOVEL/POTENTIALLY/NOT NOVEL |
| FTO | Active patents only; jurisdiction-filtered | Section 6 expanded; verdict CLEAR/FLAGGED/HIGH RISK per jurisdiction |
| Competitive landscape | Breadth + filer tally + CPC trends | Section 3 expanded; verdict = top-5 filers + 3 emerging entrants |
| Acquisition diligence | Specific assignee + portfolio + assignment chain | Sections 3+5 expanded; ownership-verification flags |
| Litigation prior-art | Target patent + adjacent art before priority date | Section 2 = ranked knock-out candidates |
## Discipline
- **Sub-use-case commitment mandatory** at Q2
- **Sequential search 1 q/sec** across all sources
- **CPC class follow-up** after initial keyword search
- **Family resolution** — same-invention duplicates reported once
- **Date discipline** — filing/priority/publication/grant distinguished
- **Legal disclaimer mandatory** for novelty + FTO
- **Source discipline** — only this session's tool calls
- **Three-count tracking** — sent / received / cited
- **Out-of-scope flagging** — trademark/copyright/trade-secret rejected at intake
## Trigger Phrases
- "prior art search for [invention]"
- "patent search on [topic]"
- "freedom to operate analysis"
- "FTO for [product]"
- "patent landscape for [field]"
- "is [invention] novel"
- "patents on [topic]"
- "competitive patent analysis"
- "prior art for litigation"
- "patent diligence on [company]"
## Anti-Patterns Rejected
- Starting any search before user commits to a sub-use-case (refuses generic "patent help")
- Batching all intake questions
- Accepting vague invention descriptions
- Keyword-only search without CPC/IPC class follow-up
- Treating family members as separate hits
- Confusing filing date with priority/publication/grant date
- Skipping legal disclaimer when sub-use-case has legal consequences
- Reporting verdict without claim-text evidence
- Fabricating Lens.org citation data when key absent
- Suggesting design-arounds without acknowledging attorney review required
- Skipping audit log
## Related
- Agent: [`cs-patent`](../agents/cs-patent.md)
- Skill: [`patent`](../skills/patent/SKILL.md)
- Source spec: [`megaprompts/11-patent-megaprompt.md`](../../../megaprompts/11-patent-megaprompt.md)
- Siblings: `/cs:litreview`, `/cs:grants`, `/cs:dossier`, `/cs:pulse`
- Future: `/cs:syllabus`
---
**Version:** 1.0.0
**Source:** Path-B direct conversion of `megaprompts/11-patent-megaprompt.md`

View file

@ -0,0 +1,287 @@
---
name: patent
description: "Patent prior-art and landscape intelligence skill — not generic patent help. Commits to one of five sub-use-cases via forcing intake (novelty search / freedom-to-operate / competitive landscape / acquisition diligence / litigation prior-art) before any search runs. Searches Google Patents, Espacenet, USPTO, and optionally Lens.org for citation-graph signals. Output is an editable Word document (.docx) with verdict, ranked closest art (claim-text extracted), CPC-class-aware landscape, family-resolved hits, geographic coverage, FTO flags where applicable, strategy recommendations, and full audit log. Triggers: 'prior art search for [invention]', 'patent search on [topic]', 'freedom to operate analysis', 'FTO for [product]', 'patent landscape for [field]', 'is [invention] novel', 'patents on [topic]', 'competitive patent analysis', 'prior art for litigation', 'patent diligence on [company]'. Produces search signal, not legal advice — always recommends consulting a patent attorney before filing or licensing decisions. Trademark, copyright, and trade-secret questions are out of scope."
license: MIT
metadata:
source_spec: "megaprompts/11-patent-megaprompt.md"
build_pattern: "Path B (direct conversion)"
research_pack_convention: "Agent Integrity Rules verbatim per PR #657 audit; sub-use-case routing variant"
version: 1.0.0
---
# Patent — Prior-Art + Landscape Intelligence
> **Portability:** Requires `web_fetch` (Google Patents, Espacenet, USPTO), `WebSearch` (adjacent academic art), Node.js with `docx` package, and optionally Lens.org API key for citation-graph signals. Works in Claude Code CLI natively. In Claude.ai with web tools + Code Execution + BYOK Lens.org, the workflow is supported.
> **Out of scope:** trademark, copyright, trade-secret. These are flagged at intake. Use a different skill or qualified counsel.
> **Legal disclaimer:** This skill produces search signal, not legal advice. Verdicts are technical assessments. **Always consult a patent attorney before filing or licensing decisions.**
## Non-Generic Framing — The Differentiator
This skill is **prior-art + landscape intelligence**. It **refuses to be a bucket**. Every invocation commits to one of five sub-use-cases via the grill-me intake before any search runs. The chosen sub-use-case dictates the entire search strategy, ranking heuristics, and DOCX emphasis.
| Sub-use-case | Search strategy | DOCX emphasis |
|---|---|---|
| **Novelty search** | Narrow + claims-text focused; pre-filing date irrelevant | Closest art + claim-differentiation |
| **Freedom-to-operate** | Broad + active patents only; jurisdiction-filtered | FTO flags + claim-by-claim risk |
| **Competitive landscape** | Breadth + filer tally + CPC trends | Filer map + investment hotspots |
| **Acquisition diligence** | Specific assignee + portfolio scope + assignment chain | Portfolio table + ownership verification |
| **Litigation prior-art** | Specific target patent + adjacent art before priority date | Knock-out candidates ranked by relevance |
See [`references/sub_use_case_routing.md`](references/sub_use_case_routing.md) for the canon.
## Agent Integrity Rules (Research-Pack Convention)
Locked verbatim per PR #657 audit.
- **Execution discipline.** Sequential search calls only. **1 query/sec rate limit.** Confirm response received before next call.
- **Source discipline.** Cite only patents returned by THIS session's tool calls. Training knowledge labeled `[Not from search — reference information]` and excluded from counts.
- **Three-count tracking.** Queries sent / patents received (shown) / patents cited. Surfaced in audit log.
- **Retry policy.** On failure → wait 3s → retry once → log. After 3 consecutive failures across tools: stop, alert user, explain what's missing.
- **Plan-tier detection.** Lens.org free tier = 1000 queries/month. Google Patents has no auth but rate-limits per IP. Detect and surface caps.
## Phase 1: Grill-Me Intake (6 forcing questions, one at a time)
### Q1 (root) — Invention description
> **Describe the invention in 23 sentences. What does it do, and what's new about it?**
>
> *Why I'm asking:* Concept and keyword extraction depends entirely on a precise description. Vague descriptions ("AI for healthcare", "a better widget") will be rejected — push back and ask the user to specify what the invention does and what differentiates it from existing approaches.
**Refuse mush.** If answer is generic, ask once more: "What does it do that existing systems don't?" Then commit (with caveat in DOCX).
### Q2 (depends on Q1) — Sub-use-case commitment
> **What's the purpose of this search? Pick one:**
>
> 1. Novelty search (am I novel enough to file)
> 2. Freedom-to-operate (will I get sued if I ship)
> 3. Competitive landscape (who else plays here)
> 4. Acquisition diligence (does target really own X)
> 5. Litigation prior-art hunting (kill a specific patent)
>
> *Why I'm asking:* Each path uses a fundamentally different search strategy. I'll **refuse to start without you picking one**.
Forcing format. If user says "all of them", push for the primary purpose — secondary purposes can run as follow-up searches.
### Q3 (asked only if Q2 ∈ {FTO, landscape, diligence}) — Jurisdictions
> **Which jurisdictions matter? Pick all that apply: US / EP / CN / JP / KR / PCT / worldwide.**
>
> *Why I'm asking:* FTO only matters where you'll sell. Landscape changes radically by region. Diligence requires checking all jurisdictions where the target operates.
Skip for novelty (priority date is jurisdictionally portable) and litigation (jurisdiction is set by the target patent).
### Q4 (depends on Q1) — Known prior art
> **Have you already seen prior art close to this? Cite a patent number or paper.**
>
> *Why I'm asking:* If you know one piece of art, I can search adjacent to it — much more precise than starting cold. If you don't, that's fine — just confirm.
Anchoring. Accept "none" but ask if the user has seen *any* related work even informally.
### Q5 (depends on Q2) — Risk tolerance
> **Risk tolerance for this search: strict (one close hit means abandon the path) or signal-gathering (you want the lay of the land regardless)?**
>
> *Why I'm asking:* Strict mode ranks aggressively and surfaces verdict-grade hits; signal mode prioritizes breadth and visualizations.
Asked for novelty and FTO; skipped for pure landscape (always signal-gathering by definition).
### Q6 (asked only if Q2 ∈ {novelty, FTO}) — Attorney status
> **Have you spoken to a patent attorney? This skill produces search signal, not legal advice. Confirm you understand this is for technical assessment only.**
>
> *Why I'm asking:* Novelty and FTO have legal consequences. The skill's verdict is signal-grade; legal positions require qualified counsel.
**Triggers the legal-disclaimer footer in the DOCX.** Skipped for landscape and diligence (lower legal exposure).
**Stop condition:** After Q6 (or earlier if dependency skips applied), commit and start Phase 2. Never re-open intake after Phase 2 begins.
## Phase 2: Search Strategy Selection
Deterministic from intake answers. Use `scripts/sub_use_case_router.py`:
```bash
python ../scripts/sub_use_case_router.py \
--sub-use-case novelty \
--jurisdictions "" \
--risk strict \
--known-art "US10000000B2"
```
Returns: query plan (5-8 queries) + ranking heuristic + DOCX emphasis flags.
## Phase 3: Multi-Source Search (Sequential)
### Source priority
1. **Google Patents** (https://patents.google.com) — workhorse, no auth required, broad coverage
2. **Espacenet** (https://worldwide.espacenet.com) — global coverage, good for non-US art
3. **USPTO PPS** (https://ppubs.uspto.gov) — US deep dive
4. **Lens.org** (https://www.lens.org) — citation graph, BYOK API key required
### Per-sub-use-case query patterns
**Novelty:**
- 3 narrow queries on invention-specific terminology (Google Patents)
- 2 broad concept queries with synonyms (Google Patents + Espacenet)
- 1 CPC-class-restricted query if class identified from initial hits
**FTO:**
- Jurisdiction-filtered: only active patents (not expired, not abandoned)
- Date filter: priority < today
- Active-claim text extraction for each hit
**Competitive landscape:**
- Broader queries on the technology space
- CPC class identification → tally top filers in that class
- 10-year filing trend by year per top-5 filer
**Acquisition diligence:**
- Specific assignee searches (target company + subsidiaries + named inventors)
- Assignment chain check (USPTO assignment recordation)
- Family resolution for deduplication
**Litigation prior-art:**
- Target patent input required (number)
- Priority date extraction
- Search for art before priority date in same CPC classes
- Adjacent-claim-language search
### Sequential discipline
1 q/sec across ALL sources combined. Tracked via `scripts/citation_tracker.py` with timestamp-enforced gap.
## Phase 4: Claim Extraction + Relevance Scoring
For each closest-art hit:
- Pull **independent claim 1** (the broadest claim — primary anticipation/obviousness vehicle)
- Pull **key dependent claims** (claims that add the inventive step)
- Score relevance against invention description (overlap of claim language with Q1 terminology)
Rank by score. Verdict per sub-use-case (NOVEL / POTENTIALLY NOVEL / NOT NOVEL for novelty; CLEAR / FLAGGED / HIGH RISK per jurisdiction for FTO).
## Phase 5: Citation Graph + Family Resolution
### Citation graph (Lens.org BYOK)
If user provides Lens.org API key:
- Foundational-patent identification (cited-by count > threshold, typically 50+)
- Recent high-cite signals (citations in last 24 months as proxy for current activity)
- Forward citations from target patent (litigation prior-art) or from closest art (novelty)
If no Lens.org key: skip; note in audit log; recommend manual citation review on Google Patents.
### Family resolution
Same invention often filed in multiple jurisdictions (US + EP + JP + CN). Group by family ID or priority number to avoid double-counting. Use `scripts/family_resolver.py`:
```bash
python ../scripts/family_resolver.py --hits-file hits.json
# Returns: deduplicated family list + family-member jurisdictions
```
## CPC/IPC Classification Awareness
**Critical:** keyword search alone misses adjacent art. After initial search, extract the CPC/IPC classes from top 5 hits and run **one class-restricted query**. This consistently surfaces art that keyword search misses.
See [`references/cpc_classification_canon.md`](references/cpc_classification_canon.md) for the canon.
## Phase 6: DOCX Generation (8 Sections)
Sub-use-case-dependent emphasis. Via Node.js + `docx` library.
1. **Executive Summary + Verdict** — Sub-use-case banner + one-line verdict (NOVEL / FLAGGED / etc.) + 3-4 key findings + legal disclaimer footer
2. **Closest Prior Art** — 5-10 patents in ranked order. Per hit: hyperlinked title + assignee + filing/priority dates + independent claim 1 text (italicized) + relevance score + relevance rationale (1-2 sentences)
3. **Patent Landscape** — Top filers table (top 10 by count) + 10-year filing trend description + CPC class distribution table. Only for landscape and diligence; abbreviated otherwise.
4. **Citation Graph Signals** — Foundational patents (if Lens-enabled) + recent high-cite activity. If Lens unavailable, note "manual review recommended" and skip table.
5. **Geographic Coverage** — Filings by jurisdiction for top 10 hits. Only for FTO, landscape, diligence; skipped for novelty and litigation.
6. **FTO Flags** (FTO only) — Active patents posing infringement risk. Per flag: hyperlinked patent + jurisdiction + relevant claims + risk level (HIGH/MEDIUM/LOW) + mitigation note.
7. **Strategy + Recommendations** — Sub-use-case-specific:
- Novelty → claim differentiation suggestions
- FTO → design-around hints + jurisdiction strategy
- Landscape → who-to-watch list
- Diligence → red flags in portfolio
- Litigation → ranked knock-out candidates
- **Mandatory disclaimer to consult patent attorney** for any filing/licensing decision.
8. **Audit Log** — Searches table (#, query, source, results, status), counts (sent/shown/cited), tool constraints (plan-tier notes), failed steps, attorney-consultation reminder
### Styling
Arial 12pt body, navy headings (#1a3a5c), light blue table headers (#e8f0f8), red FTO-flag callout. `ExternalHyperlink` patterns:
- Google Patents: `https://patents.google.com/patent/[number]`
- Espacenet: `https://worldwide.espacenet.com/patent/...`
- USPTO: `https://patents.uspto.gov/patent/...`
## Date Discipline
Distinguish at every hit:
- **Filing date** — when the application was first submitted
- **Priority date** — earliest claim of priority (often earlier than filing)
- **Publication date** — when the application became public (typically 18 months after priority)
- **Grant date** — when the patent was granted (later than publication)
Surface the **legally-relevant date** per sub-use-case:
- Novelty → priority date (vs invention's anticipated filing date)
- FTO → grant date + status (active vs expired)
- Landscape → publication date (when public knowledge began)
- Diligence → grant date + assignment date
- Litigation → priority date of target patent (sets the prior-art cutoff)
## Phase 7: Deliver
- Save: `<output-dir>/patent_<invention-slug>_<sub-use-case>_<YYYY-MM-DD>.docx`
- Chat summary: file path + sub-use-case + verdict + audit counts + plan-tier
- Validate: `python scripts/office/validate.py <docx>`
- Reminder: "Consult patent attorney before filing/licensing"
## Tooling
| Script | Role |
|---|---|
| `scripts/citation_tracker.py` | Multi-source three-count audit (Google Patents + Espacenet + USPTO + Lens.org) at `~/.patent_sessions/<session>.json` |
| `scripts/family_resolver.py` | Group same-invention filings across jurisdictions by family ID / priority number |
| `scripts/sub_use_case_router.py` | Deterministic search-strategy selection from intake answers |
## References
- [`references/sub_use_case_routing.md`](references/sub_use_case_routing.md) — 5-sub-use-case canon (7+ sources)
- [`references/cpc_classification_canon.md`](references/cpc_classification_canon.md) — CPC/IPC class follow-up rationale (7+ sources)
- [`references/legal_disclaimer_discipline.md`](references/legal_disclaimer_discipline.md) — when + why disclaimer mandatory (7+ sources)
## Error Handling
| Failure | Behavior |
|---|---|
| User refuses to commit to sub-use-case | Refuse to proceed. Re-ask Q2 with examples. |
| Invention description is generic | Reject answer. Re-ask Q1 with "what does it do that existing systems don't?" |
| Google Patents rate-limits | Wait 3s, retry once. Fall back to Espacenet for that query. Log in audit. |
| Lens.org key missing | Skip citation graph section, note "manual review recommended" in DOCX. |
| Claim text extraction fails | Fall back to abstract; flag as "abstract-only" in relevance rationale. |
| Family resolution incomplete | Note in audit; same-invention duplicates may appear; suggest manual deduplication. |
| All searches return <3 hits | Surface explicitly as "either niche art or genuine gap"; never fabricate. |
| 3 consecutive tool failures | Stop, alert user, explain what's missing. |
| DOCX generation fails | Save raw data as JSON fallback so user doesn't lose work. |
| Target patent number invalid (litigation) | Validate format before search; ask user to confirm. |
## Anti-Patterns To Reject
- Starting any search before user commits to a sub-use-case (refuses generic "patent help")
- Batching all intake questions instead of one at a time
- Accepting vague invention descriptions ("AI for healthcare")
- Keyword-only search without CPC/IPC class follow-up
- Treating family members as separate hits (must be deduplicated)
- Confusing filing date with priority date with publication date
- Skipping the legal disclaimer when sub-use-case has legal consequences
- Reporting a verdict without claim-text evidence
- Fabricating Lens.org citation data when key is absent
- Suggesting design-arounds without acknowledging attorney review is required
- Skipping the audit log
---
**Version:** 1.0.0
**Source spec:** [`megaprompts/11-patent-megaprompt.md`](../../../../megaprompts/11-patent-megaprompt.md)
**Build pattern:** Path B (direct conversion). Research-pack sibling, sub-use-case routing variant.

View file

@ -0,0 +1,154 @@
# CPC/IPC Classification — Why the Class Follow-Up Catches What Keywords Miss
This reference answers exactly one decision: **why does the patent skill always run a CPC/IPC class-restricted query after initial keyword searches, and how does class follow-up systematically surface art that pure keyword search misses?**
## The Core Claim
Keyword search alone systematically misses adjacent prior art. Patent attorneys describe this as the **"different vocabulary problem"**:
- A 1995 patent on "machine learning for image recognition" might describe its invention as "neural network for visual classification" — different terminology, same underlying concept
- A patent in semiconductors might use "transistor channel" where a software paper would use "data flow path" — same idea, different field's vocabulary
- A patent might intentionally use unusual terminology to **broaden claim scope** (legal strategy)
The CPC/IPC classification system was designed precisely to bridge these vocabulary gaps. **Always-run class follow-up** is the skill's mechanical correction for keyword-only blind spots.
## What Are CPC and IPC?
| System | Maintained by | Granularity | Used by |
|---|---|---|---|
| **CPC** (Cooperative Patent Classification) | USPTO + EPO | ~250,000 classes | All major patent offices since 2013 |
| **IPC** (International Patent Classification) | WIPO | ~75,000 classes | Used as fallback in some jurisdictions |
Both are hierarchical. Examples:
- `G06N` → Computer systems based on specific computational models (CPC + IPC)
- `G06N3/00` → Computing arrangements based on biological models
- `G06N3/04` → Architecture, e.g. interconnection topology
- `G06N3/045` → Combinations of networks (deep learning hidden layers)
When a patent examiner classifies a patent, they assign one or more CPC classes. Patents in the same class are conceptually related EVEN IF they use different vocabulary.
## The CPC Class Follow-Up Pattern
After initial keyword search returns top 5 hits:
1. **Extract CPC classes** from those 5 hits
2. **Tally** to find the dominant class (1-3 classes typically)
3. **Run one class-restricted query** — same keywords + CPC class filter
4. **Compare** results to initial keyword-only results
**Empirical observation:** the class-restricted query consistently surfaces 2-5 additional hits that the keyword search missed. Some of these are highly relevant (the "vocabulary mismatch" cases).
## Concrete Examples
### Example 1: AI for medical diagnosis
| Search | Top results |
|---|---|
| Keyword "AI medical diagnosis" | 2017+ patents using "AI" / "ML" / "deep learning" + "diagnosis" |
| **+ CPC G16H50/20** (medical informatics for diagnosis) | + 1990s patents on "expert systems" + "decision support" — same concept, different era's vocabulary |
Without the class follow-up, the searcher would miss 25 years of foundational expert-system art that the USPTO clearly considers prior art.
### Example 2: 3D printing materials
| Search | Top results |
|---|---|
| Keyword "3D printing polymer" | 2010+ patents using "additive manufacturing" + "polymer" |
| **+ CPC B33Y70/00** (materials for additive manufacturing) | + 1980s patents on "stereolithography resin" — predecessor terminology |
### Example 3: Recommender systems
| Search | Top results |
|---|---|
| Keyword "recommendation algorithm" | 2010+ patents using "recommender" / "collaborative filtering" |
| **+ CPC G06Q30/0631** (recommender system for products/services) | + 1990s patents on "preference matching" + "user modeling" |
## Why This Matters Per Sub-Use-Case
### Novelty
Missing class-adjacent art = **false negative**. User concludes invention is novel; later examiner finds the missed art and rejects. Class follow-up prevents this expensive surprise.
### FTO
Missing class-adjacent active patents = **false confidence**. User ships product believing it's clear; gets sued by patent owner whose patent used different vocabulary. Class follow-up surfaces these.
### Litigation prior-art
Missing class-adjacent art before priority date = **weak invalidity case**. The art that would knock out the target patent might be using completely different vocabulary; class follow-up finds it.
### Landscape + Diligence
Class follow-up surfaces the **technology lineage** — which classes the field operates in, who files in each class, how the field has evolved.
## Operational Pattern
In `Phase 3` of patent's SKILL.md:
```
1. Run initial keyword queries (per sub-use-case)
2. Extract CPC classes from top 5 hits → identify 1-3 dominant classes
3. Run ONE class-restricted query: keywords + CPC class filter
4. Merge results, deduplicate, rank
5. (Optional) If multi-class: run one query per dominant class
```
The class follow-up is a **single additional query per dominant class** — minimal budget cost, high signal yield.
## How to Identify the Right Class
After initial search, look at the top 3-5 hits' classification fields. Most patent search interfaces (Google Patents, Espacenet, USPTO PPS) show CPC classes in the metadata.
**Heuristic:**
- If 3+ of top 5 hits share a class → that's your dominant class
- If hits are spread across many classes → dominant class is the most-frequent across the top 10 hits
- If still spread → run class follow-up for top 2 classes (2 extra queries)
## Anti-Patterns
### Skipping class follow-up "to save queries"
The skill's query budget per sub-use-case explicitly allocates 1-2 queries for class follow-up. **Skipping it to save 1 query is the most common false-economy** in patent search. The signal-per-query of class follow-up consistently exceeds keyword-only queries.
### Relying solely on top-1 hit's class
The top-1 hit might be an outlier. Look at the top 3-5 hits' shared classes for the dominant class.
### Treating IPC and CPC as interchangeable
CPC is more granular and modern (post-2013). When available, prefer CPC classes. Fall back to IPC for older patents that haven't been re-classified.
### Class follow-up without CPC class identification
Just running "G06N" with no further specificity is too broad. Use the most specific class that 2+ top hits share (e.g., `G06N3/045` not just `G06N`).
### Ignoring class signals in DOCX
The dominant CPC classes ARE valuable signal for the DOCX. Surface them in Section 3 (Patent Landscape) so the user understands the technology classification of their search space.
## Operational Checklist
- [ ] Initial keyword queries run (per sub-use-case)
- [ ] CPC classes extracted from top 5 hits
- [ ] Dominant class(es) identified (1-3)
- [ ] Class-restricted query run (additional 1-2 queries)
- [ ] Results merged + deduplicated + ranked
- [ ] Dominant CPC classes surfaced in DOCX Section 3 (or Section 2 for novelty)
- [ ] Audit log notes class follow-up as part of search strategy
## Citations (7 sources)
1. **CPC Scheme — USPTO + EPO joint maintenance.** https://www.cooperativepatentclassification.org. Authoritative source for the CPC hierarchy + per-class definitions. The skill recommends consulting CPC scheme for any class beyond top-3 frequency.
2. **WIPO IPC Strategic Plan + IPC Schema.** https://www.wipo.int/classifications/ipc. Source for IPC fallback discipline (used for pre-2013 patents that lack CPC reclassification).
3. **Mowery, D. C., Nelson, R. R., Sampat, B. N., & Ziedonis, A. A., *Ivory Tower and Industrial Innovation* (Stanford U Press, 2004).** Source for the historical analysis of how patent classifications evolve over time and why cross-era keyword search fails. Empirical evidence for the "different vocabulary problem".
4. **WIPO PATENTSCOPE search documentation.** https://patentscope.wipo.int. Source for cross-jurisdictional class search syntax (especially for non-US/EP jurisdictions).
5. **Cohen, W. M., Nelson, R. R., & Walsh, J. P., "Protecting Their Intellectual Assets" — *NBER Working Paper* 7552 (2000).** Source for the empirical evidence that keyword-only patent search systematically under-reports prior art (especially in fast-moving technology areas).
6. **MPEP §901 — *Manual of Patent Examining Procedure* (USPTO).** Source for the examiner-side discipline of using CPC classes for prior-art search. The skill mirrors examiner discipline by including class follow-up as mandatory.
7. **Lemley, M. A., & Sampat, B., "Examiner Characteristics and Patent Office Outcomes" — *Review of Economics and Statistics* 94(3), 2012.** Source for the empirical analysis showing experienced examiners use CPC classes more aggressively + produce stronger prior-art rejections. Class follow-up is the experienced-examiner technique.

View file

@ -0,0 +1,137 @@
# Legal Disclaimer Discipline — When + Why Mandatory
This reference answers exactly one decision: **for which sub-use-cases does the patent skill require a legal disclaimer in the DOCX, and what does the disclaimer need to say?**
## The Core Rule
Patent law has **immediate financial + legal consequences**. The skill produces **search signal**, not legal advice. A reader who confuses the two and skips attorney consultation can face:
- Patent infringement liability (FTO failure → lawsuit)
- Patent application rejection (novelty failure → expensive abandoned application)
- Wasted R&D investment (proceeding on a confidence the search couldn't actually justify)
**The disclaimer is a safety property**, comparable to drafts-only in inbox-triage. It prevents foreseeable user harm.
## When Disclaimer Is Mandatory
| Sub-use-case | Disclaimer mandatory? | Rationale |
|---|---|---|
| **Novelty search** | YES (Q6 triggers it) | Prosecution decisions have legal consequences |
| **Freedom-to-operate** | YES (Q6 triggers it) | Shipping decisions have liability consequences |
| **Competitive landscape** | Optional (recommended) | Lower legal exposure; more strategic than legal |
| **Acquisition diligence** | Optional (strongly recommended) | M&A context — legal review usually already part of process |
| **Litigation prior-art** | Optional (strongly recommended) | Litigation context — counsel almost always involved |
For **mandatory** sub-use-cases, the disclaimer:
1. Appears in **Executive Summary** footer (Section 1)
2. Appears in **Strategy + Recommendations** body (Section 7)
3. Appears in **Audit Log** as a final reminder (Section 8)
For **optional** sub-use-cases, the disclaimer appears only in Section 8 as a reminder.
## What the Disclaimer Says
### Mandatory version (novelty + FTO)
> **⚖️ Legal Disclaimer:** This document is **search signal, not legal advice**. The verdict ({NOVEL/CLEAR/etc.}) is a technical assessment based on the patents found in this session's tool calls. **Patent novelty and freedom-to-operate determinations have legal consequences and require qualified counsel.**
>
> **Before any filing or licensing decision:**
> - Consult a registered patent attorney in your jurisdiction(s)
> - Provide them this dossier as starting material; they will conduct independent verification + opinion
> - Their opinion is privileged and admissible; this skill's output is neither
>
> The skill does not establish attorney-client privilege. The skill's verdict does not constitute a legal opinion.
### Optional version (landscape, diligence, litigation)
> **⚖️ Reminder:** This document is search signal. Patent attorney consultation is recommended for any decisions arising from this analysis.
## Why Disclaimer Discipline Matters
### Reason 1: Legal liability framing
Without disclaimer, a user could (in extreme cases) claim the skill misled them into a legal decision. Disclaimer makes the skill's role unambiguous: **technical assessment, not legal opinion**.
### Reason 2: Setting realistic expectations
Even a well-executed patent search has limits:
- Some patents may not be indexed in queried sources (especially recent applications)
- Some patents may be classified in unexpected CPC classes
- Some art may be in non-patent literature (academic papers, products, manuals)
- Some art may be in different languages (non-English jurisdictions)
The disclaimer tells the user: "I did the best technical search I could; counsel will catch what I might have missed."
### Reason 3: Privileged communication
Attorney-client conversations are **privileged** — not admissible against the user in litigation. This skill's output is **not privileged**. If the user later faces litigation, opposing counsel can subpoena the dossier as evidence of what the user knew.
The disclaimer reminds users to NOT rely solely on the dossier for high-stakes decisions; an attorney's opinion provides privilege.
### Reason 4: Jurisdiction-specific nuance
Patent law varies by jurisdiction:
- **First-to-file vs first-to-invent** (US switched to first-to-file in 2013; some countries differ)
- **Grace periods** (US has 1-year; many countries have none)
- **Doctrine of equivalents** (varies by jurisdiction)
- **Inequitable conduct** (US-specific; may affect prosecution strategy)
The skill cannot capture all jurisdictional nuance. Counsel can.
## Anti-Patterns
### Skipping disclaimer because user said "I'm a patent attorney"
The user might be a patent attorney, but they might also be running this for a less-experienced colleague or client. The disclaimer is **always** in the document because the document might outlive the original requester.
### Burying disclaimer in fine print
Mandatory-sub-use-case disclaimer appears in Sections 1, 7, and 8 — three locations. Reader cannot miss it.
### Replacing disclaimer with "consult your attorney"
The disclaimer must be specific about what the skill does and doesn't claim. "Consult your attorney" alone is insufficient; the disclaimer also needs to clarify:
- What the verdict means (technical assessment)
- What counsel adds (privilege, jurisdiction expertise, opinion)
- What's not covered (non-patent prior art, language coverage gaps)
### Conflating disclaimer with legal-advice disclaimer
Some skills use generic "this is not legal advice" boilerplate. The patent skill's disclaimer is specifically about patent novelty/FTO determinations and the role of qualified counsel — not generic.
### Removing disclaimer to "make the document feel more authoritative"
Authority comes from technical rigor (claim-text extraction, family resolution, CPC class follow-up), not from omitting safety disclaimers. The disclaimer **enhances** authority by making the skill's role transparent.
## Operational Checklist
For novelty + FTO (mandatory):
- [ ] Disclaimer in Executive Summary (Section 1) footer
- [ ] Disclaimer in Strategy section (Section 7)
- [ ] Disclaimer in Audit Log (Section 8) as final reminder
- [ ] Disclaimer text matches template above (don't paraphrase the legal language)
- [ ] Q6 attorney-status answer recorded in audit log
For landscape + diligence + litigation (optional):
- [ ] Reminder version in Audit Log (Section 8)
- [ ] Strategy section (Section 7) includes "consult patent attorney for [decision-specific context]"
## Citations (7 sources)
1. **MPEP §1.4 + §1.5 — *Manual of Patent Examining Procedure* (USPTO).** Source for the role of registered patent attorneys/agents in prosecution. The disclaimer's reference to "qualified counsel" tracks USPTO's registration framework.
2. **AIPLA *Code of Ethics* — American Intellectual Property Law Association.** Source for the privileged-communication framing. AIPLA's guidance on lay-vs-attorney communication informs the "this skill is not privileged" disclaimer language.
3. **35 USC §282 — *Presumption of validity*.** US patent statute. Source for understanding what "valid" means legally + why a search-signal verdict is not a legal validity opinion.
4. **35 USC §271 — *Infringement of patent*.** Source for the FTO disclaimer framing. Liability for infringement is a legal determination; the skill's "CLEAR/FLAGGED/HIGH RISK" verdict is technical.
5. **EPO Guidelines for Examination, Part E.** Source for European-jurisdiction differences (no grace period, different inventive-step analysis). The disclaimer's "jurisdiction-specific nuance" caveat tracks these variations.
6. **PCT Article 39 — Patent Cooperation Treaty.** Source for the cross-jurisdiction prosecution complexity that justifies counsel involvement. PCT national-phase entries each require local-counsel coordination.
7. **Fischer, T., & Henkel, J., "Patent Trolls on Markets for Technology" — *Research Policy* 41(9), 2012.** Empirical evidence for the cost of FTO mistakes. Patent assertion entities (PAEs) have made FTO failures financially severe; the disclaimer's emphasis on counsel reflects this risk reality.

View file

@ -0,0 +1,252 @@
# Sub-Use-Case Routing — The 5 Patent Search Strategies
This reference answers exactly one decision: **given the user's Q2 commitment, which search strategy / ranking heuristic / DOCX emphasis applies?**
The patent skill refuses to be a generic "patent search". Q2 is mandatory. The 5 sub-use-cases use **fundamentally different** strategies — running a novelty search and calling it FTO produces wrong answers in dangerous ways.
## Why Sub-Use-Case Commitment Matters
| Sub-use-case | Wrong-strategy danger |
|---|---|
| Novelty | Generic search misses claim-text proximity → false negatives ("looks novel, isn't") |
| FTO | Generic search includes expired/abandoned patents → false positives ("looks blocked, isn't") |
| Landscape | Generic search misses CPC class trends → incomplete competitive picture |
| Diligence | Generic search misses assignment chain → ownership-verification gaps |
| Litigation | Generic search includes art after priority date → useless for invalidation |
The 5 strategies are **not interchangeable**. The skill enforces commitment to prevent strategy mismatches.
## Strategy 1: Novelty Search
### Question being answered
"Am I novel enough to file? Is there art that anticipates or makes obvious my invention?"
### Search emphasis
- **Narrow** queries on invention-specific terminology (don't drown in adjacent art)
- **Claims-text focused** (the legal test is claim-by-claim anticipation)
- **Pre-filing date irrelevant** — anything published before user's filing is potential art
### Query plan
1. 3 narrow Google Patents queries on invention-specific terms
2. 2 broad concept queries with synonyms (Google Patents + Espacenet)
3. 1 CPC class-restricted query (after class identification from initial hits)
4. (Optional Lens.org) forward citations from any closest art
**Total: 6-7 sequential queries.**
### Ranking heuristic
Rank by **claim-text overlap** with user's invention description (Q1). Top hits are those whose claim 1 most overlaps in technical terminology. Surface independent claim 1 verbatim for each top-5 hit.
### Verdict scale
- **NOVEL** — closest hit has <30% claim-text overlap; clear differentiation possible
- **POTENTIALLY NOVEL** — 30-60% overlap; differentiation possible but requires careful claim drafting
- **NOT NOVEL** — >60% overlap; invention as described is anticipated by closest art
### DOCX emphasis
- Section 2 (Closest Prior Art): expanded — 8-10 hits with full claim-1 text
- Section 7 (Strategy): claim-differentiation suggestions
- Sections 3 + 5 (Landscape + Geographic): abbreviated
- Mandatory legal disclaimer footer
## Strategy 2: Freedom-to-Operate
### Question being answered
"If I ship in jurisdiction X, will I get sued for infringement?"
### Search emphasis
- **Active patents only** — expired/abandoned patents can't sue
- **Jurisdiction-filtered** — FTO only matters where user sells
- **Date filter:** priority date < today (no pending applications without published claims)
- **Independent + dependent claims** — both relevant to infringement analysis
### Query plan
1. Per jurisdiction (Q3): 2-3 queries with jurisdiction filter (US: USPTO; EP: Espacenet; etc.)
2. Active-status filter applied to all
3. CPC class follow-up after initial hits
4. (Optional) assignment chain check for active assignee context
**Total: 8-15 sequential queries (scales with # of jurisdictions).**
### Ranking heuristic
Rank by **claim-by-claim infringement risk**. For each active patent: which independent claims would the user's product practice? High risk = at least one independent claim covers user's product as designed.
### Verdict scale (per jurisdiction)
- **CLEAR** — no active patents pose infringement risk
- **FLAGGED** — 1-2 active patents may pose risk; design-around viable
- **HIGH RISK** — 3+ active patents pose risk; design changes required OR licensing path needed
### DOCX emphasis
- Section 6 (FTO Flags): expanded — per-flag risk per jurisdiction
- Section 5 (Geographic Coverage): expanded
- Section 7 (Strategy): design-around hints + jurisdiction strategy
- Mandatory legal disclaimer footer
## Strategy 3: Competitive Landscape
### Question being answered
"Who else plays in this technology space? What are the trends?"
### Search emphasis
- **Broader queries** on the technology space (NOT the specific invention)
- **CPC class identification** drives the analysis
- **Top filer tally** — who files most patents in the space
- **10-year filing trend** by year per top-5 filer
### Query plan
1. 2-3 broad queries on the technology space
2. CPC class extraction from top hits
3. 1 query per top-5 filer to gauge their portfolio
4. (Optional Lens.org) citation graph for foundational patents
**Total: 8-10 sequential queries.**
### Ranking heuristic
Rank by filer count + recency. Top-5 filers + 3 emerging entrants (filers with first patent in last 2 years).
### Verdict scale
- **CONCENTRATED** — top-3 filers own >60% of patents in the space
- **COMPETITIVE** — top-10 filers own 60-90%; mature competitive market
- **EMERGING** — long tail of filers; market is still defining itself
### DOCX emphasis
- Section 3 (Patent Landscape): expanded — top filers table + 10-yr trend + CPC distribution
- Section 5 (Geographic Coverage): expanded
- Sections 2 + 4 (Closest Art + Citation Graph): abbreviated
- Section 7 (Strategy): who-to-watch list + emerging-entrants signal
- Legal disclaimer optional (lower legal exposure)
## Strategy 4: Acquisition Diligence
### Question being answered
"Does target company actually own the patents they claim? Is there portfolio depth?"
### Search emphasis
- **Specific assignee searches** — target company + subsidiaries + named inventors
- **Assignment chain check** — USPTO assignment recordation
- **Family resolution** — deduplicate same-invention across jurisdictions
- **Portfolio scope** — are patents in core business areas or peripheral?
### Query plan
1. 2-3 assignee-name queries (Google Patents + USPTO assignee search)
2. Subsidiary searches if user provides org chart
3. Inventor searches for key named inventors
4. Assignment recordation lookups for ownership verification
5. Family resolution across all hits
**Total: 6-12 sequential queries.**
### Ranking heuristic
Group by family. Within family, surface earliest priority. Across families, rank by:
- Citation count (foundational vs niche)
- Filing recency (active R&D vs legacy)
- Claim breadth (broad coverage vs narrow)
### Verdict scale
- **PORTFOLIO VERIFIED** — claimed patents owned, assignment chains clean, no orphans
- **PARTIAL VERIFICATION** — some claimed patents not found OR assignment chain unclear
- **OWNERSHIP RISK** — significant claimed patents not owned by target OR major assignment gaps
### DOCX emphasis
- Section 3 (Patent Landscape): expanded as portfolio table
- Section 5 (Geographic Coverage): expanded
- Section 7 (Strategy): red flags in portfolio + ownership-verification flags
- Legal disclaimer optional but recommended (M&A context)
## Strategy 5: Litigation Prior-Art
### Question being answered
"Can I invalidate this specific patent? What art exists before its priority date?"
### Search emphasis
- **Target patent input required** (number)
- **Priority date extraction** — sets the prior-art cutoff
- **Search before priority date in same CPC classes**
- **Adjacent-claim-language search** — art that uses similar claim language
### Query plan
1. Fetch target patent (extract priority date + claims + CPC classes)
2. CPC class queries with date filter (priority < target's priority)
3. Keyword queries on independent claim language with date filter
4. (Optional Lens.org) forward citations from target's cited art
**Total: 5-8 sequential queries.**
### Ranking heuristic
Rank by **knock-out potential** — claim-by-claim anticipation/obviousness. Highest rank: art that anticipates ALL elements of target's broadest independent claim.
### Verdict scale
- **KNOCK-OUT FOUND** — art clearly anticipates all elements of broadest claim
- **STRONG OBVIOUSNESS COMBINATION** — multiple pieces of art combine to cover all elements
- **WEAK OBVIOUSNESS** — art relevant but anticipation/obviousness argument is uphill
- **NO MATERIAL ART FOUND** — patent appears strong against this prior-art set
### DOCX emphasis
- Section 2 (Closest Prior Art): expanded — ranked knock-out candidates with claim-language overlap
- Section 7 (Strategy): per-claim invalidity analysis
- Sections 3 + 5 (Landscape + Geographic): abbreviated
- Legal disclaimer optional but recommended (litigation context)
## Out-of-Scope Topics (Flagged at Intake)
| Topic | Why out of scope |
|---|---|
| Trademark | Different legal regime, different sources (USPTO TESS not Patent Office) |
| Copyright | No formal search system; rights attach automatically |
| Trade secret | By definition, not in public records |
If user asks for any of these → halt at intake, recommend appropriate skill or attorney.
## Operational Checklist
- [ ] Q2 sub-use-case picked (no "all of them")
- [ ] `scripts/sub_use_case_router.py` returns query plan + ranking heuristic + DOCX flags
- [ ] Search emphasis matches sub-use-case (not generic)
- [ ] Verdict scale per sub-use-case applied
- [ ] DOCX emphasis adjusted (not all 8 sections expanded for every sub-use-case)
- [ ] Legal disclaimer mandatory for novelty + FTO; optional for landscape/diligence/litigation but recommended
## Citations (7 sources)
1. **MPEP (Manual of Patent Examining Procedure) — USPTO.** Source for the legal definitions of novelty (35 USC §102) vs FTO (no explicit USC; case law) vs anticipation/obviousness. The verdict scales follow MPEP terminology.
2. **35 USC §102 + §103 — US patent statute.** Source for the priority-date-as-cutoff rule for novelty (§102) and the obviousness combination doctrine (§103) that drives litigation prior-art ranking.
3. **WIPO Patent Cooperation Treaty (PCT) procedural docs.** Source for the cross-jurisdiction family discipline. PCT applications generate national-phase entries in many jurisdictions; the family resolver follows WIPO's family-id taxonomy.
4. **EPO Guidelines for Examination — European Patent Office.** Source for the EP-specific FTO discipline. EP active-status filtering uses EPO's "in force" status field.
5. **USPTO Patent Public Search documentation.** Source for the USPTO PPS query syntax and assignment-recordation lookup endpoints used in acquisition diligence.
6. **Google Patents search documentation + advanced operators.** Source for the keyword + CPC class + date filter syntax. Google Patents indexes all PCT national-phase entries plus most jurisdictions' grant data.
7. **Lens.org API documentation (https://docs.api.lens.org).** Source for the citation-graph queries. Lens.org's citation API exposes forward + backward citations with citation-count thresholds for foundational-patent identification.

View file

@ -0,0 +1,241 @@
#!/usr/bin/env python3
"""citation_tracker.py — Patent skill three-count audit across multi-source patent search.
Stdlib-only. Mirrors litreview/grants/dossier trackers but adapted for patent's
4-source workflow:
- Google Patents (workhorse, no auth)
- Espacenet (global)
- USPTO PPS (US deep dive)
- Lens.org (BYOK, citation graph)
Tracked counts:
- searches_per_source (broken out by source)
- patents_received_total
- patents_cited_total
- patents_cited_by_source
- sub_use_case (recorded at start; drives audit verbatim)
- lens_byok_used (boolean surfaced in audit log)
Enforces 1s sequential discipline across ALL sources combined.
Usage:
python citation_tracker.py --action start --session patent-MS-novelty-20260515 --invention "..." --sub-use-case novelty
python citation_tracker.py --action record_search --session ... --source google_patents --query "..."
python citation_tracker.py --action record_received --session ... --source google_patents --count 10
python citation_tracker.py --action record_cited --session ... --source google_patents --patent-num "US10000000B2"
python citation_tracker.py --action record_lens_byok --session ...
python citation_tracker.py --action status --session ...
python citation_tracker.py --action close --session ...
"""
import argparse
import json
import sys
from datetime import datetime, timezone
from pathlib import Path
from typing import Any, Dict, List, Optional
SESSIONS_DIR = Path.home() / ".patent_sessions"
MIN_GAP_SECONDS = 1.0
VALID_SOURCES = ["google_patents", "espacenet", "uspto", "lens", "websearch"]
VALID_SUB_USE_CASES = ["novelty", "fto", "landscape", "diligence", "litigation"]
def session_path(name: str) -> Path:
return SESSIONS_DIR / f"{name}.json"
def load_session(name: str) -> Dict[str, Any]:
p = session_path(name)
if not p.exists():
raise FileNotFoundError(f"Session not found: {name}")
return json.loads(p.read_text(encoding="utf-8"))
def save_session(name: str, data: Dict[str, Any]) -> None:
SESSIONS_DIR.mkdir(parents=True, exist_ok=True)
session_path(name).write_text(json.dumps(data, indent=2), encoding="utf-8")
def now_iso() -> str:
return datetime.now(timezone.utc).isoformat()
def now_ts() -> float:
return datetime.now(timezone.utc).timestamp()
def action_start(name: str, invention: Optional[str], sub_use_case: Optional[str]) -> Dict[str, Any]:
if session_path(name).exists():
raise FileExistsError(f"Session already exists: {name}")
if sub_use_case and sub_use_case not in VALID_SUB_USE_CASES:
raise ValueError(f"Invalid sub-use-case '{sub_use_case}'. Pick from: {VALID_SUB_USE_CASES}")
data: Dict[str, Any] = {
"session": name,
"invention": invention or "",
"sub_use_case": sub_use_case or "",
"started_at": now_iso(),
"ended_at": None,
"lens_byok_used": False,
"searches": [],
"received_log": [],
"cited": [],
"counts": {
"searches_total": 0,
"searches_by_source": {s: 0 for s in VALID_SOURCES},
"received_total": 0,
"received_by_source": {s: 0 for s in VALID_SOURCES},
"cited_total": 0,
"cited_by_source": {s: 0 for s in VALID_SOURCES},
},
}
save_session(name, data)
return data
def action_record_search(name: str, source: str, query: str) -> Dict[str, Any]:
data = load_session(name)
if source not in VALID_SOURCES:
raise ValueError(f"Invalid source '{source}'. Pick from: {VALID_SOURCES}")
if data["searches"]:
last_ts = data["searches"][-1].get("ts", 0)
gap = now_ts() - last_ts
if gap < MIN_GAP_SECONDS:
raise RuntimeError(
f"Sequential discipline violated: {gap:.2f}s gap (need >= {MIN_GAP_SECONDS}s). "
f"Wait {MIN_GAP_SECONDS - gap:.2f}s more."
)
data["searches"].append({"source": source, "query": query, "at": now_iso(), "ts": now_ts()})
data["counts"]["searches_total"] += 1
data["counts"]["searches_by_source"][source] += 1
save_session(name, data)
return data
def action_record_received(name: str, source: str, count: int) -> Dict[str, Any]:
data = load_session(name)
if source not in VALID_SOURCES:
raise ValueError(f"Invalid source '{source}'")
data["received_log"].append({"source": source, "count": count, "at": now_iso()})
data["counts"]["received_total"] += count
data["counts"]["received_by_source"][source] += count
save_session(name, data)
return data
def action_record_cited(name: str, source: str, patent_num: str, title: Optional[str]) -> Dict[str, Any]:
data = load_session(name)
if source not in VALID_SOURCES:
raise ValueError(f"Invalid source '{source}'")
if any(c["patent_num"] == patent_num for c in data["cited"]):
return data
data["cited"].append({"source": source, "patent_num": patent_num, "title": title, "at": now_iso()})
data["counts"]["cited_total"] += 1
data["counts"]["cited_by_source"][source] += 1
save_session(name, data)
return data
def action_record_lens_byok(name: str) -> Dict[str, Any]:
data = load_session(name)
data["lens_byok_used"] = True
save_session(name, data)
return data
def action_status(name: str) -> Dict[str, Any]:
return load_session(name)
def action_close(name: str) -> Dict[str, Any]:
data = load_session(name)
if data.get("ended_at") is None:
data["ended_at"] = now_iso()
save_session(name, data)
return data
def render_status_human(data: Dict[str, Any]) -> str:
out: List[str] = []
out.append(f"Session: {data['session']}")
out.append(f"Invention: {data.get('invention', '(unset)')}")
out.append(f"Sub-use-case: {data.get('sub_use_case', '(unset)')}")
out.append(f"Lens.org BYOK: {'YES' if data.get('lens_byok_used') else 'no (citation graph skipped)'}")
out.append(f"Started: {data['started_at']}")
out.append(f"Ended: {data.get('ended_at') or '(active)'}")
out.append("")
c = data["counts"]
out.append(f"Total searches: {c['searches_total']}")
out.append("By source:")
for src, n in c["searches_by_source"].items():
if n > 0:
out.append(f" {src:<18s} {n}")
out.append("")
out.append(f"Patents received: {c['received_total']}")
out.append(f"Patents cited: {c['cited_total']}")
out.append("Cited by source:")
for src, n in c["cited_by_source"].items():
if n > 0:
out.append(f" {src:<18s} {n}")
out.append("")
out.append("Audit block (paste in DOCX Section 8):")
out.append(
f" Searches: {c['searches_total']} (Google Patents: {c['searches_by_source'].get('google_patents', 0)}, "
f"Espacenet: {c['searches_by_source'].get('espacenet', 0)}, "
f"USPTO: {c['searches_by_source'].get('uspto', 0)}, "
f"Lens.org: {c['searches_by_source'].get('lens', 0)}). "
f"Patents received: {c['received_total']}. Patents cited: {c['cited_total']}. "
f"Lens.org BYOK: {'used' if data.get('lens_byok_used') else 'not provided (citation graph skipped)'}."
)
return "\n".join(out)
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--action", required=True, choices=["start", "record_search", "record_received", "record_cited", "record_lens_byok", "status", "list", "close"])
parser.add_argument("--session")
parser.add_argument("--invention")
parser.add_argument("--sub-use-case", choices=VALID_SUB_USE_CASES)
parser.add_argument("--source", choices=VALID_SOURCES)
parser.add_argument("--query")
parser.add_argument("--count", type=int)
parser.add_argument("--patent-num")
parser.add_argument("--title")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
try:
if args.action == "start":
result = action_start(args.session, args.invention, args.sub_use_case)
elif args.action == "record_search":
result = action_record_search(args.session, args.source, args.query)
elif args.action == "record_received":
result = action_record_received(args.session, args.source, args.count)
elif args.action == "record_cited":
result = action_record_cited(args.session, args.source, args.patent_num, args.title)
elif args.action == "record_lens_byok":
result = action_record_lens_byok(args.session)
elif args.action == "status":
result = action_status(args.session)
elif args.action == "close":
result = action_close(args.session)
else:
SESSIONS_DIR.mkdir(parents=True, exist_ok=True)
result = [{"session": p.stem, "data": json.loads(p.read_text(encoding="utf-8"))} for p in sorted(SESSIONS_DIR.glob("*.json"))]
except (FileNotFoundError, FileExistsError, ValueError, RuntimeError) as e:
print(f"error: {e}", file=sys.stderr); return 2
if args.output == "json":
print(json.dumps(result, indent=2, default=str))
else:
if args.action == "list":
print(json.dumps(result, indent=2, default=str))
else:
print(render_status_human(result))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))

View file

@ -0,0 +1,260 @@
#!/usr/bin/env python3
"""family_resolver.py — Deduplicate same-invention patent filings across jurisdictions.
Stdlib-only. The same invention is often filed in multiple jurisdictions
(US + EP + JP + CN of one underlying invention). They share a "family" identifier
or a common priority application number.
Without family resolution, a multi-jurisdiction search returns the same invention
multiple times inflating the perceived prior-art set and wasting reviewer
attention.
Family resolution rules:
1. If two patents share the same `family_id`, they're family members
2. If two patents share the same `priority_number`, they're family members
3. If two patents share the same `priority_date` AND have 80% applicant overlap
AND 80% inventor overlap likely family (heuristic, flag with confidence)
For each family: surface ONE representative member (typically earliest priority OR
US member for US-context users) and list all family-member jurisdictions.
NO LLM CALLS. Pure JSON aggregation + heuristic clustering.
Input file format (`--hits-file`):
[
{
"patent_num": "US10000000B2",
"title": "...",
"family_id": "F12345678",
"priority_number": "US15/123,456",
"priority_date": "2018-03-15",
"filing_date": "2019-03-14",
"publication_date": "2020-09-15",
"grant_date": "2022-01-10",
"jurisdiction": "US",
"assignee": "Acme Corp",
"inventors": ["Smith, J", "Jones, K"]
}
]
Usage:
python family_resolver.py --hits-file /tmp/hits.json
python family_resolver.py --hits-file /tmp/hits.json --output json
python family_resolver.py --sample
"""
import argparse
import json
import sys
from collections import defaultdict
from pathlib import Path
from typing import Any, Dict, List, Optional, Set
SAMPLE_HITS = [
{
"patent_num": "US10000000B2",
"title": "Machine learning sepsis prediction system",
"family_id": "F12345678",
"priority_number": "US15/123,456",
"priority_date": "2018-03-15",
"filing_date": "2019-03-14",
"jurisdiction": "US",
"assignee": "Acme Corp",
"inventors": ["Smith, J", "Jones, K"],
},
{
"patent_num": "EP3500000B1",
"title": "Système de prédiction sépticémie par apprentissage automatique",
"family_id": "F12345678", # same family
"priority_number": "US15/123,456", # same priority
"priority_date": "2018-03-15",
"filing_date": "2019-03-14",
"jurisdiction": "EP",
"assignee": "Acme Corp",
"inventors": ["Smith, J", "Jones, K"],
},
{
"patent_num": "JP2020100000A",
"title": "敗血症予測システム",
"family_id": "F12345678", # same family
"priority_number": "US15/123,456",
"priority_date": "2018-03-15",
"filing_date": "2019-03-14",
"jurisdiction": "JP",
"assignee": "Acme Corp",
"inventors": ["Smith, J", "Jones, K"],
},
{
"patent_num": "US10500000B2",
"title": "Different sepsis prediction method using LSTMs",
"family_id": "F87654321", # DIFFERENT family
"priority_number": "US16/200,000",
"priority_date": "2019-08-22",
"filing_date": "2020-08-21",
"jurisdiction": "US",
"assignee": "Beta Inc",
"inventors": ["Lee, M"],
},
{
"patent_num": "WO2020/123456",
"title": "PCT application: sepsis prediction with multi-modal data",
"family_id": "F87654321", # same family as US10500000
"priority_number": "US16/200,000",
"priority_date": "2019-08-22",
"jurisdiction": "WO",
"assignee": "Beta Inc",
"inventors": ["Lee, M"],
},
{
"patent_num": "US11000000B1",
"title": "Yet another sepsis ML approach",
"family_id": "F11111111", # alone
"priority_number": "US17/300,000",
"priority_date": "2021-01-10",
"jurisdiction": "US",
"assignee": "Gamma LLC",
"inventors": ["Park, S", "Kim, J"],
},
]
def jaccard_similarity(set1: Set[str], set2: Set[str]) -> float:
if not set1 and not set2:
return 1.0
if not set1 or not set2:
return 0.0
inter = len(set1 & set2)
union = len(set1 | set2)
return inter / union if union > 0 else 0.0
def normalize_name(name: str) -> str:
"""Normalize assignee/inventor names for comparison."""
return name.lower().strip().replace(",", "").replace(".", "")
def resolve_families(hits: List[Dict[str, Any]]) -> Dict[str, Any]:
# Pass 1: group by exact family_id
by_family: Dict[str, List[Dict[str, Any]]] = defaultdict(list)
no_family_id: List[Dict[str, Any]] = []
for h in hits:
fid = h.get("family_id")
if fid:
by_family[fid].append(h)
else:
no_family_id.append(h)
# Pass 2: group remaining by exact priority_number
if no_family_id:
by_priority: Dict[str, List[Dict[str, Any]]] = defaultdict(list)
unmatched: List[Dict[str, Any]] = []
for h in no_family_id:
pn = h.get("priority_number")
if pn:
by_priority[pn].append(h)
else:
unmatched.append(h)
# Add to family map using priority_number as fallback ID
for pn, group in by_priority.items():
by_family[f"PRI:{pn}"] = group
# Pass 3: heuristic clustering for unmatched (priority_date + applicant + inventor overlap)
for h in unmatched:
matched = False
for fid, group in list(by_family.items()):
rep = group[0]
if (h.get("priority_date") == rep.get("priority_date")
and jaccard_similarity({normalize_name(h.get("assignee", ""))}, {normalize_name(rep.get("assignee", ""))}) >= 0.8
and jaccard_similarity({normalize_name(i) for i in h.get("inventors", [])}, {normalize_name(i) for i in rep.get("inventors", [])}) >= 0.8):
by_family[fid].append(h)
matched = True
break
if not matched:
# Solo — use patent_num as family_id
by_family[f"SOLO:{h.get('patent_num', 'unknown')}"] = [h]
# For each family: pick representative (earliest priority date, prefer US member if available)
families: List[Dict[str, Any]] = []
for fid, members in by_family.items():
sorted_members = sorted(members, key=lambda m: (m.get("priority_date", "9999"), 0 if m.get("jurisdiction") == "US" else 1))
rep = sorted_members[0]
jurisdictions = sorted({m.get("jurisdiction", "?") for m in members})
family = {
"family_id": fid,
"representative": {
"patent_num": rep.get("patent_num"),
"title": rep.get("title"),
"assignee": rep.get("assignee"),
"priority_date": rep.get("priority_date"),
"filing_date": rep.get("filing_date"),
"jurisdiction": rep.get("jurisdiction"),
},
"family_member_count": len(members),
"jurisdictions": jurisdictions,
"all_patent_nums": [m.get("patent_num") for m in members],
}
families.append(family)
families.sort(key=lambda f: f["representative"].get("priority_date", "9999"))
return {
"input_hits": len(hits),
"unique_families": len(families),
"deduplication_savings": len(hits) - len(families),
"families": families,
}
def render_human(result: Dict[str, Any]) -> str:
out: List[str] = []
out.append(f"Family resolution complete:")
out.append(f" Input hits: {result['input_hits']}")
out.append(f" Unique families: {result['unique_families']}")
out.append(f" Deduplication savings: {result['deduplication_savings']} duplicate hits removed")
out.append("")
out.append("Families (representative + all jurisdictions):")
for i, f in enumerate(result["families"], 1):
rep = f["representative"]
out.append(f"")
out.append(f" Family {i} (id: {f['family_id']}):")
out.append(f" Representative: {rep['patent_num']} ({rep['jurisdiction']}, priority {rep['priority_date']})")
out.append(f" Title: {rep['title'][:80]}")
out.append(f" Assignee: {rep['assignee']}")
out.append(f" Family size: {f['family_member_count']} member(s) across {len(f['jurisdictions'])} jurisdiction(s)")
out.append(f" Jurisdictions: {', '.join(f['jurisdictions'])}")
if f['family_member_count'] > 1:
out.append(f" All members: {', '.join(f['all_patent_nums'])}")
return "\n".join(out)
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--hits-file", help="Path to JSON file with patent hits")
parser.add_argument("--sample", action="store_true", help="Run on embedded sample")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
if args.sample:
result = resolve_families(SAMPLE_HITS)
elif args.hits_file:
p = Path(args.hits_file)
if not p.exists():
print(f"error: {args.hits_file} not found", file=sys.stderr); return 2
try:
hits = json.loads(p.read_text(encoding="utf-8"))
except json.JSONDecodeError as e:
print(f"error: invalid JSON: {e}", file=sys.stderr); return 2
result = resolve_families(hits)
else:
parser.print_help(); return 0
if args.output == "json":
print(json.dumps(result, indent=2, default=str))
else:
print(render_human(result))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))

View file

@ -0,0 +1,255 @@
#!/usr/bin/env python3
"""sub_use_case_router.py — Deterministic search-strategy from intake answers.
Stdlib-only. Routes to one of 5 patent search strategies based on grill-me
intake answers, returning a query plan + ranking heuristic + DOCX emphasis.
The 5 sub-use-cases:
- novelty am I novel enough to file
- fto will I get sued if I ship
- landscape who else plays here
- diligence does target really own X
- litigation kill a specific patent
Each gets a fundamentally different search strategy, ranking heuristic, and
DOCX emphasis (which sections expand vs abbreviate).
NO LLM CALLS. Pure rule-based routing.
Usage:
python sub_use_case_router.py --sub-use-case novelty --jurisdictions "" --risk strict --known-art "US10000000B2"
python sub_use_case_router.py --sub-use-case fto --jurisdictions "US,EP" --risk strict
python sub_use_case_router.py --sample
"""
import argparse
import json
import sys
from typing import Any, Dict, List, Optional
VALID_SUB_USE_CASES = ["novelty", "fto", "landscape", "diligence", "litigation"]
VALID_RISK = ["strict", "signal-gathering"]
# Strategy templates per sub-use-case
STRATEGIES = {
"novelty": {
"query_count": 6,
"sources": ["google_patents", "espacenet"],
"filters": {"date_filter": "any", "active_only": False},
"queries": [
{"type": "narrow_keyword", "count": 3, "source": "google_patents"},
{"type": "broad_concept", "count": 2, "source": "google_patents+espacenet"},
{"type": "cpc_class", "count": 1, "source": "google_patents", "after_initial": True},
],
"ranking_heuristic": "claim_text_overlap_with_invention_description",
"verdict_scale": ["NOVEL", "POTENTIALLY NOVEL", "NOT NOVEL"],
"docx_emphasis": {
"executive_summary": "expanded",
"closest_prior_art": "expanded",
"patent_landscape": "abbreviated",
"citation_graph_signals": "if_lens_only",
"geographic_coverage": "abbreviated",
"fto_flags": "skip",
"strategy_recommendations": "claim_differentiation_focus",
"audit_log": "standard",
},
"legal_disclaimer_mandatory": True,
},
"fto": {
"query_count": 12, # scales with jurisdiction count
"sources": ["google_patents", "espacenet", "uspto"],
"filters": {"date_filter": "priority_lt_today", "active_only": True, "jurisdiction_filtered": True},
"queries": [
{"type": "jurisdiction_filtered", "count": "2-3 per jurisdiction"},
{"type": "active_status_filter", "applied_to_all": True},
{"type": "cpc_class", "count": 1, "after_initial": True},
],
"ranking_heuristic": "claim_by_claim_infringement_risk",
"verdict_scale": ["CLEAR (per jurisdiction)", "FLAGGED", "HIGH RISK"],
"docx_emphasis": {
"executive_summary": "expanded",
"closest_prior_art": "abbreviated",
"patent_landscape": "abbreviated",
"citation_graph_signals": "if_lens_only",
"geographic_coverage": "expanded",
"fto_flags": "expanded_main_section",
"strategy_recommendations": "design_around_jurisdiction_focus",
"audit_log": "standard",
},
"legal_disclaimer_mandatory": True,
},
"landscape": {
"query_count": 9,
"sources": ["google_patents", "espacenet", "lens"],
"filters": {"date_filter": "10_year_window"},
"queries": [
{"type": "broad_technology", "count": "2-3"},
{"type": "cpc_class_extraction", "count": 1, "after_initial": True},
{"type": "per_top_filer", "count": "1 per top-5 filer"},
{"type": "lens_citation_graph", "count": "if_byok_available"},
],
"ranking_heuristic": "filer_count_plus_recency",
"verdict_scale": ["CONCENTRATED", "COMPETITIVE", "EMERGING"],
"docx_emphasis": {
"executive_summary": "standard",
"closest_prior_art": "abbreviated",
"patent_landscape": "expanded_main_section",
"citation_graph_signals": "expanded_if_lens",
"geographic_coverage": "expanded",
"fto_flags": "skip",
"strategy_recommendations": "who_to_watch_focus",
"audit_log": "standard",
},
"legal_disclaimer_mandatory": False,
},
"diligence": {
"query_count": 10,
"sources": ["google_patents", "uspto"],
"filters": {"date_filter": "any", "assignee_focused": True},
"queries": [
{"type": "assignee_search", "count": "2-3"},
{"type": "subsidiary_search", "count": "if_org_chart_provided"},
{"type": "inventor_search", "count": "for_key_inventors"},
{"type": "assignment_recordation", "count": "for_ownership_verification"},
{"type": "family_resolution", "applied_to_all": True},
],
"ranking_heuristic": "family_grouped_then_citation_count",
"verdict_scale": ["PORTFOLIO VERIFIED", "PARTIAL VERIFICATION", "OWNERSHIP RISK"],
"docx_emphasis": {
"executive_summary": "expanded",
"closest_prior_art": "abbreviated",
"patent_landscape": "expanded_as_portfolio_table",
"citation_graph_signals": "if_lens_only",
"geographic_coverage": "expanded",
"fto_flags": "skip",
"strategy_recommendations": "red_flags_in_portfolio",
"audit_log": "standard",
},
"legal_disclaimer_mandatory": False,
},
"litigation": {
"query_count": 7,
"sources": ["google_patents", "espacenet", "lens"],
"filters": {"date_filter": "before_target_priority_date"},
"queries": [
{"type": "fetch_target_patent", "extract": ["priority_date", "claims", "cpc_classes"]},
{"type": "cpc_class_with_date_filter", "count": 2},
{"type": "claim_language_with_date_filter", "count": 2},
{"type": "lens_forward_citations", "count": "if_byok_available"},
],
"ranking_heuristic": "knock_out_potential_claim_by_claim",
"verdict_scale": ["KNOCK-OUT FOUND", "STRONG OBVIOUSNESS COMBINATION", "WEAK OBVIOUSNESS", "NO MATERIAL ART"],
"docx_emphasis": {
"executive_summary": "expanded",
"closest_prior_art": "expanded_as_knock_out_candidates",
"patent_landscape": "abbreviated",
"citation_graph_signals": "expanded_if_lens",
"geographic_coverage": "abbreviated",
"fto_flags": "skip",
"strategy_recommendations": "per_claim_invalidity_analysis",
"audit_log": "standard",
},
"legal_disclaimer_mandatory": False,
},
}
def route(sub_use_case: str, jurisdictions: List[str], risk: Optional[str], known_art: Optional[str]) -> Dict[str, Any]:
if sub_use_case not in STRATEGIES:
raise ValueError(f"Invalid sub-use-case '{sub_use_case}'. Pick from: {list(STRATEGIES.keys())}")
strategy = STRATEGIES[sub_use_case].copy()
strategy["sub_use_case"] = sub_use_case
strategy["jurisdictions_input"] = jurisdictions
strategy["risk_input"] = risk
strategy["known_art_input"] = known_art
notes: List[str] = []
# FTO scales query count with jurisdictions
if sub_use_case == "fto" and jurisdictions:
per_jurisdiction = 3
strategy["query_count"] = len(jurisdictions) * per_jurisdiction + 2 # + CPC + active filter
notes.append(f"FTO query count scaled to {strategy['query_count']} for {len(jurisdictions)} jurisdiction(s)")
# Risk modifies ranking
if risk == "strict":
notes.append("Strict risk: aggressive ranking; surface verdict-grade hits only")
elif risk == "signal-gathering":
notes.append("Signal-gathering risk: prioritize breadth + visualization over verdict")
# Known art enables anchored search
if known_art and known_art.lower() != "none":
notes.append(f"Known art anchor: {known_art} — adjacent searches will reference this hit")
# Lens.org availability check (not asked here; flag in audit only)
notes.append("Lens.org BYOK: required for Citation Graph section. Check at runtime.")
if strategy["legal_disclaimer_mandatory"]:
notes.append("LEGAL DISCLAIMER MANDATORY: include in DOCX Sections 1, 7, 8")
strategy["operational_notes"] = notes
return strategy
def render_human(result: Dict[str, Any]) -> str:
out: List[str] = []
out.append(f"Sub-use-case: {result['sub_use_case']}")
out.append(f"Jurisdictions: {result.get('jurisdictions_input', []) or '(N/A for this sub-use-case)'}")
out.append(f"Risk tolerance: {result.get('risk_input', '(not specified)')}")
out.append(f"Known art: {result.get('known_art_input', '(none)')}")
out.append("")
out.append(f"Total query count: {result['query_count']}")
out.append(f"Sources: {', '.join(result['sources'])}")
out.append(f"Filters: {result['filters']}")
out.append("")
out.append("Query plan:")
for q in result["queries"]:
out.append(f" - {q}")
out.append("")
out.append(f"Ranking heuristic: {result['ranking_heuristic']}")
out.append(f"Verdict scale: {' / '.join(result['verdict_scale'])}")
out.append(f"Legal disclaimer mandatory: {result['legal_disclaimer_mandatory']}")
out.append("")
out.append("DOCX section emphasis:")
for section, emphasis in result["docx_emphasis"].items():
out.append(f" {section:<30s} {emphasis}")
out.append("")
if result.get("operational_notes"):
out.append("Operational notes:")
for n in result["operational_notes"]:
out.append(f" - {n}")
return "\n".join(out)
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--sub-use-case", choices=VALID_SUB_USE_CASES)
parser.add_argument("--jurisdictions", help="Comma-separated jurisdiction codes (US,EP,CN,JP,KR,PCT,worldwide)")
parser.add_argument("--risk", choices=VALID_RISK)
parser.add_argument("--known-art", help="Patent number or paper citation if user has seen prior art")
parser.add_argument("--sample", action="store_true", help="Run sample (FTO with US+EP jurisdictions, strict risk)")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
if args.sample:
result = route("fto", ["US", "EP"], "strict", "US10000000B2")
elif args.sub_use_case:
jurisdictions = [j.strip() for j in args.jurisdictions.split(",") if j.strip()] if args.jurisdictions else []
try:
result = route(args.sub_use_case, jurisdictions, args.risk, args.known_art)
except ValueError as e:
print(f"error: {e}", file=sys.stderr); return 2
else:
parser.print_help(); return 0
if args.output == "json":
print(json.dumps(result, indent=2, default=str))
else:
print(render_human(result))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))

View file

@ -0,0 +1,15 @@
{
"name": "syllabus",
"description": "Generates a curated supplementary reading list from any course syllabus using Consensus academic search. Grill-me intake (syllabus input format + course audience + year range) plus a grouping forcing-options checkpoint before any search runs — so the reading list matches the course's level and recency need. Parses the syllabus to extract topics and learning outcomes, searches Consensus for recent peer-reviewed papers per topic, and produces a professionally formatted .docx with clickable Consensus links, plain-language summaries calibrated to audience level, and Bloom-higher-order discussion questions tied to course learning goals. Triggers whenever a user uploads a syllabus, course outline, or curriculum document and wants supplementary readings. Also triggers on: 'syllabus reading list', 'find papers for my course', 'create a reading list from this syllabus', 'recent research for my class', 'supplementary readings', 'find journal articles for these topics', 'what recent papers cover this material', 'any new research on these course topics', 'update my syllabus with recent papers'. Even casual mentions when a syllabus is attached should trigger this skill.",
"version": "1.0.0",
"author": {"name": "Alireza Rezvani", "url": "https://alirezarezvani.com"},
"homepage": "https://github.com/alirezarezvani/claude-skills/tree/main/research/syllabus",
"repository": "https://github.com/alirezarezvani/claude-skills",
"license": "MIT",
"skills": ["./skills/syllabus"],
"source": {
"spec": "megaprompts/10-syllabus-megaprompt.md",
"build_pattern": "Path B (direct conversion). Research-pack shape, BUNDLED-JS-DOCX-GENERATOR variant — ships scripts/generate_reading_list.js for 300+ line DOCX assembly logic (token-efficient: skill doesn't re-derive layout each run).",
"sibling_of": "research/litreview, research/grants, research/patent, research/dossier, research/pulse"
}
}

View file

@ -0,0 +1,68 @@
# syllabus
Course supplementary reading list generator. Takes any course syllabus (PDF / DOCX / text / image) and produces a curated `.docx` of recent peer-reviewed papers via Consensus search, with plain-language summaries calibrated to audience level + Bloom-higher-order discussion questions tied to learning outcomes.
## What this skill does
1. **Phase 0 grill-me** (3 forcing Qs): syllabus input format + course audience + year range
2. **Phase 1**: parse syllabus (PDF/DOCX/text/image) → extract topics + learning outcomes
3. **Phase 2**: group topics into 6-12 sections + grill-me forcing checkpoint (proceed/merge/split/add/remove)
4. **Phase 3**: targeted Consensus searches per section (1-2 queries each, sequential at 1 q/sec, **applied-domain weaving**)
5. **Phase 4**: write summaries (audience-calibrated jargon) + discussion questions (Bloom higher-order)
6. **Phase 5**: generate .docx via **bundled JS script** (`scripts/generate_reading_list.js`)
7. **Phase 6**: deliver file + audit summary
## Architectural pattern: bundled JS
This skill uses a **bundled JavaScript helper script** (`scripts/generate_reading_list.js`) for DOCX generation rather than inlining the 300+ lines of layout code in SKILL.md. Rationale:
- DOCX generation logic is reusable + complex
- Better separation of concerns: skill = orchestration + intelligence; script = mechanical document assembly
- Token-efficient: skill doesn't re-derive layout each run
- Easier to maintain and version
The skill orchestrates the pipeline and invokes the script with JSON input.
## Sibling skill relationship
Part of the **research pack** (sibling of `pulse`, `litreview`, `grants`, `patent`, `dossier`). Shares Agent Integrity Rules. Adds:
- **Audience calibration** — undergrad summaries define every term; grad summaries assume technical fluency
- **Applied-domain weaving** — search "X applications" not just "X" (boosts relevance dramatically)
- **Bundled JS script pattern** — first research-pack skill to use this layout
## Source spec
[`megaprompts/10-syllabus-megaprompt.md`](../../megaprompts/10-syllabus-megaprompt.md) (PR #657).
## Plugin layout
```
research/syllabus/
├── .claude-plugin/plugin.json
├── README.md
├── agents/cs-syllabus.md
├── commands/cs-syllabus.md
└── skills/syllabus/
├── SKILL.md
├── references/
│ ├── applied_domain_weaving.md ← search-quality canon (7+ sources)
│ ├── audience_calibration.md ← undergrad vs grad summary jargon (7+ sources)
│ └── bundled_script_pattern.md ← why bundle vs inline (7+ sources)
└── scripts/
├── citation_tracker.py ← stdlib: Consensus three-count + 1s sequential
├── topic_grouper.py ← stdlib: heuristic 6-12 section grouping
├── discussion_question_validator.py ← stdlib: Bloom higher-order quality check
└── generate_reading_list.js ← bundled Node.js: DOCX assembly (~300 lines)
```
## Dependencies
- **Consensus MCP** — Required for literature search
- **Node.js with `docx` package** — Required (`npm install docx`)
- **Bundled script**`scripts/generate_reading_list.js` (shipped with skill, not external)
- **File reading** — PDF reader / DOCX parser via pandoc / vision for images
## License
MIT.

View file

@ -0,0 +1,84 @@
---
name: cs-syllabus
description: Course supplementary reading list persona. Walks 3 forcing intake questions (syllabus input format + course audience + year range) before parsing. Halts at grouping checkpoint after Phase 2 (proceed/merge/split/add/remove). Searches Consensus sequentially at 1 q/sec with applied-domain weaving (e.g., 'enzyme kinetics food processing' not just 'enzyme kinetics'). Calibrates summary jargon to audience (undergrad defines every term; grad assumes technical fluency). Writes Bloom higher-order discussion questions tied to learning outcomes. Generates .docx via bundled JS script.
skills: research/syllabus/skills/syllabus
domain: research
model: opus
tools: [Read, Write, Bash]
---
# Syllabus Agent
## Voice
**Opening:** "Drop your syllabus — file path, pasted text, or image. I'll grill you on audience and year range, parse the syllabus into 6-12 sections, halt for your confirmation, then search Consensus per section with applied-domain weaving."
**Refusing missing syllabus:** Q1 force; can't proceed without input.
**Audience calibration reminder (mid-Phase 4):**
> "Audience: Q2=undergrad-intro. Calibrating summaries to define jargon, not assume fluency. Discussion questions test analysis, not critique."
**Group-and-confirm checkpoint:**
> "Proposed sections: [list]. **Pick one:** proceed / merge X+Y / split X / add section for Y / remove X. This is the last cheap moment before search budget is consumed."
**Closing:**
> "Saved: <path>/reading_list_<course>_<date>.docx via bundled JS script. Audit: 12 searches × 47 papers / 22 cited. Plan tier: free (3/search). Sections: 8. Each paper has: hyperlinked title + audience-calibrated summary + Bloom-tied discussion question."
Sequential, audience-aware, applied-domain-weaving discipline.
## Purpose
The cs-syllabus agent orchestrates the `syllabus` skill across course-reading-list generation:
1. **Phase 0 intake** — Q1 input format, Q2 audience, Q3 year range
2. **Phase 1 parse** — PDF/DOCX/text/image → topics + learning outcomes
3. **Phase 2 group** — 6-12 sections + checkpoint
4. **Phase 3 search** — Consensus sequential 1 q/sec with applied-domain angle
5. **Phase 4 write** — audience-calibrated summaries + Bloom higher-order questions
6. **Phase 5 generate** — bundled JS DOCX
7. **Phase 6 deliver** — file + audit summary
**Hard rules:**
1. **One intake Q per turn.** Never bundle.
2. **Refuse missing syllabus** at Q1.
3. **Halt at grouping checkpoint.** No Phase 3 without explicit user choice.
4. **Sequential Consensus.** 1 q/sec.
5. **Applied-domain weaving** on every query (not "enzyme kinetics" alone — "enzyme kinetics food processing").
6. **Audience-calibrated summaries.** Undergrad defines jargon; grad assumes fluency.
7. **Bloom higher-order discussion questions.** Apply / analyze / evaluate. NOT recall ("what did the authors find?").
8. **Source discipline.** Consensus-only; training knowledge labeled.
9. **Three-count tracking.** Sent / received / cited.
10. **Bundled JS for DOCX.** Don't inline.
## Skill Integration
**Skill Location:** `../skills/syllabus/`
### Python Tools (Stdlib)
1. **Citation Tracker**`scripts/citation_tracker.py` — Consensus three-count + 1s sequential at `~/.syllabus_sessions/<session>.json`
2. **Topic Grouper**`scripts/topic_grouper.py` — heuristic 6-12 section grouping from extracted topics
3. **Discussion Question Validator**`scripts/discussion_question_validator.py` — Bloom higher-order quality check (rejects recall questions)
### Bundled Node.js Script
**Generate Reading List** — `scripts/generate_reading_list.js` — JSON-input → .docx output. ~300 lines. Handles `docx` package require with multi-location fallback. Uses `ExternalHyperlink` with full Consensus URLs (never truncated). `LevelFormat.BULLET` for lists.
### Knowledge Bases
- `references/applied_domain_weaving.md` — search-quality canon (7+ sources)
- `references/audience_calibration.md` — undergrad vs grad summary jargon (7+ sources)
- `references/bundled_script_pattern.md` — why bundle vs inline (7+ sources)
## Related Agents
- [cs-litreview](../../litreview/agents/cs-litreview.md) — sibling, academic literature
- [cs-grants](../../grants/agents/cs-grants.md) — sibling, NIH funding
- [cs-patent](../../patent/agents/cs-patent.md) — sibling, patent prior-art
- [cs-dossier](../../dossier/agents/cs-dossier.md) — sibling, entity research
---
**Version:** 1.0.0
**Source:** Path-B direct conversion of `megaprompts/10-syllabus-megaprompt.md`

View file

@ -0,0 +1,143 @@
---
name: "cs-syllabus"
description: "/cs:syllabus <syllabus-file-or-paste> — Generate curated supplementary reading list from any course syllabus. 3-Q grill-me (input format + audience + year range) + grouping checkpoint → Consensus searches per section with applied-domain weaving → .docx via bundled JS script with audience-calibrated summaries + Bloom higher-order discussion questions."
---
# /cs:syllabus — Course Supplementary Reading List
**Command:** `/cs:syllabus <syllabus-file-or-paste>`
The `cs-syllabus` persona produces a `.docx` reading list of recent peer-reviewed research per course section.
## When to Run
- Adding supplementary readings to an existing course
- Updating a syllabus with current research
- Checking what's recent in your field for course planning
- Even casual mentions when a syllabus is attached
## Forcing Intake (3 Questions, One at a Time)
| Q | Asks | Notes |
|---|---|---|
| Q1 | Syllabus input: file path / pasted content / image | refuses missing syllabus |
| Q2 | Course audience: undergrad-intro / undergrad-advanced / grad-masters / grad-doctoral / professional / mixed | drives summary jargon + discussion-question complexity |
| Q3 | Year range: 1 / 2 / 5 years | drives `year_min` on every Consensus search; default 2 |
## What You Get
```
reading_list_<course-slug>_<YYYY-MM-DD>.docx
Structure:
- Title page (course name, subtitle, date)
- Introduction (with Consensus app link)
- Course Learning Outcomes (boxed section)
- Sections (6-12, from grouping):
Each section = numbered papers, each with:
- Clickable hyperlinked title
- Author / journal / year (italic)
- Summary (plain language, audience-calibrated)
- Discussion Question (Bloom higher-order, tied to learning outcome)
- Footer (generation metadata)
```
## Grouping Checkpoint (After Phase 2)
After parsing the syllabus, the skill **halts** with a forcing-options prompt:
```
Proposed sections: [list with item counts]. Pick one:
1. Looks good — proceed with these sections
2. Merge sections [X] and [Y]
3. Split section [X] into two
4. Add a section for [topic]
5. Remove section [X]
```
This is the last cheap moment to correct course before search budget is consumed. **Refuses to start Phase 3 without explicit user choice.**
## Discipline
- **One intake Q per turn.** Never bundle.
- **Halt at grouping checkpoint.** No Phase 3 without user.
- **Sequential Consensus.** 1 q/sec.
- **Applied-domain weaving** — search "enzyme kinetics food processing" not just "enzyme kinetics". Boosts paper relevance dramatically.
- **Audience-calibrated summaries** — undergrad-intro defines every term; grad-doctoral assumes technical fluency.
- **Bloom higher-order discussion questions** — apply / analyze / evaluate. NOT recall ("what did the authors find?").
- **Source discipline** — only Consensus session results. Training knowledge labeled.
- **Three-count tracking** — sent / received / cited.
- **Bundled JS DOCX generator** — don't inline 300 lines of layout code.
## Quality Bars
### Summary
| ✅ Good | ❌ Bad |
|---|---|
| "This review maps how different diets — Mediterranean, Nordic, vegetarian — reshape the types of fat molecules circulating in your blood, with implications for heart disease risk." | "This paper reviews lipidomic profiles across dietary interventions and their cardiometabolic implications." (Too jargon-heavy) |
### Discussion Question
| ✅ Good | ❌ Bad |
|---|---|
| "If dietary fat quality can reshape your lipoprotein lipidome, what does this suggest about the biochemical basis for dietary guidelines recommending unsaturated over saturated fats?" | "What did the authors find?" (Just recall) |
## Workflow
```bash
# Phase 0 intake (Q1-Q3)
python ../skills/syllabus/scripts/citation_tracker.py --action start --session NAME
# Phase 1 parse (PDF/DOCX/text/image-appropriate reader)
# Phase 2 group + CHECKPOINT (wait for user)
python ../skills/syllabus/scripts/topic_grouper.py --topics-file /tmp/topics.json
# Phase 3 search (sequential Consensus 1 q/sec, applied-domain weaving)
# Phase 4 write summaries + discussion questions
python ../skills/syllabus/scripts/discussion_question_validator.py --questions-file /tmp/qs.json
# Phase 5 generate .docx via bundled script
node ../skills/syllabus/scripts/generate_reading_list.js \
--input /tmp/data.json \
--output /path/to/reading_list_<course>_<date>.docx
# Phase 6 deliver
python ../skills/syllabus/scripts/citation_tracker.py --action close --session NAME
```
## Trigger Phrases
- "syllabus reading list"
- "find papers for my course"
- "create a reading list from this syllabus"
- "recent research for my class"
- "supplementary readings"
- "find journal articles for these topics"
- "what recent papers cover this material"
- "any new research on these course topics"
- "update my syllabus with recent papers"
- Casual mentions when syllabus is attached
## Anti-Patterns Rejected
- Parallelizing Consensus calls (rate limit)
- Searching topics without applied-domain angle (poor relevance)
- Padding sections with fabricated entries when Consensus thin
- Generic discussion questions ("What did the authors find?")
- Jargon-heavy summaries unsuitable for course audience
- Skipping group-and-confirm step
- Truncating Consensus URLs in hyperlinks
- Inlining 300 lines of docx-generation JavaScript in skill body
## Related
- Agent: [`cs-syllabus`](../agents/cs-syllabus.md)
- Skill: [`syllabus`](../skills/syllabus/SKILL.md)
- Source spec: [`megaprompts/10-syllabus-megaprompt.md`](../../../megaprompts/10-syllabus-megaprompt.md)
- Siblings: `/cs:litreview`, `/cs:grants`, `/cs:patent`, `/cs:dossier`, `/cs:pulse`
---
**Version:** 1.0.0
**Source:** Path-B direct conversion of `megaprompts/10-syllabus-megaprompt.md`

View file

@ -0,0 +1,293 @@
---
name: syllabus
description: "Generates a curated supplementary reading list from any course syllabus using Consensus academic search. Grill-me intake (syllabus input format + course audience + year range) plus a grouping forcing-options checkpoint before any search runs — so the reading list matches the course's level and recency need. Parses the syllabus to extract topics and learning outcomes, searches Consensus for recent peer-reviewed papers per topic, and produces a professionally formatted .docx with clickable Consensus links, plain-language summaries calibrated to audience level, and Bloom-higher-order discussion questions tied to course learning goals. Triggers whenever a user uploads a syllabus, course outline, or curriculum document and wants supplementary readings. Also triggers on: 'syllabus reading list', 'find papers for my course', 'create a reading list from this syllabus', 'recent research for my class', 'supplementary readings', 'find journal articles for these topics', 'what recent papers cover this material', 'any new research on these course topics', 'update my syllabus with recent papers'. Even casual mentions when a syllabus is attached should trigger this skill."
license: MIT
metadata:
source_spec: "megaprompts/10-syllabus-megaprompt.md"
build_pattern: "Path B (direct conversion)"
research_pack_convention: "Agent Integrity Rules verbatim per PR #657 audit; bundled-JS-DOCX-generator variant"
version: 1.0.0
---
# Syllabus — Course Supplementary Reading List
> **Portability:** Requires a Consensus MCP connection, Node.js with `docx` package, and file reading capability for the syllabus. Works in Claude Code CLI natively. In Claude.ai with Consensus MCP + Code Execution + file upload, the workflow is supported.
For an instructor or student with a course syllabus, produce a professional supplementary reading list as `.docx` containing recent peer-reviewed papers per course section.
## Architectural Pattern: Bundled Script
This skill uses a **bundled JavaScript helper script** for DOCX generation rather than inlining the 300+ lines of layout code:
- DOCX generation logic is reusable + complex
- Better separation of concerns: skill = orchestration + intelligence; script = mechanical document assembly
- Token-efficient: skill doesn't re-derive layout each run
- Easier to maintain and version
The bundled script is at `scripts/generate_reading_list.js`. The skill orchestrates the pipeline + invokes the script with JSON input.
## Agent Integrity Rules (Research-Pack Convention)
Locked verbatim per PR #657 audit.
- **Only use what Consensus returns.** Every paper title, author, journal, year, URL must come from this session's tool calls. Training-knowledge papers labeled `[Not from Consensus — model knowledge]` and excluded.
- **Confirm before moving on.** A search isn't complete until response received and inspected.
- **Track three counts.** Queries sent / papers received / papers cited. Surface in audit summary.
- **Surface gaps, don't fill them.** Section with one paper + note about limited results > section padded with fabrications.
## Phase 0: Grill-Me Intake (3 forcing questions)
### Q1 (root) — Syllabus input
> **Provide the syllabus — pick one:**
>
> 1. File path (PDF, DOCX, text) — I'll read it
> 2. Pasted content — paste below
> 3. Image of a printed syllabus — attach the image
>
> *Why I'm asking:* Each format needs a different reader (PDF / DOCX parser / vision). Picking upfront prevents wasted attempts.
Forcing choice. Refuse to start without a syllabus.
### Q2 (depends on Q1) — Course audience
> **Course audience — pick one:**
>
> 1. Undergraduate (intro level)
> 2. Undergraduate (advanced / upper division)
> 3. Graduate (Masters / early PhD)
> 4. Graduate (doctoral / advanced)
> 5. Professional / continuing education
> 6. Mixed
>
> *Why I'm asking:* Audience dictates summary jargon level and discussion-question complexity. Undergrad summaries define every term; grad summaries assume technical fluency. Discussion questions for undergrads test analysis; for grads test critique and extension.
See [`references/audience_calibration.md`](references/audience_calibration.md) for the canon.
### Q3 (depends on Q1) — Year range
> **Year range for papers — pick one:**
>
> 1. Last 1 year (most recent only)
> 2. Last 2 years (default — recent + a year of context)
> 3. Last 5 years (broader, includes foundational recent work)
>
> *Why I'm asking:* Reading lists go stale fast. 1-year filters keep things fresh; 5-year filters surface foundational recent work that's already standard. Drives the year_min parameter on every Consensus search.
Forcing choice with default (last 2 years).
**Stop condition:** 3 questions max before Phase 1. The post-Phase-2 group-and-confirm checkpoint is its own grill-me moment.
## Phase 1: Parse the Syllabus
Per Q1 input format:
- **PDF**: use PDF reader; extract text
- **DOCX**: use pandoc or DOCX parser; extract text
- **Text/pasted**: read directly
- **Image**: use vision; extract text
From extracted text:
1. Course title + instructor + term
2. Topic list (lecture titles, week-by-week breakdown, etc.)
3. Learning outcomes (if explicit; if missing, infer 3-5 from description)
Mark inferred learning outcomes as `[inferred]` in the DOCX.
## Phase 2: Group Topics + Confirm with User
### Group via topic_grouper.py
Use `scripts/topic_grouper.py` to cluster related topics into 6-12 sections. Heuristic: closely-related topics merge; cross-cutting topics get their own section.
### Group-and-Confirm Checkpoint (Forcing Options)
After grouping, present:
> **Proposed sections: [list with item counts]. Pick one:**
>
> 1. "Looks good — proceed with these sections"
> 2. "Merge sections [X] and [Y]"
> 3. "Split section [X] into two"
> 4. "Add a section for [topic]"
> 5. "Remove section [X]"
>
> *Why I'm asking:* Grouping drives search allocation. Wrong grouping wastes the search budget on bad clusters. This is the **last cheap moment** to correct course before searches consume Consensus calls.
**Refuse to start Phase 3 without explicit user choice.**
## Phase 3: Search Consensus per Section
Sequential, 1 q/sec. 1-2 queries per section.
### Applied-Domain Weaving (Critical)
Don't just search the topic — **search the topic + applied domain**:
| ❌ Generic | ✅ Applied-domain |
|---|---|
| "enzyme kinetics" | "enzyme kinetics food processing applications" |
| "machine learning" | "machine learning clinical decision support" |
| "thermodynamics" | "thermodynamics renewable energy systems" |
| "social network analysis" | "social network analysis public health interventions" |
Boosts paper relevance dramatically. See [`references/applied_domain_weaving.md`](references/applied_domain_weaving.md) for the canon.
### Per-Section Pattern
```
For each section:
1. Construct query: "{topic-keywords} {applied-domain-angle}" + year_min from Q3
2. Submit to Consensus (sequential, 1 q/sec gap enforced by citation_tracker)
3. Receive results
4. (If thin) submit one fallback query without applied-domain angle
5. Select 1-3 papers per section (15-25 total across all sections)
```
### Selection Priorities
1. **Relevance** — paper directly addresses the section topic
2. **Reviews / meta-analyses** — synthesize the field
3. **Citation count** — established work
4. **Applied-domain connection** — tied to the course's domain (e.g., engineering vs theory)
## Phase 4: Write Summaries + Discussion Questions
### Summary writing
Per paper:
- Plain language (calibrated to audience from Q2)
- 2-3 sentences
- Define jargon if undergraduate audience; assume fluency if graduate
### Quality bars
| ✅ Good summary | ❌ Bad summary |
|---|---|
| "This review maps how different diets — Mediterranean, Nordic, vegetarian — reshape the types of fat molecules circulating in your blood, with implications for heart disease risk." | "This paper reviews lipidomic profiles across dietary interventions and their cardiometabolic implications." |
### Discussion question writing
Per paper:
- Bloom **higher-order** (apply / analyze / evaluate)
- Tied to a specific course learning outcome
- Promotes discussion, not just recall
| ✅ Good question | ❌ Bad question |
|---|---|
| "If dietary fat quality can reshape your lipoprotein lipidome, what does this suggest about the biochemical basis for dietary guidelines recommending unsaturated over saturated fats?" | "What did the authors find?" (Just recall) |
Use `scripts/discussion_question_validator.py` to flag recall-only questions.
## Phase 5: Generate .docx via Bundled Script
```bash
node ../scripts/generate_reading_list.js \
--input /tmp/syllabus_data.json \
--output /path/to/reading_list_<course>_<date>.docx
```
The script accepts JSON with this schema:
```json
{
"courseTitle": "string",
"courseSubtitle": "string",
"generatedDate": "string",
"yearRange": "string",
"introText": "string",
"learningOutcomes": ["string", ...],
"sections": [
{
"heading": "string",
"papers": [
{
"title": "string",
"authors": "string",
"journal": "string",
"year": number,
"url": "string",
"summary": "string",
"question": "string"
}
]
}
],
"auditLog": {
"totalQueriesSent": number,
"totalPapersReceived": number,
"totalPapersCited": number,
"toolConstraints": "string",
"searchDetails": [
{
"section": "string",
"query": "string",
"papersReturned": number,
"papersSelected": number,
"status": "string"
}
],
"failures": []
}
}
```
The script handles:
- `docx` package require with multi-location fallback
- Title page, intro with Consensus link, learning outcomes box, numbered papers per section
- `ExternalHyperlink` with full Consensus URLs (never truncated)
- `LevelFormat.BULLET` for lists (not unicode bullets)
- Footer with generation metadata
- Input validation (missing fields → graceful error)
See [`references/bundled_script_pattern.md`](references/bundled_script_pattern.md) for why bundled vs inline.
## Phase 6: Deliver
- File path
- Audit summary in chat: "Saved {file}. {N} sections × {M} papers / {K} cited. Plan tier: {tier}."
- Validate: `python scripts/office/validate.py <docx>`
## Tooling
| Script | Role |
|---|---|
| `scripts/citation_tracker.py` | Consensus three-count audit + 1s sequential discipline at `~/.syllabus_sessions/<session>.json` |
| `scripts/topic_grouper.py` | Heuristic 6-12 section grouping from extracted topics |
| `scripts/discussion_question_validator.py` | Bloom higher-order quality check; flags recall-only questions |
| `scripts/generate_reading_list.js` | **Bundled Node.js DOCX generator** — JSON input → .docx output |
## References
- [`references/applied_domain_weaving.md`](references/applied_domain_weaving.md) — search-quality canon (7+ sources)
- [`references/audience_calibration.md`](references/audience_calibration.md) — undergrad vs grad summary jargon (7+ sources)
- [`references/bundled_script_pattern.md`](references/bundled_script_pattern.md) — why bundle vs inline (7+ sources)
## Error Handling
| Failure | Behavior |
|---|---|
| Consensus rate-limit hit | Wait 3s, retry once, log |
| Search returns 0 for a section | Note section as "limited results — consider manual supplementation" |
| 3 consecutive failures | Stop, alert user, share collected so far |
| `docx` package not installed | Script attempts `npm install`; if still failing, fail with clear message |
| DOCX validation fails | Unpack XML, log issue, ask user to retry |
| Syllabus format unsupported | List supported formats, ask user to convert |
| Learning outcomes can't be extracted | Infer 3-5 from course description; mark as inferred in document |
## Anti-Patterns To Reject
- Parallelizing Consensus calls (rate limit)
- Searching topics without applied-domain angle (poor relevance)
- Padding sections with fabricated entries when Consensus returns thin
- Generic discussion questions ("What did the authors find?")
- Jargon-heavy summaries unsuitable for the course's audience level
- Skipping the group-and-confirm step (wastes searches)
- Truncating Consensus URLs in hyperlinks
- Inlining 300 lines of docx-generation JavaScript in the skill body (use bundled script)
---
**Version:** 1.0.0
**Source spec:** [`megaprompts/10-syllabus-megaprompt.md`](../../../../megaprompts/10-syllabus-megaprompt.md)
**Build pattern:** Path B (direct conversion). Bundled-JS-DOCX-generator variant.

View file

@ -0,0 +1,148 @@
# Applied-Domain Weaving — The Search-Quality Multiplier
This reference answers exactly one decision: **why does the syllabus skill always weave the applied domain into Consensus queries, and what makes a generic search produce thin results?**
## The Core Insight
A query like `"enzyme kinetics"` returns **review papers and theoretical treatments** — useful for a biochemistry course but unhelpful for a *food science* course where students need to know how enzyme kinetics applies to bread fermentation, cheese ripening, and meat tenderization.
The query `"enzyme kinetics food processing applications"` returns the SAME field but from the angle the course actually needs.
> **Applied-domain weaving = search the topic + the course's applied domain.**
This is the single highest-leverage technique in the skill. Boosts paper relevance dramatically — typically 3-5x more course-appropriate papers per query.
## Concrete Examples by Discipline
### Engineering / Applied Sciences
| Topic | Generic search | Applied-domain search |
|---|---|---|
| Thermodynamics | "thermodynamics" | "thermodynamics renewable energy systems" |
| Fluid mechanics | "fluid mechanics" | "fluid mechanics biomedical device design" |
| Control systems | "PID control" | "PID control HVAC building automation" |
| Materials science | "polymer composites" | "polymer composites aerospace structural" |
### Health Sciences
| Topic | Generic search | Applied-domain search |
|---|---|---|
| Pharmacology | "drug interactions" | "drug interactions pediatric oncology" |
| Public health | "social determinants" | "social determinants rural health disparities" |
| Nutrition | "lipid metabolism" | "lipid metabolism Mediterranean diet" |
| Immunology | "innate immunity" | "innate immunity vaccine development" |
### Computer Science / Data Science
| Topic | Generic search | Applied-domain search |
|---|---|---|
| Machine learning | "neural networks" | "neural networks medical imaging diagnosis" |
| Distributed systems | "consensus algorithms" | "consensus algorithms blockchain finance" |
| Database systems | "query optimization" | "query optimization warehouse analytics" |
| HCI | "user interface design" | "user interface design accessibility" |
### Business / Social Sciences
| Topic | Generic search | Applied-domain search |
|---|---|---|
| Game theory | "Nash equilibrium" | "Nash equilibrium auction design" |
| Behavioral econ | "loss aversion" | "loss aversion retirement savings" |
| Org psychology | "team dynamics" | "team dynamics remote engineering" |
| Marketing | "consumer behavior" | "consumer behavior subscription services" |
### Physical Sciences
| Topic | Generic search | Applied-domain search |
|---|---|---|
| Quantum mechanics | "entanglement" | "entanglement quantum computing applications" |
| Astrophysics | "stellar evolution" | "stellar evolution exoplanet habitability" |
| Geology | "plate tectonics" | "plate tectonics earthquake hazard" |
## Why This Works
The applied-domain term:
1. **Filters Consensus to applied-research papers** — practical reviews, case studies, applied benchmarks
2. **Shifts citation network into your course's lineage** — papers other applied-domain researchers also cite
3. **Surfaces papers in the right journals** — domain-specific journals over pure-theory ones
4. **Gives papers students can connect to** — abstract theory → "I see how this matters"
## How to Identify the Applied Domain
The applied domain comes from one or more of:
1. **Course title** — "Food Science 301" → "food processing applications"
2. **Department / college** — Engineering → "engineering applications"
3. **Course description** — explicit "applied to X" / "for Y industry"
4. **Learning outcomes** — operational outcomes signal applied focus
If the syllabus is genuinely theoretical (e.g., a pure-math course), use **methodological angle** instead:
- Theoretical CS → "theoretical CS algorithm complexity"
- Pure math → "pure math applications" (or skip — pure-theory queries are fine here)
## When to Skip Applied-Domain Weaving
- **Pure theory courses** — no applied angle. Search topic only.
- **Survey courses** — broad coverage needed; applied-domain may narrow too much.
- **Topic genuinely doesn't have a natural applied domain** — e.g., "intro to research methods" — skip and search the topic + "review" or "introduction".
If applied-domain search returns < 3 papers, **fall back to generic search** for that section. Don't pad with fabrications.
## Operational Pattern
In Phase 3 of the skill:
```
For each section in [proposed sections]:
1. Construct primary query: "{topic} {applied-domain-keyword}" + year_min
2. Submit to Consensus (sequential, 1 q/sec gap)
3. If results >= 3: select papers, move on
4. If results < 3: submit fallback "{topic}" + year_min
5. Select 1-3 papers from combined results
```
## Anti-Patterns
### "Just search the topic"
Most common mistake. Produces theoretically rigorous but unhelpful papers for an applied course. Students can't connect them to course goals. Engagement drops.
### "Search the applied domain alone"
Without the topic anchor, query is too broad. "Food processing" returns 10,000+ papers across all subfields. Topic + applied-domain is the sweet spot.
### "Use multiple applied domains in one query"
"Enzyme kinetics food processing biomedical industrial applications" overconstrains. Each query targets ONE applied domain. If a section spans multiple domains, run separate queries.
### "Weave domain into queries even for pure-theory courses"
Pure-theory courses don't have applied domains. Forcing one in produces awkward queries that miss the actual theoretical literature.
### "Skip applied-domain weaving to save query budget"
The applied-domain weaving doesn't add queries — it modifies them. Same query budget, dramatically better relevance.
## Operational Checklist
- [ ] Course's applied domain identified (from title / department / description / learning outcomes)
- [ ] Each Phase 3 query: `{topic} + {applied-domain}` format
- [ ] Fallback to generic search if applied-domain returns < 3 papers
- [ ] Pure-theory courses: skip applied-domain weaving (use generic)
- [ ] Multi-domain sections: separate query per domain (don't stack in one query)
## Citations (7 sources)
1. **Bloom, B. S. (ed.), *Taxonomy of Educational Objectives* (1956).** Source for the application-tier of learning that justifies the applied-domain framing. Higher-tier learning (apply / analyze / evaluate) requires applied examples; pure-theory readings only support recall + comprehension.
2. **Mayer, R. E., *Multimedia Learning* (Cambridge, 2nd ed. 2009).** Empirical research on how applied examples accelerate learning vs abstract presentation. Source for the engagement-drop signal that pure-theory readings produce in applied courses.
3. **Fink, L. D., *Creating Significant Learning Experiences* (Jossey-Bass, 2003).** Source for the "integration" learning category — the discipline of connecting course content to students' applied contexts. Applied-domain weaving operationalizes this.
4. **Donald, J. G., *Learning to Think: Disciplinary Perspectives* (Jossey-Bass, 2002).** Empirical study of disciplinary thinking patterns. Justifies the per-discipline query-pattern table — engineering thinks differently from biology thinks differently from CS.
5. **Lave, J. & Wenger, E., *Situated Learning* (Cambridge, 1991).** Source for "situated cognition" — knowledge is best learned in the context of its application. Applied-domain weaving brings the readings into the situated context.
6. **Chickering, A. W. & Gamson, Z. F., "Seven Principles for Good Practice in Undergraduate Education" — *AAHE Bulletin*, 1987.** Principle #5 ("Emphasize Time on Task") + Principle #7 ("Respect Diverse Talents") favor applied-domain readings over pure-theory abstracts that don't connect to student backgrounds.
7. **Boyer, E. L., *Scholarship Reconsidered* (Carnegie Foundation, 1990).** Source for the "Scholarship of Application" framing. Applied-domain papers represent this scholarship category; weaving them into reading lists honors that scholarship.

View file

@ -0,0 +1,167 @@
# Audience Calibration — Undergrad vs Grad Summary Jargon + Question Complexity
This reference answers exactly one decision: **how does the syllabus skill calibrate summary jargon and discussion question complexity to the course's audience (Q2)?**
## The Core Rule
The same paper needs **different summaries** for different audiences:
- **Undergrad-intro**: define every technical term; assume zero prior knowledge
- **Undergrad-advanced**: assume foundational vocabulary; explain field-specific terms
- **Grad-Masters**: assume technical fluency; brief context for novel concepts
- **Grad-doctoral**: assume technical + methodological fluency; brief mention only of established context
Same paper, different summaries. Generic summaries miss the engagement target.
## Audience Buckets (Q2)
| Bucket | Vocabulary assumption | Method assumption | Discussion question complexity |
|---|---|---|---|
| Undergraduate (intro) | Zero specialized | Zero | Recall + comprehension + simple application |
| Undergraduate (advanced) | Foundational vocab | Common methods | Application + analysis |
| Graduate (Masters / early PhD) | Technical fluency | Common research methods | Analysis + evaluation |
| Graduate (doctoral / advanced) | Technical + methodological fluency | Methods specifics | Evaluation + critique + synthesis |
| Professional / continuing ed | Field-specific assumed | Methods context-dependent | Application to practice |
| Mixed | Lowest bucket present | Same | Same |
## Summary Calibration
### Undergrad-intro
Every technical term defined. Plain language. Connects to common experience.
| ❌ Too jargon | ✅ Calibrated |
|---|---|
| "This RCT compared lipidomic profiles across dietary interventions to assess cardiometabolic risk modulation." | "This randomized study compared what happens to fat molecules in the blood when people eat different diets — Mediterranean, Nordic, vegetarian — and looked at how those changes might affect heart disease risk." |
| "The phylogenetic analysis identified convergent evolution of toxin-resistant Na+ channels across reptilian lineages." | "Researchers compared sodium-channel genes across snake species and found that snakes from very different evolutionary branches independently developed similar resistance to toxic prey." |
### Undergrad-advanced
Foundational vocabulary assumed. Explain field-specific terms briefly.
| ❌ Too dumbed-down | ✅ Calibrated |
|---|---|
| "This randomized study compared what happens to fat molecules in the blood..." | "This RCT (n=240) tracked lipidomic shifts across three dietary patterns — Mediterranean, Nordic, vegetarian — over 12 weeks. Cardiometabolic markers improved most in the Mediterranean arm." |
| "Researchers compared sodium-channel genes..." | "Phylogenetic analysis across 47 reptilian lineages identifies convergent evolution of Na+ channel modifications conferring resistance to neurotoxic prey." |
### Grad (Masters or doctoral)
Technical fluency assumed. Brief context for novel concepts. Method specifics if relevant.
| ❌ Too verbose | ✅ Calibrated |
|---|---|
| "This RCT (n=240) tracked lipidomic shifts across three dietary patterns over 12 weeks. Cardiometabolic markers improved most in Mediterranean." | "RCT (n=240, 12-week, parallel-arm) comparing Mediterranean / Nordic / vegetarian. Mediterranean → 14% lower LDL-particle count, 22% lower oxidized LDL; differences plausibly mediated by MUFA:SFA ratio." |
| "Phylogenetic analysis across 47 reptilian lineages identifies convergent evolution..." | "Bayesian phylogenetic analysis (47 lineages, BEAST 2.7) supports independent emergence of Na+ channel S6-domain modifications in 6 lineages; convergence rate inconsistent with neutral drift (PP > 0.95)." |
### Professional / continuing ed
Field-specific terms assumed. Emphasize practice implications.
| ❌ Too academic | ✅ Calibrated |
|---|---|
| "RCT (n=240, 12-week)... LDL-particle count down 14%..." | "12-week RCT shows Mediterranean diet improves LDL-particle metrics 14-22% vs comparators. Practice implication: nutritional counseling for cardiovascular-risk patients should emphasize MUFA-rich foods specifically, not just 'low-fat'." |
## Discussion Question Calibration
Use Bloom's revised taxonomy (Anderson & Krathwohl 2001):
| Level | Action verbs | Question pattern |
|---|---|---|
| Remember | identify, list, recall | "What is X?" "Name the components" |
| Understand | explain, summarize, classify | "Why does X happen?" "How would you describe Y?" |
| Apply | use, apply, demonstrate | "How could this method be applied to...?" "What would happen if we used X for Y?" |
| Analyze | compare, contrast, examine | "What patterns connect X and Y?" "Why do X and Y produce different results?" |
| Evaluate | judge, critique, defend | "Is this study's conclusion warranted by its methods?" "Which approach better serves goal Z, and why?" |
| Create | design, propose, construct | "Design a study that would test the limits of X." "Propose a novel application of Y to Z." |
### Calibration by audience
| Audience | Question levels | Avoid |
|---|---|---|
| Undergrad-intro | Remember + Understand + simple Apply | Pure recall ("what did authors find?") |
| Undergrad-advanced | Understand + Apply + simple Analyze | Sophisticated Evaluate / Create |
| Grad-Masters | Apply + Analyze + Evaluate | Pure recall (insulting) |
| Grad-doctoral | Analyze + Evaluate + Create | Anything below Apply |
### Examples per audience
#### Undergrad-intro
| ❌ Recall only | ✅ Calibrated |
|---|---|
| "What did the authors find?" | "If you wanted to lower your heart disease risk through diet, what does this study suggest you should change?" (Apply) |
#### Grad-doctoral
| ❌ Below level | ✅ Calibrated |
|---|---|
| "What did this RCT show?" | "How would you redesign this RCT to test whether MUFA:SFA ratio specifically (vs total fat composition) drives the lipidomic shift?" (Create) |
## Discussion Question Validator
`scripts/discussion_question_validator.py` flags:
- **Recall-only questions** (any audience): "what did authors find?", "summarize", "describe"
- **Below-audience questions**: undergrad-intro questions in grad course → flag
- **Above-audience questions**: doctoral-level questions in undergrad-intro → flag
Validator suggests upgrades by replacing verbs with audience-appropriate Bloom verbs.
## Tying Discussion Questions to Learning Outcomes
Beyond audience calibration, each question should **explicitly tie to a learning outcome**:
| Without LO tie | With LO tie |
|---|---|
| "How could this approach be applied to...?" | "Course outcome 3 says students should be able to design enzymatic processes. How would the kinetics described in this paper inform a process design for cheese ripening?" |
The LO tie:
- Reinforces course goals
- Shows students why the reading matters
- Creates assessable discussion behaviors
If learning outcomes were inferred (`[inferred]`), still tie discussion questions to them — flag both as inferred.
## Anti-Patterns
### "Same summary for all audiences"
The biggest engagement killer. Undergrad summaries that read like graduate abstracts produce blank stares; graduate summaries that read like K-12 explainers feel patronizing.
### "Add jargon to look academic in undergrad summaries"
Engagement signal: students underline / highlight content. Jargon-heavy summaries get less highlighting in undergrad classes. Plain-language summaries get more.
### "Generic discussion questions"
"What did the authors find?" works for any audience — and serves none. The discussion question is the engagement hook; generic questions waste it.
### "All discussion questions at the highest Bloom level"
In a grad-doctoral course, even one Create-level question per paper is taxing. Mix Analyze, Evaluate, Create. Don't make every reading require students to design a follow-up study.
## Operational Checklist
- [ ] Q2 audience parsed → calibration bucket selected
- [ ] All summaries calibrated to bucket
- [ ] All discussion questions calibrated to bucket's Bloom range
- [ ] Each discussion question tied to a learning outcome (explicit or inferred)
- [ ] Validator (`discussion_question_validator.py`) run on all questions
- [ ] Recall-only questions rejected
- [ ] Below-audience or above-audience questions reworked
## Citations (7 sources)
1. **Bloom, B. S. (1956); Anderson, L. W. & Krathwohl, D. R. (2001), *A Taxonomy for Learning, Teaching, and Assessing*.** The revised Bloom's taxonomy. Source for the 6-level question hierarchy + action verb lexicon.
2. **Marzano, R. J. & Kendall, J. S., *The New Taxonomy of Educational Objectives* (Corwin, 2007).** Modern alternative to Bloom; emphasizes meta-cognitive and self-system levels. Source for the validator's "below-level vs above-level" distinction.
3. **Hattie, J., *Visible Learning* (Routledge, 2008/2023 update).** Meta-meta-analysis of educational interventions. Effect size 0.6+ for "teacher clarity" justifies the audience-calibrated summary discipline (clarity is audience-relative).
4. **Bain, K., *What the Best College Teachers Do* (Harvard, 2004).** Source for the "tied to learning outcome" discipline. Bain's research found great teachers connect every reading explicitly to course-level goals; generic readings produce engagement drop.
5. **Walvoord, B. E. & Anderson, V. J., *Effective Grading* (Jossey-Bass, 2nd ed. 2010).** Source for the "discussion question is assessable behavior" framing. Each discussion question = an opportunity to assess whether learning outcomes are being met.
6. **Brookfield, S. D. & Preskill, S., *Discussion as a Way of Teaching* (Jossey-Bass, 2nd ed. 2005).** Source for the engagement-vs-jargon trade-off in summary writing. Brookfield's research: students engage with content they can paraphrase; jargon-heavy summaries reduce paraphrase capability.
7. **Bjork, R. A. & Bjork, E. L., "Making Things Hard on Yourself, but in a Good Way" — *Psychology and the Real World* (FABBS Foundation, 2011).** Source for the "desirable difficulty" framing. Discussion questions should be challenging at the audience's edge, not below it (insulting) or above it (defeating).

View file

@ -0,0 +1,150 @@
# Bundled Script Pattern — Why JS for DOCX Generation, Not Inline
This reference answers exactly one decision: **why does the syllabus skill ship a bundled `generate_reading_list.js` script rather than inlining the DOCX generation logic in SKILL.md?**
## The Core Trade
DOCX generation requires ~300 lines of `docx`-package boilerplate (table layouts, hyperlink patterns, list formatting, page setup, etc.). This logic is:
1. **Reusable** across runs — every reading list uses the same DOCX layout
2. **Mechanical** — no LLM judgment required; just JSON-in / DOCX-out
3. **Long-lived** — the layout doesn't change between runs
Inlining 300 lines of mechanical layout code in SKILL.md means:
- The skill prompt is much longer (token cost on every invocation)
- Layout changes require editing the skill prompt (high-risk)
- The skill body has to re-derive the same logic each run
Bundling the logic in `scripts/generate_reading_list.js` means:
- The skill body is ~200 lines lighter (token-efficient)
- Layout changes are isolated to one file
- The skill orchestrates; the script executes mechanically
## When to Bundle (vs Inline)
### Bundle when:
- ✅ The logic is mechanical (no LLM judgment)
- ✅ The logic is reusable across runs (same layout / same algorithm)
- ✅ The logic is non-trivial (>50 lines)
- ✅ The logic is in a non-Python language (JS, Go, Rust, etc.)
- ✅ The logic has external dependencies (`docx` package, `requests`, etc.)
### Inline (in SKILL.md) when:
- The logic requires LLM judgment per run (e.g., paper-summary writing)
- The logic is short (<20 lines) and run-specific
- The logic is in-context-only (uses session-specific tool calls)
- The logic varies significantly per invocation
## The Pattern Used Here
`scripts/generate_reading_list.js`:
1. **Accepts JSON input + output path as CLI args**
```bash
node generate_reading_list.js --input data.json --output result.docx
```
2. **Has a documented JSON schema** (in SKILL.md so the orchestrator knows what to produce)
3. **Handles `docx` require with multi-location fallback** (works whether `docx` is installed locally, globally, or in a parent dir)
4. **Validates input** (missing fields → graceful error, not silent failure)
5. **Produces a clean professional DOCX** with:
- Title page
- Introduction (with Consensus link)
- Learning outcomes box
- Numbered papers per section
- Footer with metadata
6. **Uses canonical `docx` patterns**:
- `ExternalHyperlink` with full URLs
- `LevelFormat.BULLET` for lists
- Dual-width tables (`columnWidths` + cell `width`)
## Skill Orchestrator's Role
The skill body (SKILL.md):
1. Walks Phase 0 intake
2. Parses syllabus + extracts topics
3. Walks group-and-confirm checkpoint
4. Runs Consensus searches (LLM judgment per query)
5. Writes summaries + discussion questions (LLM judgment per paper)
6. **Constructs the JSON payload** matching the bundled script's schema
7. **Invokes the script** with the JSON
8. Validates output + delivers
The skill body is responsible for **what goes in the document**. The script is responsible for **how it's laid out**.
## Why Node.js Specifically
The `docx` library is a JavaScript library (npm package). Could the skill use a Python `docx` library (`python-docx`)? Yes, but:
- The repo's other research-pack DOCX-generating skills (litreview, grants, dossier) all use Node.js + `docx`
- Consistency: one DOCX library across the research pack
- The `docx` JS library is more actively maintained + has richer features
- `python-docx` doesn't support all the features the skill needs (advanced hyperlinks, table styling)
## File Structure
```
research/syllabus/skills/syllabus/scripts/
├── citation_tracker.py ← stdlib Python (orchestration helper)
├── topic_grouper.py ← stdlib Python (orchestration helper)
├── discussion_question_validator.py ← stdlib Python (orchestration helper)
└── generate_reading_list.js ← BUNDLED Node.js (mechanical DOCX assembly)
```
The Python scripts are stateless helpers (per-run). The JS script is the bundled mechanical assembler (called once per run).
## Anti-Patterns
### "Inline the JS into a Python script via subprocess"
Adds an unnecessary layer. The skill should call `node` directly.
### "Convert JS logic to Python to keep all scripts in one language"
Loses access to the better-maintained `docx` JS library. Worse: would diverge from sibling skills (litreview, grants, dossier all use `docx` JS).
### "Keep the JS script but inline the JSON schema in the script"
The JSON schema needs to be IN SKILL.md so the orchestrator knows what to construct. Documenting it in the script alone hides it from the orchestrator's prompt context.
### "Inline 300 lines of docx code in SKILL.md"
The original anti-pattern. Bloats the prompt, makes layout changes risky, makes the skill body harder to read.
### "Import the script from another skill"
Cross-skill dependencies break the per-skill self-contained discipline (per CLAUDE.md anti-patterns). Even though it would save duplication, the bundled script lives within syllabus's own folder.
## Operational Checklist
- [ ] `scripts/generate_reading_list.js` exists in syllabus's scripts/ folder
- [ ] Script accepts `--input <json>` + `--output <docx>` CLI args
- [ ] Script handles `docx` require with multi-location fallback
- [ ] Script validates input (missing fields → graceful error)
- [ ] JSON schema documented in SKILL.md (not just in the script)
- [ ] Skill orchestrator constructs JSON matching the schema
- [ ] Skill orchestrator invokes the script via `node` (not `python`)
- [ ] DOCX output validated post-generation
## Citations (7 sources)
1. **Karpathy-coder discipline + write-a-skill conventions** (this repo's `engineering/write-a-skill/`). Source for the "stdlib-only Python tools, bundled non-Python scripts allowed for mechanical jobs" pattern.
2. **CLAUDE.md anti-pattern: "Don't add features beyond what the task requires."** The bundled script honors this — it does ONE thing (DOCX layout) and does it mechanically.
3. **`docx` Node.js package — github.com/dolanmiu/docx (MIT).** Authoritative source for the API patterns the bundled script uses. Active maintenance, comprehensive feature set.
4. **CommonJS / Node.js module resolution algorithm.** Source for the "multi-location fallback" pattern in the require statement. Ensures the script works in development (local node_modules) and production (global install).
5. **Twelve-Factor App principles — III. Config: store config in the environment.** Source for the CLI-args-not-config pattern. Script accepts input/output as args, not via env vars or config files.
6. **Brian Kernighan & P. J. Plauger, *Software Tools* (1976).** Source for the "do one thing well + compose" pattern. The bundled script does exactly one thing (mechanical DOCX assembly); the skill body composes it with the rest of the pipeline.
7. **Doug McIlroy / Unix philosophy.** Source for the broader pattern: "Write programs that do one thing and do it well. Write programs to work together. Write programs to handle text streams, because that is a universal interface." JSON-in / DOCX-out is the modern equivalent.

View file

@ -0,0 +1,217 @@
#!/usr/bin/env python3
"""citation_tracker.py — Syllabus three-count audit + 1s sequential discipline.
Stdlib-only. Mirrors litreview's citation_tracker (research-pack convention)
adapted for syllabus's per-section search budget.
Tracked counts:
- searches_total
- searches_per_section
- papers_received
- papers_cited
Per-section detail recorded for DOCX audit log.
Enforces 1s sequential gap.
Usage:
python citation_tracker.py --action start --session syllabus-bio101-20260515 --course "Intro Biology"
python citation_tracker.py --action record_search --session ... --section "Cell Biology" --query "..."
python citation_tracker.py --action record_received --session ... --section "Cell Biology" --count 3
python citation_tracker.py --action record_cited --session ... --section "Cell Biology" --url "..."
python citation_tracker.py --action status --session ...
"""
import argparse
import json
import sys
from datetime import datetime, timezone
from pathlib import Path
from typing import Any, Dict, List, Optional
SESSIONS_DIR = Path.home() / ".syllabus_sessions"
MIN_GAP_SECONDS = 1.0
def session_path(name: str) -> Path:
return SESSIONS_DIR / f"{name}.json"
def load_session(name: str) -> Dict[str, Any]:
p = session_path(name)
if not p.exists():
raise FileNotFoundError(f"Session not found: {name}")
return json.loads(p.read_text(encoding="utf-8"))
def save_session(name: str, data: Dict[str, Any]) -> None:
SESSIONS_DIR.mkdir(parents=True, exist_ok=True)
session_path(name).write_text(json.dumps(data, indent=2), encoding="utf-8")
def now_iso() -> str:
return datetime.now(timezone.utc).isoformat()
def now_ts() -> float:
return datetime.now(timezone.utc).timestamp()
def action_start(name: str, course: Optional[str], audience: Optional[str], year_range: Optional[str]) -> Dict[str, Any]:
if session_path(name).exists():
raise FileExistsError(f"Session already exists: {name}")
data: Dict[str, Any] = {
"session": name,
"course": course or "",
"audience": audience or "",
"year_range": year_range or "",
"consensus_tier": None,
"started_at": now_iso(),
"ended_at": None,
"searches": [],
"received_log": [],
"cited": [],
"counts": {
"searches_total": 0,
"papers_received_total": 0,
"papers_cited_total": 0,
},
"by_section": {},
}
save_session(name, data)
return data
def action_record_search(name: str, section: str, query: str, tier: Optional[str]) -> Dict[str, Any]:
data = load_session(name)
if data["searches"]:
last_ts = data["searches"][-1].get("ts", 0)
gap = now_ts() - last_ts
if gap < MIN_GAP_SECONDS:
raise RuntimeError(
f"Sequential discipline violated: {gap:.2f}s gap (need >= {MIN_GAP_SECONDS}s). "
f"Wait {MIN_GAP_SECONDS - gap:.2f}s more."
)
if tier and not data["consensus_tier"]:
data["consensus_tier"] = tier
data["searches"].append({"section": section, "query": query, "tier": tier, "at": now_iso(), "ts": now_ts()})
data["counts"]["searches_total"] += 1
if section not in data["by_section"]:
data["by_section"][section] = {"searches": 0, "received": 0, "cited": 0}
data["by_section"][section]["searches"] += 1
save_session(name, data)
return data
def action_record_received(name: str, section: str, count: int) -> Dict[str, Any]:
data = load_session(name)
data["received_log"].append({"section": section, "count": count, "at": now_iso()})
data["counts"]["papers_received_total"] += count
if section not in data["by_section"]:
data["by_section"][section] = {"searches": 0, "received": 0, "cited": 0}
data["by_section"][section]["received"] += count
save_session(name, data)
return data
def action_record_cited(name: str, section: str, url: str, title: Optional[str]) -> Dict[str, Any]:
data = load_session(name)
if any(c["url"] == url for c in data["cited"]):
return data
data["cited"].append({"section": section, "url": url, "title": title, "at": now_iso()})
data["counts"]["papers_cited_total"] += 1
if section not in data["by_section"]:
data["by_section"][section] = {"searches": 0, "received": 0, "cited": 0}
data["by_section"][section]["cited"] += 1
save_session(name, data)
return data
def action_status(name: str) -> Dict[str, Any]:
return load_session(name)
def action_close(name: str) -> Dict[str, Any]:
data = load_session(name)
if data.get("ended_at") is None:
data["ended_at"] = now_iso()
save_session(name, data)
return data
def render_status_human(data: Dict[str, Any]) -> str:
out: List[str] = []
out.append(f"Session: {data['session']}")
out.append(f"Course: {data.get('course', '(unset)')}")
out.append(f"Audience: {data.get('audience', '(unset)')}")
out.append(f"Year range: {data.get('year_range', '(unset)')}")
out.append(f"Consensus tier: {data.get('consensus_tier') or '(not detected)'}")
out.append(f"Started: {data['started_at']}")
out.append(f"Ended: {data.get('ended_at') or '(active)'}")
out.append("")
c = data["counts"]
out.append(f"Total searches: {c['searches_total']}")
out.append(f"Total received: {c['papers_received_total']}")
out.append(f"Total cited: {c['papers_cited_total']}")
out.append("")
if data["by_section"]:
out.append("Per-section breakdown:")
for section, stats in data["by_section"].items():
out.append(f" {section:<40s} {stats['searches']} searches → {stats['received']} received → {stats['cited']} cited")
out.append("")
out.append("Audit block (paste in DOCX audit-log section):")
out.append(
f" Total queries: {c['searches_total']}. Papers received: {c['papers_received_total']}. "
f"Papers cited: {c['papers_cited_total']}. "
f"Plan tier: {data.get('consensus_tier') or 'undetected'}."
)
return "\n".join(out)
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--action", required=True, choices=["start", "record_search", "record_received", "record_cited", "status", "list", "close"])
parser.add_argument("--session")
parser.add_argument("--course")
parser.add_argument("--audience")
parser.add_argument("--year-range")
parser.add_argument("--section")
parser.add_argument("--query")
parser.add_argument("--tier")
parser.add_argument("--count", type=int)
parser.add_argument("--url")
parser.add_argument("--title")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
try:
if args.action == "start":
result = action_start(args.session, args.course, args.audience, args.year_range)
elif args.action == "record_search":
result = action_record_search(args.session, args.section, args.query, args.tier)
elif args.action == "record_received":
result = action_record_received(args.session, args.section, args.count)
elif args.action == "record_cited":
result = action_record_cited(args.session, args.section, args.url, args.title)
elif args.action == "status":
result = action_status(args.session)
elif args.action == "close":
result = action_close(args.session)
else:
SESSIONS_DIR.mkdir(parents=True, exist_ok=True)
result = [{"session": p.stem, "data": json.loads(p.read_text(encoding="utf-8"))} for p in sorted(SESSIONS_DIR.glob("*.json"))]
except (FileNotFoundError, FileExistsError, RuntimeError) as e:
print(f"error: {e}", file=sys.stderr); return 2
if args.output == "json":
print(json.dumps(result, indent=2, default=str))
else:
if args.action == "list":
print(json.dumps(result, indent=2, default=str))
else:
print(render_status_human(result))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))

View file

@ -0,0 +1,196 @@
#!/usr/bin/env python3
"""discussion_question_validator.py — Bloom higher-order quality check.
Stdlib-only. Validates each discussion question against Bloom's revised
taxonomy (Anderson & Krathwohl 2001). Flags:
- Recall-only questions (any audience): "what did authors find?", "summarize", etc.
- Below-audience questions (e.g., grad-doctoral course with undergrad-intro questions)
- Above-audience questions (e.g., undergrad-intro course with doctoral-level questions)
Suggests upgrades by replacing low-tier verbs with audience-appropriate Bloom verbs.
NO LLM CALLS. Pure regex + verb classification.
Usage:
python discussion_question_validator.py --questions-file /tmp/questions.json --audience grad_masters
python discussion_question_validator.py --question "What did the authors find?" --audience undergrad_intro
python discussion_question_validator.py --sample
"""
import argparse
import json
import re
import sys
from typing import Any, Dict, List, Optional
VALID_AUDIENCES = ["undergrad_intro", "undergrad_advanced", "grad_masters", "grad_doctoral", "professional", "mixed"]
# Bloom's revised taxonomy verb classification
BLOOM_VERBS = {
"remember": ["identify", "list", "recall", "name", "define", "label", "match", "recognize", "state", "what is", "what are", "what did", "describe what"],
"understand": ["explain", "summarize", "classify", "compare", "contrast", "describe how", "interpret", "paraphrase", "translate"],
"apply": ["use", "apply", "demonstrate", "implement", "execute", "carry out", "how could you use", "how would you apply", "how could this be applied", "what would happen if"],
"analyze": ["compare", "contrast", "examine", "differentiate", "organize", "what patterns", "why do", "what connections", "deconstruct"],
"evaluate": ["judge", "critique", "defend", "justify", "argue", "is this", "should we", "which is better", "do you agree", "evaluate the"],
"create": ["design", "propose", "construct", "develop", "formulate", "create a", "design a", "what would you propose", "how would you redesign"],
}
# Audience → minimum acceptable Bloom level
AUDIENCE_MIN_BLOOM = {
"undergrad_intro": 1, # Remember+ acceptable, but apply+ preferred
"undergrad_advanced": 2, # Understand+
"grad_masters": 3, # Apply+
"grad_doctoral": 4, # Analyze+
"professional": 3, # Apply+ (practice-oriented)
"mixed": 2, # Understand+ (lowest bucket present)
}
BLOOM_LEVEL_ORDER = ["remember", "understand", "apply", "analyze", "evaluate", "create"]
def classify_question(question: str) -> Dict[str, Any]:
"""Classify question by Bloom level."""
q_lower = question.lower()
detected_levels: List[str] = []
matched_phrases: Dict[str, List[str]] = {}
for level, verbs in BLOOM_VERBS.items():
for verb in verbs:
if re.search(rf"\b{re.escape(verb)}\b", q_lower):
if level not in detected_levels:
detected_levels.append(level)
matched_phrases.setdefault(level, []).append(verb)
if not detected_levels:
# Default heuristic: if starts with "what/why/how", probably understand or apply
if q_lower.strip().startswith(("what", "why", "how")):
detected_levels = ["understand"]
matched_phrases["understand"] = ["(inferred from interrogative)"]
else:
detected_levels = ["unknown"]
# Highest Bloom level detected
highest_level = "unknown"
highest_idx = -1
for level in detected_levels:
if level in BLOOM_LEVEL_ORDER:
idx = BLOOM_LEVEL_ORDER.index(level)
if idx > highest_idx:
highest_idx = idx
highest_level = level
return {
"question": question,
"detected_levels": detected_levels,
"highest_level": highest_level,
"highest_level_index": highest_idx,
"matched_phrases": matched_phrases,
}
def validate_against_audience(question: str, audience: str) -> Dict[str, Any]:
if audience not in VALID_AUDIENCES:
raise ValueError(f"Invalid audience '{audience}'. Pick from: {VALID_AUDIENCES}")
classification = classify_question(question)
min_required_idx = AUDIENCE_MIN_BLOOM[audience] - 1 # convert level to 0-indexed
detected_idx = classification["highest_level_index"]
if detected_idx == -1:
verdict = "WARN"
message = f"Could not detect Bloom level. Manual review recommended."
elif detected_idx < min_required_idx:
verdict = "FAIL"
required_level = BLOOM_LEVEL_ORDER[min_required_idx]
message = (
f"Question level '{classification['highest_level']}' is BELOW required minimum "
f"'{required_level}' for {audience}. Rework with verbs from higher Bloom levels."
)
elif detected_idx > min_required_idx + 2:
verdict = "WARN"
target_level = BLOOM_LEVEL_ORDER[min_required_idx]
message = (
f"Question level '{classification['highest_level']}' may be ABOVE typical "
f"{audience} level. Consider whether students can engage at {target_level} level."
)
else:
verdict = "PASS"
message = f"Question level '{classification['highest_level']}' appropriate for {audience}."
suggested_upgrades: List[str] = []
if verdict == "FAIL":
target_level = BLOOM_LEVEL_ORDER[min_required_idx]
suggested_upgrades = [
f"Replace verb with: {', '.join(BLOOM_VERBS[target_level][:5])}",
f"Pattern: '{BLOOM_VERBS[target_level][0]} [the {target_level} concept]...'",
]
return {
"verdict": verdict,
"audience": audience,
"min_required_level": BLOOM_LEVEL_ORDER[min_required_idx] if min_required_idx >= 0 else "unknown",
"classification": classification,
"message": message,
"suggested_upgrades": suggested_upgrades,
}
SAMPLE_QUESTIONS = [
{"question": "What did the authors find?", "audience": "undergrad_intro"},
{"question": "What did the authors find?", "audience": "grad_doctoral"},
{"question": "How could you apply this method to clinical decision support for sepsis?", "audience": "grad_masters"},
{"question": "Design a follow-up study that would test whether MUFA:SFA ratio specifically drives the lipidomic shift.", "audience": "grad_doctoral"},
{"question": "Why does the Mediterranean diet improve lipoprotein profiles?", "audience": "undergrad_intro"},
]
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--question", help="Single question to validate")
parser.add_argument("--questions-file", help="JSON file with [{question, audience}, ...] entries")
parser.add_argument("--audience", choices=VALID_AUDIENCES, help="Course audience for the question(s)")
parser.add_argument("--sample", action="store_true")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
results: List[Dict[str, Any]] = []
try:
if args.sample:
for sq in SAMPLE_QUESTIONS:
results.append(validate_against_audience(sq["question"], sq["audience"]))
elif args.question and args.audience:
results.append(validate_against_audience(args.question, args.audience))
elif args.questions_file:
from pathlib import Path
p = Path(args.questions_file)
if not p.exists():
print(f"error: {args.questions_file} not found", file=sys.stderr); return 2
data = json.loads(p.read_text(encoding="utf-8"))
for item in data:
results.append(validate_against_audience(item["question"], item["audience"]))
else:
parser.print_help(); return 0
except ValueError as e:
print(f"error: {e}", file=sys.stderr); return 2
if args.output == "json":
print(json.dumps(results, indent=2))
else:
for r in results:
marker = {"PASS": "[ok]", "WARN": "[warn]", "FAIL": "[FAIL]"}[r["verdict"]]
print(f"{marker} ({r['audience']:<20s}) {r['classification']['question'][:80]}")
print(f" Highest Bloom level: {r['classification']['highest_level']}; required: {r['min_required_level']}")
print(f"{r['message']}")
if r["suggested_upgrades"]:
print(f" Suggested upgrades:")
for s in r["suggested_upgrades"]:
print(f" - {s}")
print()
fail_count = sum(1 for r in results if r["verdict"] == "FAIL")
return 1 if fail_count > 0 else 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))

View file

@ -0,0 +1,395 @@
#!/usr/bin/env node
/**
* generate_reading_list.js Bundled DOCX generator for syllabus skill.
*
* Accepts JSON input + output path as CLI args. Produces a clean professional
* .docx reading list with title page, learning outcomes, sections of papers
* (each with hyperlinked title + audience-calibrated summary + Bloom-tied
* discussion question), and footer.
*
* Path-B build: this is the bundled mechanical layout logic. The skill
* orchestrator constructs JSON; this script assembles the DOCX. ~300 lines.
*
* Handles `docx` package require with multi-location fallback (works whether
* `docx` is installed locally, globally, or in a parent dir).
*
* JSON schema (documented in SKILL.md):
* { courseTitle, courseSubtitle, generatedDate, yearRange, introText,
* learningOutcomes: [], sections: [{ heading, papers: [...] }],
* auditLog: { totalQueriesSent, totalPapersReceived, totalPapersCited,
* toolConstraints, searchDetails: [], failures: [] } }
*
* Usage:
* node generate_reading_list.js --input data.json --output result.docx
*/
'use strict';
const fs = require('fs');
const path = require('path');
// Multi-location require for docx package
function loadDocx() {
const candidates = [
'docx', // Local node_modules
path.join(process.cwd(), 'node_modules', 'docx'), // Explicit local
'/usr/lib/node_modules/docx', // Global Linux
'/usr/local/lib/node_modules/docx', // Global macOS / brew
path.join(process.env.HOME || '', '.npm-global', 'lib', 'node_modules', 'docx'),
];
for (const candidate of candidates) {
try {
return require(candidate);
} catch (e) {
// try next
}
}
console.error('error: cannot find `docx` npm package. Install with: npm install docx');
process.exit(2);
}
const docx = loadDocx();
const {
Document, Paragraph, TextRun, Packer, AlignmentType, HeadingLevel,
ExternalHyperlink, Table, TableRow, TableCell, WidthType, ShadingType,
LevelFormat, Footer, Header, PageNumber, PageBreak, BorderStyle,
} = docx;
// ----------------------------------------------------------------------------
// CLI args
// ----------------------------------------------------------------------------
function parseArgs() {
const args = process.argv.slice(2);
const opts = {};
for (let i = 0; i < args.length; i++) {
if (args[i] === '--input') opts.input = args[++i];
else if (args[i] === '--output') opts.output = args[++i];
else if (args[i] === '--help' || args[i] === '-h') {
console.log('Usage: node generate_reading_list.js --input <data.json> --output <result.docx>');
process.exit(0);
}
}
if (!opts.input || !opts.output) {
console.error('error: both --input and --output are required');
console.error('Usage: node generate_reading_list.js --input <data.json> --output <result.docx>');
process.exit(2);
}
return opts;
}
// ----------------------------------------------------------------------------
// Input validation
// ----------------------------------------------------------------------------
function validateInput(data) {
const required = ['courseTitle', 'sections'];
for (const field of required) {
if (!data[field]) {
console.error(`error: missing required field '${field}' in input JSON`);
process.exit(2);
}
}
if (!Array.isArray(data.sections) || data.sections.length === 0) {
console.error('error: sections must be a non-empty array');
process.exit(2);
}
for (const section of data.sections) {
if (!section.heading || !Array.isArray(section.papers)) {
console.error('error: each section must have heading + papers array');
process.exit(2);
}
for (const paper of section.papers) {
if (!paper.title || !paper.url) {
console.error('error: each paper must have title + url');
process.exit(2);
}
}
}
}
// ----------------------------------------------------------------------------
// DOCX building blocks
// ----------------------------------------------------------------------------
const NAVY = '1A3A5C';
const LIGHT_BLUE = 'E8F0F8';
const ACCENT_BLUE = '2E5C8A';
const GRAY = '808080';
const DARK_GRAY = '404040';
function buildTitlePage(data) {
return [
new Paragraph({
children: [new TextRun({ text: data.courseTitle, bold: true, size: 48, color: NAVY })],
alignment: AlignmentType.CENTER,
spacing: { before: 2400, after: 200 },
}),
new Paragraph({
children: [new TextRun({ text: 'Supplementary Reading List', bold: false, size: 28, color: ACCENT_BLUE })],
alignment: AlignmentType.CENTER,
spacing: { after: 200 },
}),
data.courseSubtitle ? new Paragraph({
children: [new TextRun({ text: data.courseSubtitle, italics: true, size: 22, color: DARK_GRAY })],
alignment: AlignmentType.CENTER,
spacing: { after: 800 },
}) : null,
new Paragraph({
children: [new TextRun({ text: `Generated: ${data.generatedDate || new Date().toISOString().split('T')[0]}`, size: 18, color: GRAY })],
alignment: AlignmentType.CENTER,
spacing: { after: 100 },
}),
new Paragraph({
children: [new TextRun({ text: `Year range: ${data.yearRange || 'last 2 years'}`, size: 18, color: GRAY })],
alignment: AlignmentType.CENTER,
spacing: { after: 200 },
}),
new Paragraph({ children: [new PageBreak()] }),
].filter(Boolean);
}
function buildIntroSection(data) {
const introText = data.introText || 'This supplementary reading list collects recent peer-reviewed research relevant to each section of the course. Each entry includes a plain-language summary calibrated to the course audience and a discussion question tied to the course learning outcomes.';
return [
new Paragraph({
heading: HeadingLevel.HEADING_1,
children: [new TextRun({ text: 'Introduction', color: NAVY, bold: true, size: 32 })],
spacing: { after: 200 },
}),
new Paragraph({
children: [new TextRun({ text: introText, size: 22 })],
spacing: { after: 200 },
}),
new Paragraph({
children: [
new TextRun({ text: 'Papers sourced via ', size: 20 }),
new ExternalHyperlink({
link: 'https://consensus.app',
children: [new TextRun({ text: 'Consensus', style: 'Hyperlink', size: 20 })],
}),
new TextRun({ text: ' academic search. URLs in this document link directly to Consensus paper records.', size: 20 }),
],
spacing: { after: 400 },
}),
];
}
function buildLearningOutcomesBox(outcomes) {
if (!outcomes || outcomes.length === 0) return [];
const cells = [
new TableRow({
children: [
new TableCell({
width: { size: 9000, type: WidthType.DXA },
shading: { type: ShadingType.CLEAR, color: 'auto', fill: LIGHT_BLUE },
children: [
new Paragraph({
children: [new TextRun({ text: 'Course Learning Outcomes', bold: true, size: 24, color: NAVY })],
spacing: { after: 100 },
}),
...outcomes.map(outcome => new Paragraph({
children: [new TextRun({ text: '• ' + outcome, size: 20 })],
spacing: { after: 60 },
})),
],
}),
],
}),
];
return [
new Table({
columnWidths: [9000],
rows: cells,
}),
new Paragraph({ children: [new TextRun({ text: '', size: 4 })], spacing: { after: 400 } }),
];
}
function buildSection(section, sectionIndex) {
const elements = [
new Paragraph({
heading: HeadingLevel.HEADING_1,
children: [new TextRun({ text: `${sectionIndex}. ${section.heading}`, color: NAVY, bold: true, size: 28 })],
spacing: { before: 400, after: 200 },
}),
];
for (let i = 0; i < section.papers.length; i++) {
const paper = section.papers[i];
const paperNum = `${sectionIndex}.${i + 1}`;
// Title (hyperlinked)
elements.push(new Paragraph({
children: [
new TextRun({ text: `${paperNum}. `, bold: true, size: 22 }),
new ExternalHyperlink({
link: paper.url,
children: [new TextRun({ text: paper.title, style: 'Hyperlink', size: 22, bold: true })],
}),
],
spacing: { after: 60 },
}));
// Author / journal / year (italic gray)
const meta = `${paper.authors || ''}${paper.journal ? ' • ' + paper.journal : ''}${paper.year ? ' (' + paper.year + ')' : ''}`;
if (meta.trim()) {
elements.push(new Paragraph({
children: [new TextRun({ text: meta, italics: true, size: 18, color: GRAY })],
spacing: { after: 60 },
}));
}
// Summary
if (paper.summary) {
elements.push(new Paragraph({
children: [
new TextRun({ text: 'Summary: ', bold: true, size: 20 }),
new TextRun({ text: paper.summary, size: 20 }),
],
spacing: { after: 60 },
}));
}
// Discussion question (blue accent)
if (paper.question) {
elements.push(new Paragraph({
children: [
new TextRun({ text: 'Discussion: ', bold: true, size: 20, color: ACCENT_BLUE }),
new TextRun({ text: paper.question, size: 20 }),
],
spacing: { after: 200 },
}));
}
}
return elements;
}
function buildAuditLogSection(audit) {
if (!audit) return [];
const elements = [
new Paragraph({ children: [new PageBreak()] }),
new Paragraph({
heading: HeadingLevel.HEADING_1,
children: [new TextRun({ text: 'Audit Log', color: NAVY, bold: true, size: 28 })],
spacing: { after: 200 },
}),
new Paragraph({
children: [
new TextRun({ text: `Total queries sent: `, bold: true, size: 20 }),
new TextRun({ text: `${audit.totalQueriesSent || 0}`, size: 20 }),
],
spacing: { after: 60 },
}),
new Paragraph({
children: [
new TextRun({ text: `Total papers received: `, bold: true, size: 20 }),
new TextRun({ text: `${audit.totalPapersReceived || 0}`, size: 20 }),
],
spacing: { after: 60 },
}),
new Paragraph({
children: [
new TextRun({ text: `Total papers cited in this list: `, bold: true, size: 20 }),
new TextRun({ text: `${audit.totalPapersCited || 0}`, size: 20 }),
],
spacing: { after: 200 },
}),
];
if (audit.toolConstraints) {
elements.push(new Paragraph({
children: [
new TextRun({ text: 'Tool constraints: ', bold: true, size: 20 }),
new TextRun({ text: audit.toolConstraints, size: 20 }),
],
spacing: { after: 200 },
}));
}
if (Array.isArray(audit.searchDetails) && audit.searchDetails.length > 0) {
elements.push(new Paragraph({
children: [new TextRun({ text: 'Per-search detail:', bold: true, size: 22, color: NAVY })],
spacing: { after: 100 },
}));
for (const sd of audit.searchDetails) {
elements.push(new Paragraph({
children: [
new TextRun({ text: `${sd.section || 'Unassigned'}: `, bold: true, size: 18 }),
new TextRun({ text: `"${sd.query}" → ${sd.papersReturned || 0} returned, ${sd.papersSelected || 0} selected (${sd.status || 'OK'})`, size: 18 }),
],
spacing: { after: 40 },
}));
}
}
if (Array.isArray(audit.failures) && audit.failures.length > 0) {
elements.push(new Paragraph({
children: [new TextRun({ text: 'Failures:', bold: true, size: 22, color: 'AA0000' })],
spacing: { before: 200, after: 100 },
}));
for (const f of audit.failures) {
elements.push(new Paragraph({
children: [new TextRun({ text: `${f}`, size: 18, color: '880000' })],
spacing: { after: 40 },
}));
}
}
return elements;
}
function buildFooter(data) {
return new Footer({
children: [
new Paragraph({
children: [new TextRun({ text: `${data.courseTitle} — Supplementary Reading List`, size: 16, color: GRAY })],
alignment: AlignmentType.CENTER,
}),
],
});
}
// ----------------------------------------------------------------------------
// Main
// ----------------------------------------------------------------------------
function main() {
const opts = parseArgs();
let data;
try {
data = JSON.parse(fs.readFileSync(opts.input, 'utf-8'));
} catch (e) {
console.error(`error: cannot read input JSON ${opts.input}: ${e.message}`);
process.exit(2);
}
validateInput(data);
const sections = data.sections.map((s, i) => buildSection(s, i + 1)).flat();
const doc = new Document({
creator: 'syllabus skill',
title: `${data.courseTitle} — Supplementary Reading List`,
description: 'Generated by syllabus skill via bundled generate_reading_list.js',
sections: [
{
properties: {
page: {
margin: { top: 1440, right: 1440, bottom: 1440, left: 1440 }, // 1 inch
size: { width: 12240, height: 15840 }, // US Letter
},
},
footers: { default: buildFooter(data) },
children: [
...buildTitlePage(data),
...buildIntroSection(data),
...buildLearningOutcomesBox(data.learningOutcomes),
...sections,
...buildAuditLogSection(data.auditLog),
],
},
],
});
Packer.toBuffer(doc).then(buffer => {
fs.writeFileSync(opts.output, buffer);
console.log(`Generated: ${opts.output} (${buffer.length} bytes, ${data.sections.length} sections, ${data.sections.reduce((sum, s) => sum + s.papers.length, 0)} papers)`);
}).catch(e => {
console.error(`error: DOCX packing failed: ${e.message}`);
process.exit(2);
});
}
main();

View file

@ -0,0 +1,198 @@
#!/usr/bin/env python3
"""topic_grouper.py — Heuristic 6-12 section grouping from extracted syllabus topics.
Stdlib-only. Given a list of extracted course topics, produce a proposed
grouping into 6-12 sections by detecting shared keywords.
The output feeds the Phase 2 group-and-confirm checkpoint where the user
can override (proceed / merge / split / add / remove).
Algorithm:
1. Tokenize each topic into significant words (stop-words removed)
2. Build word topics inverted index
3. Greedy clustering: topics sharing 2+ significant words same section
4. Cap at 12 sections (over-cap merge smallest); ensure minimum 6 (under split largest)
5. Each section gets a heading derived from its dominant shared keyword
NO LLM CALLS. Pure tokenization + clustering.
Usage:
python topic_grouper.py --topics "Cell biology, DNA replication, Protein synthesis, ..."
python topic_grouper.py --topics-file /tmp/topics.json
python topic_grouper.py --sample
"""
import argparse
import json
import re
import sys
from collections import Counter, defaultdict
from typing import Any, Dict, List, Set
STOP_WORDS = {
"the", "a", "an", "and", "or", "but", "if", "of", "in", "on", "at", "to",
"for", "with", "by", "from", "is", "are", "was", "were", "be", "been",
"this", "that", "these", "those", "introduction", "overview", "basics",
"fundamentals", "principles", "concepts", "topics", "review", "advanced",
"intermediate", "i", "ii", "iii", "iv", "v", "1", "2", "3", "4", "5",
"6", "7", "8", "9", "10", "11", "12", "week", "lecture", "chapter", "unit",
"module", "lesson", "section",
}
MIN_SECTIONS = 6
MAX_SECTIONS = 12
SHARED_WORD_THRESHOLD = 2
def tokenize(topic: str) -> Set[str]:
"""Extract significant words from a topic string."""
words = re.findall(r"\b[a-z]{3,}\b", topic.lower())
return {w for w in words if w not in STOP_WORDS}
def cluster_topics(topics: List[str]) -> List[Dict[str, Any]]:
"""Cluster topics by shared significant words."""
topic_tokens = [(i, t, tokenize(t)) for i, t in enumerate(topics)]
clusters: List[List[int]] = [] # list of topic-index lists
assigned: Set[int] = set()
for i, _, tokens_i in topic_tokens:
if i in assigned:
continue
# Start a new cluster with topic i
cluster = [i]
assigned.add(i)
# Try to add other topics that share >= SHARED_WORD_THRESHOLD tokens
for j, _, tokens_j in topic_tokens:
if j in assigned or j == i:
continue
shared = tokens_i & tokens_j
if len(shared) >= SHARED_WORD_THRESHOLD:
cluster.append(j)
assigned.add(j)
clusters.append(cluster)
return _normalize_to_size(clusters, topic_tokens)
def _normalize_to_size(clusters: List[List[int]], topic_tokens: List[tuple]) -> List[Dict[str, Any]]:
"""Ensure 6-12 sections by merging smallest or splitting largest."""
# Merge smallest if over MAX_SECTIONS
while len(clusters) > MAX_SECTIONS:
clusters.sort(key=len)
smallest = clusters.pop(0)
# Merge into next-smallest
if clusters:
clusters[0].extend(smallest)
else:
clusters.append(smallest)
# Split largest if under MIN_SECTIONS (and largest has >= 4 items)
while len(clusters) < MIN_SECTIONS and clusters:
clusters.sort(key=len, reverse=True)
largest = clusters.pop(0)
if len(largest) >= 4:
mid = len(largest) // 2
clusters.extend([largest[:mid], largest[mid:]])
else:
clusters.insert(0, largest)
break # Can't split further
# Generate section heading per cluster (most-common shared word)
sections: List[Dict[str, Any]] = []
topic_lookup = {i: (t, tokens) for i, t, tokens in topic_tokens}
for cluster_indices in clusters:
all_tokens: Counter = Counter()
cluster_topics: List[str] = []
for idx in cluster_indices:
topic, tokens = topic_lookup[idx]
cluster_topics.append(topic)
all_tokens.update(tokens)
# Heading = top 1-3 most common tokens, capitalized
top_words = [w for w, _ in all_tokens.most_common(2)]
heading = " + ".join(w.capitalize() for w in top_words) if top_words else f"Section {len(sections) + 1}"
sections.append({
"heading": heading,
"topic_count": len(cluster_indices),
"topics": cluster_topics,
})
return sections
SAMPLE_TOPICS = [
"Cell Biology Fundamentals",
"DNA Replication",
"Protein Synthesis",
"Cell Division and Mitosis",
"Mendelian Genetics",
"Population Genetics",
"Evolution and Natural Selection",
"Speciation",
"Ecology Basics",
"Ecosystem Dynamics",
"Energy Flow in Ecosystems",
"Conservation Biology",
"Plant Anatomy",
"Plant Physiology",
"Animal Anatomy Overview",
"Animal Behavior",
"Microbiology Introduction",
"Bacterial Genetics",
"Viruses and Pathogens",
]
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--topics", help="Comma-separated topic list")
parser.add_argument("--topics-file", help="Path to JSON file with topics array")
parser.add_argument("--sample", action="store_true")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
if args.sample:
topics = SAMPLE_TOPICS
elif args.topics:
topics = [t.strip() for t in args.topics.split(",") if t.strip()]
elif args.topics_file:
from pathlib import Path
p = Path(args.topics_file)
if not p.exists():
print(f"error: {args.topics_file} not found", file=sys.stderr); return 2
topics = json.loads(p.read_text(encoding="utf-8"))
else:
parser.print_help(); return 0
sections = cluster_topics(topics)
result = {
"input_topic_count": len(topics),
"section_count": len(sections),
"sections": sections,
}
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(f"Input topics: {len(topics)}")
print(f"Output sections: {len(sections)} (target: {MIN_SECTIONS}-{MAX_SECTIONS})")
print()
print("Proposed sections (present this at Phase 2 checkpoint):")
for i, s in enumerate(sections, 1):
print(f"")
print(f" Section {i}: {s['heading']} ({s['topic_count']} topics)")
for t in s["topics"]:
print(f" - {t}")
print()
print("Group-and-confirm checkpoint forcing options:")
print(" 1. Looks good — proceed with these sections")
print(" 2. Merge sections [X] and [Y]")
print(" 3. Split section [X] into two")
print(" 4. Add a section for [topic]")
print(" 5. Remove section [X]")
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))