Connects each research-ops sub-skill to engineering/autoresearch-agent and adds
per-skill onboarding with a customization approach that the tools actually consume.
Each sub-skill (clinical-research, research-finance, market-research,
product-research) now ships three integration scripts:
- onboard.py — its own onboarding questionnaire; interactive or
--defaults / --set key=value / --reset / --scope; writes
~/.config/research-ops/<skill>.json (global) or
./.research-ops/<skill>.json (project).
- config_loader.py — loads that config (project > global > defaults;
RESEARCH_OPS_NO_CONFIG=1 bypass). Every scoring tool now
reads it so saved answers change behavior: default profile,
thresholds (alpha/power/dropout, F&A rate, runway, confidence,
MoE, insight source-threshold), and named owners printed on
clinical/finance outputs. CLI flags always override.
- ar_evaluator.py — an isolated, OPT-IN ground-truth evaluator bridging to
autoresearch. The loop edits the skill's input file and the
evaluator (never edited) scores it: clinical
feasibility_composite (higher), finance runway_months (higher),
market tam_divergence (lower), product validated_insights
(higher). No cross-skill coupling; invoked only on explicit
user request.
SKILL.md (each), the orchestrator, the agent, per-skill commands, and the domain
CLAUDE.md document the onboarding + opt-in autoresearch handoff. plugin.json /
marketplace describe 24 tools (12 analysis + 12 integration).
Builds the cross-platform plugins: codex symlinks + gemini SKILL.md mirrors +
vibe symlinks materialized for all 5 research-ops skills; indices regenerated.
Root CLAUDE.md scope/highlights updated.
https://claude.ai/code/session_01PUNmQVE4WYvcrzpq2anC3D
Builds on @mitnick2012's universal+per-language restructure (PR #742).
- Add Java as a first-class deterministic language in code_quality_checker.py
(LANGUAGE_EXTENSIONS + function/class/method patterns + check_java_specific_smells),
so the documented `--language java` command works instead of erroring on an
invalid choice. Add Java debug + @SuppressWarnings signals to pr_analyzer.py.
- Add Java regression fixtures (sample_java_smells/clean.java) with committed
expected_outputs JSON, mirroring the existing C# fixtures.
- Delete references/{code_review_checklist,coding_standards,common_antipatterns}.md,
now duplicated by rules/universal.md + languages/*.md; repoint README and the
C# clean fixture header at the new structure.
- Document the optional analyzer-wiring + fixture steps in the "Adding a New
Language" guide and restore a Regression Fixtures section in SKILL.md.
https://claude.ai/code/session_01DjuELpoFdFbFscr3kAatni
Both ./engineering/handoff and ./productivity/handoff were registered as
"handoff" in marketplace.json. Duplicate names pass the local CLI but fail
Claude.ai marketplace sync (used by Cowork), causing "Failed to add
marketplace". Renamed to handoff-engineering and handoff-productivity.
https://claude.ai/code/session_01DjuELpoFdFbFscr3kAatni
- Extract commun languages rules in a separate rules/universal.md containing all cross-language rules in one place
- Move language-specific rules inline into each languages/*.md file,
organised into consistent sections: Security / Async / Resource
Management / Exception Handling / Performance / Idioms
- Add Java support: languages/java.md with full section coverage
- Every review now requires exactly 2 file reads: universal.md +
one language file
- Add "Adding a new language" guide to SKILL.md: one file to create,
nothing else changes
The previous vibe sync was run with --domain productivity, which rewrote
.vibe/skills/claude-skills/skills-index.json to contain only the 6 productivity
skills, dropping the other ~317. The symlink tree was unaffected (all 14 domain
dirs intact) — only the index JSON was clobbered. Re-running the full sync
restores total_skills 6 -> 323 across all 14 domains. andreessen present.
https://claude.ai/code/session_01SF6MzfjHurZMt5JUFET9h3
Genuine, repo-consistent additions that lift the real quality gaps (not doc
padding):
- assets/forcing_question_worksheet.md — fillable 6-question interrogation
- assets/blank_3x5_card.md — blank daily card template
- assets/example_market_verdict.md — full worked market-first verdict
- assets/example_pmf_check.md — worked before/after PMF check
Both worked examples' tool invocations are verified against the actual scripts
(MARKET-FIRST-DERISK at composite 6.36; BEFORE-PMF at composite 4.35). SKILL.md
Assets section updated to reference all five.
Quality scorer: 56.2 -> 65.7 (clears the 60 gate). examples 60->100%,
assets 12->60%, practical_examples 40->80%. Audit now a clean PASS on quality
alongside structure 91.3/EXCELLENT, scripts 3/3, security 0/0.
https://claude.ai/code/session_01SF6MzfjHurZMt5JUFET9h3
Post-merge audit fixes for the andreessen productivity skill:
- Add skills/andreessen/README.md (inner skill README). Lifts structure
91.3/EXCELLENT and quality 56.2; the README was the one genuine doc gap
vs sibling skills.
- Run the gemini cross-platform sync (codex ran at merge time; gemini was
missed). Adds the andreessen symlink + index entry and reconciles a
pre-existing stale claude-coach entry the generator surfaced.
Audit verdict: PASS WITH WARNINGS. Structure 91.3 EXCELLENT, scripts 3/3
PASS, security PASS (0 critical/0 high), marketplace + ecosystem clean. The
single warning is the quality scorer's title-case section/frontmatter schema
that no Path-B productivity skill uses (andreessen 56.2 vs merged siblings
reflect 44.6 / capture 46.4).
https://claude.ai/code/session_01SF6MzfjHurZMt5JUFET9h3
New productivity/andreessen plugin — a Marc Andreessen-mode operator that
pressure-tests ventures/ideas/features/bets through his documented frameworks
(market > team > product; product/market fit is the only milestone; bias to
build) and runs his 3x5-card + Anti-Todo daily routine. Built as the
Andreessen-lens counterpart to a founder-operating-system plugin.
Runs on the user-supplied anti-sycophancy operating prompt, preserved verbatim
in references/operating_prompt.md (counterargument first, no premise validation,
no disclaimers, explicit confidence levels, no capitulation without new
evidence). The second emphasis block is operationalized as a posture-mapping
table so each instruction changes behavior rather than sitting as decoration.
Ships 3 stdlib-only deterministic tools (market_first_evaluator with a hard
sub-4 market kill gate, pmf_signal_scorer with the Sean Ellis 40% gate,
anti_todo_card enforcing the 3-5 cap), 4 references each citing 5-7 sources with
explicit confidence levels on every Andreessen attribution, cs-andreessen agent,
/cs:andreessen + /cs:pmf-check commands, and a worked 3x5-card asset.
Registered in marketplace.json + .codex skills index (productivity 5 -> 6).
.codex review/run/status symlinks reflect the sync generator's standard
collision resolution on the current tree.
https://claude.ai/code/session_01SF6MzfjHurZMt5JUFET9h3
Addresses the Phase 3 quality_scorer roadmap items from the plugin audit:
adds the bundled fixtures, sample outputs, and quick-reference README that
the scorer expects, without diverging from the project's minimal-frontmatter
SKILL.md convention.
assets/:
- sample_csharp_smells.cs: a C# fixture with every pattern the skill
detects (async void, blocking on Task, swallowed Exception, undisposed
IDisposable, new HttpClient(), missing await, null-forgiving, hardcoded
connection string, unsafe, dynamic, #pragma warning disable,
[SuppressMessage], SQL concatenation), each smell labelled inline
- sample_csharp_clean.cs: the same code refactored per the standards in
references/coding_standards.md — verifies the analyzer produces 0 HIGH
smells on idiomatic code
expected_outputs/:
- sample_csharp_smells_quality.json: committed analyzer output for the
smells fixture (F/45, 3 HIGH smells)
- sample_csharp_clean_quality.json: committed analyzer output for the
clean fixture (A/98, 0 HIGH smells)
These act as a regression harness: diff the live output against the
committed JSON to detect any behaviour change in the analyzer.
scripts/code_quality_checker.py:
- Add _strip_csharp_comments() that removes // line and /* */ block
comments before running C#-specific regex detectors. Fixes false
positives where comment prose ("// FIX: await instead of .Result")
matched a detection pattern.
SKILL.md:
- New ## Examples section pointing at the fixtures + showing how to
reproduce the expected output with diff
- TOC updated to list "C# / .NET Review Notes" and "Examples"
README.md (new):
- Quick-reference card with how-to, 3 worked examples (one per script),
pointer to fixtures, pointer to references
Phase re-scores after this change:
- Structure: 86.4/GOOD → 91.3/EXCELLENT (+4.9)
- Quality: 54.8/D → 72.5/B- (+17.7)
- Scripts: 3/3 PASS (unchanged)
- Security: 0/0 (unchanged)
Phase 7 of the plugin audit on engineering-team/skills/code-reviewer
found stale entries in the cross-platform indices — they still showed
the pre-#723 description without C# / .NET. Re-running sync also
refreshed two unrelated drifted entries (review/run/status codex
symlinks, claude-coach gemini symlink).
v2.8.1 was already taken by the engineering role-skill upgrade
(senior-fullstack / senior-frontend / senior-backend with karpathy-coder
+ Matt Pocock decision engines), released 2026-05-20 — before the
handoff PRs even merged. The auto-release workflow created the v2.8.1
tag from that work via CHANGELOG.md parsing.
The productivity/handoff skill is the next minor on top of v2.8.1:
v2.8.2.
Changes:
- CHANGELOG.md: prepend a new [2.8.2] entry documenting the handoff
skill (PRs #724, #728, #729). The auto-release workflow
(.github/workflows/release.yml) will pick up this entry and create
the v2.8.2 git tag + GitHub Release on the next push to main.
- productivity/handoff/.claude-plugin/plugin.json: 2.8.1 -> 2.8.2
- .claude-plugin/marketplace.json (handoff entry): 2.8.1 -> 2.8.2
- CLAUDE.md: 4 spots bumped to v2.8.2; v2.8.1 references kept where
they correctly point to the engineering role-skill release
- README.md: Productivity table row ✨v2.8.1 -> ✨v2.8.2
- docs/index.md: description, hero subtitle, "329 Skills" card text
- docs/getting-started.md: description meta + FAQ count text
- mkdocs.yml: site_description
The narrative across all top-level docs now reads correctly:
v2.8.0 (bizops + commercial) -> v2.8.1 (engineering role-skills) ->
v2.8.2 (productivity/handoff).
Verified:
- 0 v2.7.5 references remain (earlier typo)
- All v2.8.1 references that remain point to engineering role-skills
- CHANGELOG topmost entry: [2.8.2] - 2026-05-23
- plugin.json + marketplace.json both at 2.8.2
- mkdocs build clean (will re-verify in CI)
https://claude.ai/code/session_01KLhHBAfEDXdQMeRe6G8sRa
Ships the three improvements judged most impactful in v1.1 design review:
1. SessionEnd hook (hooks/session_end.py)
Pairs with SessionStart. When a session ends with no handoff in the
last 30 minutes, prints a one-line reminder. Cannot prompt
interactively or block session end — surfaces text via stdout.
Disable per-session with HANDOFF_SESSIONEND=0. hooks.json updated to
wire both SessionStart and SessionEnd.
2. handoff_self_check.py — fidelity script (~300 LOC, stdlib-only)
Operationalizes handoff_prompt.md. Six checks:
- All 5 sections present
- Goal is non-empty and non-placeholder
- State-of-play bullets reference at least one artifact (commit hash,
PR/issue number, file path, URL)
- Open decisions are present (or explicit "- None.") when git is dirty
or has recent commits
- Skills to use: 3-5 entries, hard cap enforced
- Artifacts contain paths/URLs only, no inline content
Severity: high/medium/low. Strict mode exits 1 only on HIGH findings.
--sample fixture has 3 planted issues (2 high + 1 medium) and exits 1.
Canonical example_handoff.md passes clean (exit 0).
/cs:handoff command updated to run self-check between scaffold-fill
and redaction linter.
3. --refresh flag on handoff_template_generator.py
Reuses the most recent handoff in the configured save location
instead of creating a new file. Falls through to create-if-missing
when no existing handoff is found. Keeps the save location
uncluttered when work continues past the original handoff time;
ensures the SessionStart hook always loads the up-to-date version.
Version bump: 2.7.4 -> 2.7.5. Marketplace description and keywords
updated. README v1.1 section added. SKILL.md gains "Refreshing an
Existing Handoff" and "SessionEnd Reminder" subsections.
Verified:
- All 9 Python files compile clean
- self-check --sample correctly fails (3 findings, exit 1)
- self-check passes clean against assets/example_handoff.md (exit 0)
- --refresh finds the latest /tmp/handoff-*.md and prints its path
- SessionEnd hook prints the reminder when no recent handoff exists
- check_plugin_json.py + marketplace.json + hooks.json all parse
- Plugin audit re-run: structure 84.2 -> 86.0, quality 62.2 -> 63.0,
security PASS (0 critical, 0 high)
- Codex + Gemini sync re-ran clean
https://claude.ai/code/session_01KLhHBAfEDXdQMeRe6G8sRa