mirror of
https://github.com/alirezarezvani/claude-skills.git
synced 2026-10-06 02:50:08 +00:00
- Resolve conflicts: keep redesigned skills index, take dev's cs-aeo link fix, union of DOMAIN_SEO_CONTEXT entries in generate-docs.py - Regenerate catalog on the merged tree (dev's agent/command description updates, removed ai-seo/release-manager/command-guide, restructured universal-scraping-architect) - Update counters to post-merge truth from scripts/derive_counters.py: 345 skills, 78 plugins (14 bundles + 64 standalone), 570+ Python tools - Add redirects for upstream-removed pages (ai-seo -> aeo, release-manager -> changelog-generator, command-guide -> engineering index) - Add compliance-os bundle to bundle tables; rebuild 78-plugin table from marketplace.json - Teach the generator to rewrite repo-root-relative source links to GitHub URLs — mkdocs build --strict now passes with zero warnings https://claude.ai/code/session_015bYZ97nV4oRb3LbxCRFVcP
41 lines
2.5 KiB
Markdown
41 lines
2.5 KiB
Markdown
---
|
|
title: "/cs-scrape — Slash Command for AI Coding Agents"
|
|
description: "Route, extract, and validate a scraping job (URL or local file) via the universal-scraping-architect skill — refuses to deliver unvalidated data.. Slash command for Claude Code, Codex CLI, Gemini CLI."
|
|
---
|
|
|
|
# /cs-scrape
|
|
|
|
<div class="page-meta" markdown>
|
|
<span class="meta-badge">:material-console: Slash Command</span>
|
|
<span class="meta-badge">:material-github: <a href="https://github.com/alirezarezvani/2-claude-skills/tree/main/engineering/universal-scraping-architect/commands/cs-scrape.md">Source</a></span>
|
|
</div>
|
|
|
|
|
|
Run a gated extraction pipeline for `$ARGUMENTS` using `skills/universal-scraping-architect/SKILL.md`.
|
|
|
|
## Pre-flight gates (stop if any fails)
|
|
|
|
1. **Target stated?** If `$ARGUMENTS` is empty, ask for the URL or file path plus the desired output format — do not guess.
|
|
2. **Live-site etiquette:** for URLs, check `robots.txt` and plan rate limits; refuse disallowed targets.
|
|
3. **Privacy:** if the target is a local/sensitive file, do not send it to an external API — force Mode 2 (local Python).
|
|
4. **Secrets:** Firecrawl key only via `os.getenv('FIRECRAWL_API_KEY')`; if a key appears inline anywhere, fix that first.
|
|
|
|
## Workflow
|
|
|
|
1. **Route** — state the mode and why (per the skill's routing rules):
|
|
Mode 1 Firecrawl (public/JS-heavy URL, bulk crawl) · Mode 2 local Python (local files, private data, simple static HTML) · Mode 3 hybrid (Firecrawl extract + pandas clean).
|
|
2. **Budget** — estimate API quota / token limits before multi-page jobs; add checkpointing + pagination.
|
|
3. **Extract** — start from the matching runner template (run from the plugin root; `--sample` previews the summary shape offline):
|
|
```bash
|
|
python3 skills/universal-scraping-architect/scripts/firecrawl_example.py --sample
|
|
python3 skills/universal-scraping-architect/scripts/local_bs4_example.py --sample
|
|
```
|
|
4. **Validate (mandatory, exit-code gated):**
|
|
```bash
|
|
python3 skills/universal-scraping-architect/scripts/validate_extraction.py extracted_output.json --json
|
|
```
|
|
- exit 0 (`status: ok`) → continue
|
|
- exit 1 (`warning` = empty output, `error` = malformed JSON) → fix and re-extract; **never deliver unvalidated data**
|
|
|
|
Then check required fields and duplicates against the job spec.
|
|
5. **Deliver** — CSV (tabular) / JSON (nested) / Markdown (docs, chunked), per the user's requested format, with a summary of mode chosen, row counts, empty values, and the validation verdict.
|