From 41b116c8c085d4848b1cbf5745bc311dd8bd39cf Mon Sep 17 00:00:00 2001 From: mehanshbarthwal-lab Date: Wed, 20 May 2026 11:17:02 +0530 Subject: [PATCH 1/3] feat(engineering): add universal-scraping-architect skill --- .../universal-scraping-architect/LICENSE | 21 ++ .../universal-scraping-architect/README.md | 212 ++++++++++++++++ .../universal-scraping-architect/SKILL.md | 16 ++ .../scripts/firecrawl_example.py | 157 ++++++++++++ .../scripts/local_bs4_example.py | 226 ++++++++++++++++++ 5 files changed, 632 insertions(+) create mode 100644 skills/engineering/universal-scraping-architect/LICENSE create mode 100644 skills/engineering/universal-scraping-architect/README.md create mode 100644 skills/engineering/universal-scraping-architect/SKILL.md create mode 100644 skills/engineering/universal-scraping-architect/scripts/firecrawl_example.py create mode 100644 skills/engineering/universal-scraping-architect/scripts/local_bs4_example.py diff --git a/skills/engineering/universal-scraping-architect/LICENSE b/skills/engineering/universal-scraping-architect/LICENSE new file mode 100644 index 00000000..52def8d8 --- /dev/null +++ b/skills/engineering/universal-scraping-architect/LICENSE @@ -0,0 +1,21 @@ +MIT License + +Copyright (c) 2026 Mehansh Barthwal + +Permission is hereby granted, free of charge, to any person obtaining a copy +of this software and associated documentation files (the "Software"), to deal +in the Software without restriction, including without limitation the rights +to use, copy, modify, merge, publish, distribute, sublicense, and/or sell +copies of the Software, and to permit persons to whom the Software is +furnished to do so, subject to the following conditions: + +The above copyright notice and this permission notice shall be included in all +copies or substantial portions of the Software. + +THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR +IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, +FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE +AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER +LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, +OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE +SOFTWARE. diff --git a/skills/engineering/universal-scraping-architect/README.md b/skills/engineering/universal-scraping-architect/README.md new file mode 100644 index 00000000..53aeef04 --- /dev/null +++ b/skills/engineering/universal-scraping-architect/README.md @@ -0,0 +1,212 @@ +# Universal Scraping Architect + +> A robust, general-purpose scraping and data extraction framework designed as a reusable **Skill** for AI Agents and LLMs. + +[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT) +[![Python 3.8+](https://img.shields.io/badge/python-3.8+-blue.svg)](https://www.python.org/downloads/) +[![Firecrawl Ready](https://img.shields.io/badge/Firecrawl-Ready-orange.svg)](https://firecrawl.dev) + +--- + +## What This Is + +Most scraping scripts are brittle one-offs that break the moment a page layout changes, a column gets renamed, or a file format shifts slightly. This framework is built differently — it treats every extraction task as a **complete data pipeline**, with intelligent routing, validation, token tracking, and clean outputs baked in from the start. + +It was built specifically to be dropped into an AI agent's skill library (like Claude's skill system), so that any scraping or data extraction task — whether it's a live public website, a local PDF, an Excel file, or a nested JSON — gets handled with the same consistent, robust approach every time. + +--- + +## Key Features + +**Intelligent Approach Routing** — Before writing a single line of code, the framework decides whether the task calls for Firecrawl (dynamic web, search-first, or bulk crawl), traditional local scraping (local files, static HTML, private data), or a hybrid of both. It states the choice and the reason. + +**Firecrawl Integration (5 Paths)** — Full support for Firecrawl's CLI, REST API, and SDK across five clearly defined paths: live web data, app/product integration, finished deliverables, auth-only setup, and direct REST without installation. + +**Token Budget Tracking** — Every extraction that feeds an LLM estimates token volume against a configurable context limit before processing, warning early if the output is over budget. + +**Firecrawl Quota Safety** — Single-key rule enforced by default. Estimates page/credit usage before large crawl jobs and only prompts for a new key when the current one is actually exhausted or a job would genuinely exceed ~1000 requests. + +**Validation at Every Step** — Required fields are checked, row counts are logged, empty outputs are caught, and duplicate rows are flagged before anything gets saved. + +**Checkpointing for Large Jobs** — Progress is saved during multi-page or multi-file extractions so a failure mid-job doesn't mean starting from scratch. + +**Security and Ethical Scraping** — API keys are loaded from environment variables only, never hardcoded or printed. robots.txt, rate limits, and terms of service are respected. Private files are never sent to external APIs without explicit approval. + +--- + +## Supported Sources + +| Category | Examples | +|---|---| +| **Live Web** | Public URLs, dynamic pages, SPAs, paginated sites | +| **Search-First** | Topic queries, keyword discovery, entity research | +| **Bulk Crawl** | Full site sections, documentation sites, large domains | +| **Local Files** | PDF, DOCX, Excel (.xlsx/.xls), CSV, JSON, XML, ZIP | +| **Scanned Docs** | OCR pipelines for image-based PDFs | +| **APIs** | REST APIs, paginated API responses, JSON data feeds | +| **Databases** | CSV exports, SQLite, structured data dumps | + +--- + +## Output Formats + +The framework can save clean outputs as CSV, Excel, JSON, Markdown, TXT, SQLite, Parquet, or HTML depending on what the task needs. If not specified, it defaults sensibly — CSV for structured tabular data, JSON for nested structures, Markdown for clean web page text. + +--- + +## Project Structure + +``` +universal-scraping-architect/ +├── SKILL.md # Core LLM skill definition (the agent's system prompt) +├── README.md # This file +├── LICENSE # MIT License +├── .gitignore # Standard Python/project ignores +├── requirements.txt # Python dependencies +└── examples/ + ├── firecrawl_example.py # Path C workflow: Firecrawl → clean markdown output + └── local_bs4_example.py # Traditional scraping: static HTML → validated CSV +``` + +--- + +## How To Use This As An AI Skill + +Drop `SKILL.md` into your agent's system prompt or tool context. The agent will then adopt the Universal Scraping Architect approach for any extraction task — routing correctly, validating outputs, tracking tokens, and producing copy-paste-ready pipeline code rather than fragile one-offs. + +This is how it appears in Claude's skill system: + +```yaml +name: universal-scraping-architect +description: Use this skill for any scraping, crawling, extraction, parsing, web + research, document processing, dataset preparation, Firecrawl workflow, + local file extraction, API extraction, PDF/Excel/CSV/JSON/XML parsing, + validation-heavy data pipeline, or repeatable clean-output scraping task. +``` + +--- + +## Quick Start (For Developers Running the Examples) + +**Install dependencies:** + +```bash +pip install -r requirements.txt +``` + +**For Firecrawl workflows, set your API key:** + +```bash +# Linux / macOS +export FIRECRAWL_API_KEY="fc-YOUR_API_KEY_HERE" + +# Windows PowerShell +$env:FIRECRAWL_API_KEY = "fc-YOUR_API_KEY_HERE" + +# Windows Command Prompt +set FIRECRAWL_API_KEY=fc-YOUR_API_KEY_HERE +``` + +Or add it to a `.env` file in your project root (never commit this file): + +```dotenv +FIRECRAWL_API_KEY=fc-YOUR_API_KEY_HERE +``` + +**Run the Firecrawl example:** + +```bash +python examples/firecrawl_example.py +``` + +**Run the local scraping example:** + +```bash +python examples/local_bs4_example.py +``` + +--- + +## The 15-Step Pipeline + +Every task the framework handles follows this sequence: + +1. Understand the source +2. Choose the most appropriate extraction approach +3. Configure task-specific settings +4. Extract safely +5. Handle pagination, layout changes, dynamic content, or file variations +6. Clean the extracted data +7. Normalize structure and field names +8. Validate the result +9. Track token/data volume if LLM processing is involved +10. Estimate Firecrawl usage/quota before large Firecrawl jobs +11. Handle errors clearly +12. Save clean outputs +13. Save logs and checkpoints when useful +14. Print a final summary +15. Explain what was done and what can be customized + +--- + +## Firecrawl Path Reference + +| Path | When To Use | +|---|---| +| **Path A** | Need live web data right now during the current session | +| **Path B** | Building an app or product that calls Firecrawl from code | +| **Path C** | Need a finished deliverable — research brief, SEO audit, lead list, etc. | +| **Path D** | Need to set up an account or API key first | +| **Path E** | Don't want to install anything — use the REST API directly | + +Install command (covers all paths): + +```bash +npx -y firecrawl-cli@latest init --all --browser +``` + +--- + +## When To Use Firecrawl vs. Local Scraping + +**Use Firecrawl when:** +- The source is a public URL and you want clean, reliable extraction +- The page is dynamic (JavaScript-rendered, SPA, requires interaction) +- You need search-first discovery before you know the URLs +- You're crawling many pages across a domain +- You want to wire Firecrawl into an app or agentic workflow + +**Use local/traditional scraping when:** +- The source is a local file (PDF, Excel, CSV, JSON, XML) +- The data is private or sensitive and shouldn't leave your machine +- You're using official downloads/APIs where scraping isn't needed +- Simple static HTML where Firecrawl would be overkill +- Custom parsing logic with pandas, pdfplumber, openpyxl, etc. is the right fit + +**Use a hybrid when:** +- Firecrawl handles web extraction, then Python cleans and structures the output +- Firecrawl discovers URLs, then local code processes and saves the dataset +- Web content gets merged with local files + +--- + +## Security Notes + +- API keys are always loaded from environment variables — never hardcoded +- `.env` files are in `.gitignore` and should never be committed +- Real keys are never printed in logs or included in output files +- Private or sensitive files are never sent to Firecrawl without explicit user approval +- One active Firecrawl API key is used at a time (no rotation by default) +- robots.txt and rate limits are respected + +--- + +## Author + +**Mehansh Barthwal** + +--- + +## License + +This project is licensed under the MIT License — see [LICENSE](LICENSE) for details. diff --git a/skills/engineering/universal-scraping-architect/SKILL.md b/skills/engineering/universal-scraping-architect/SKILL.md new file mode 100644 index 00000000..288a9126 --- /dev/null +++ b/skills/engineering/universal-scraping-architect/SKILL.md @@ -0,0 +1,16 @@ +--- +name: universal-scraping-architect +description: Use this skill for any scraping, crawling, extraction, parsing, web research, document processing, dataset preparation, Firecrawl workflow, local file extraction, API extraction, PDF/Excel/CSV/JSON/XML parsing, validation-heavy data pipeline, or repeatable clean-output scraping task. +--- +# Universal Scraping Architect +## Purpose +This skill is a reusable, general-purpose scraping and data extraction framework. +Use it whenever the user asks for scraping, crawling, extraction, parsing, or data collection. + +## Mandatory Approach Selection Rule +1. **Firecrawl** — public web data, dynamic content, search-first discovery +2. **Traditional/local scraping** — local files, private data, static HTML +3. **Hybrid approach** — Firecrawl for web extraction, local tools for processing + +## Master Rule +The scraping framework must always remain customizable, approach-aware, token-aware, Firecrawl-quota-aware, validated, and error-safe. diff --git a/skills/engineering/universal-scraping-architect/scripts/firecrawl_example.py b/skills/engineering/universal-scraping-architect/scripts/firecrawl_example.py new file mode 100644 index 00000000..f7b0a9c8 --- /dev/null +++ b/skills/engineering/universal-scraping-architect/scripts/firecrawl_example.py @@ -0,0 +1,157 @@ +""" +Universal Scraping Architect: Firecrawl Example (Path C — Repeatable Deliverable) +=================================================================================== +Extracts clean markdown content from a target URL using the Firecrawl SDK. + +Demonstrates: + - Safe API key loading from environment + - Token budget tracking before LLM processing + - Firecrawl quota awareness logging + - Structured error handling + - Clean output saving with validation + +Usage: + export FIRECRAWL_API_KEY="fc-YOUR_KEY_HERE" # Linux/macOS + $env:FIRECRAWL_API_KEY = "fc-YOUR_KEY_HERE" # Windows PowerShell + python examples/firecrawl_example.py +""" + +import os +import sys +from datetime import datetime +from firecrawl import FirecrawlApp + +# ============================================================================= +# CONFIG — Edit these for your task +# ============================================================================= +TARGET_URL = "https://example.com/research-data" +OUTPUT_FILE = "clean_extraction.md" +LOG_FILE = "firecrawl_run.log" + +# Firecrawl / Token settings +FIRECRAWL_API_KEY_ENV = "FIRECRAWL_API_KEY" +TOKEN_CONTEXT_LIMIT = 100_000 # Adjust to your model's context window +RESERVED_OUTPUT_TOKENS = 4_000 # Tokens held back for the model's response + + +# ============================================================================= +# HELPERS +# ============================================================================= + +def log(message: str, level: str = "INFO"): + """Prints a timestamped log line and appends it to the log file.""" + timestamp = datetime.now().strftime("%Y-%m-%d %H:%M:%S") + line = f"[{timestamp}] [{level}] {message}" + print(line) + with open(LOG_FILE, "a", encoding="utf-8") as f: + f.write(line + "\n") + + +def check_environment() -> str: + """Loads the Firecrawl API key from the environment. Fails fast if missing.""" + api_key = os.getenv(FIRECRAWL_API_KEY_ENV) + if not api_key: + log( + f"Missing environment variable: {FIRECRAWL_API_KEY_ENV}. " + "Set it before running this script.", + level="ERROR" + ) + sys.exit(1) + log("API key loaded successfully from environment.") + return api_key + + +def estimate_tokens(text: str) -> int: + """Rough token estimate based on character count (characters / 4).""" + return len(text) // 4 + + +def check_token_budget(text: str) -> bool: + """ + Estimates token usage and warns if over budget. + Returns True if within budget, False if over. + """ + estimated = estimate_tokens(text) + available = TOKEN_CONTEXT_LIMIT - RESERVED_OUTPUT_TOKENS + log(f"Token estimate: {estimated:,} / {available:,} available tokens.") + if estimated > available: + log( + f"OVER TOKEN BUDGET by {estimated - available:,} tokens. " + "Consider chunking the output before passing to an LLM.", + level="WARNING" + ) + return False + log("Within token budget.") + return True + + +def validate_content(content: str) -> bool: + """Basic validation — checks the response is non-empty.""" + if not content or not content.strip(): + log("Validation failed: empty content received from Firecrawl.", level="ERROR") + return False + log(f"Validation passed. Content length: {len(content):,} characters.") + return True + + +def save_output(content: str, path: str): + """Saves validated content to the output path.""" + with open(path, "w", encoding="utf-8") as f: + f.write(content) + log(f"Output saved to: {path}") + + +# ============================================================================= +# MAIN +# ============================================================================= + +def main(): + log("=" * 60) + log("Universal Scraping Architect — Firecrawl Path C") + log("=" * 60) + + # Step 1: Environment check + api_key = check_environment() + + # Step 2: Initialise Firecrawl client + app = FirecrawlApp(api_key=api_key) + log(f"Firecrawl client initialised. Target: {TARGET_URL}") + + # Step 3: Scrape + try: + log("Starting scrape...") + result = app.scrape_url(TARGET_URL, params={"formats": ["markdown"]}) + markdown_content = result.get("markdown", "") + except Exception as e: + log(f"Firecrawl scrape failed: {e}", level="ERROR") + log( + "Tip: if this is a quota or auth error, check your FIRECRAWL_API_KEY " + "and run `firecrawl --status` to inspect account state.", + level="WARNING" + ) + sys.exit(1) + + # Step 4: Validate + if not validate_content(markdown_content): + sys.exit(1) + + # Step 5: Token budget check + check_token_budget(markdown_content) + + # Step 6: Save clean output + save_output(markdown_content, OUTPUT_FILE) + + # Step 7: Final summary + log("=" * 60) + log("EXTRACTION COMPLETE") + log(f" Source URL : {TARGET_URL}") + log(f" Output file : {OUTPUT_FILE}") + log(f" Characters : {len(markdown_content):,}") + log(f" Est. tokens : {estimate_tokens(markdown_content):,}") + log(f" Log file : {LOG_FILE}") + log("=" * 60) + log("Customisable: TARGET_URL, OUTPUT_FILE, TOKEN_CONTEXT_LIMIT, formats.") + + +if __name__ == "__main__": + main() diff --git a/skills/engineering/universal-scraping-architect/scripts/local_bs4_example.py b/skills/engineering/universal-scraping-architect/scripts/local_bs4_example.py new file mode 100644 index 00000000..0d70c4ce --- /dev/null +++ b/skills/engineering/universal-scraping-architect/scripts/local_bs4_example.py @@ -0,0 +1,226 @@ +""" +Universal Scraping Architect: Traditional / Local Scraping Example +================================================================== +Extracts a data table from a static HTML page using Requests and BeautifulSoup, +then validates, cleans, and saves to CSV. + +Demonstrates: + - Safe HTTP fetching with headers, timeouts, and retry logic + - HTML table parsing with pandas + - Column name normalisation to snake_case + - Required field validation before saving + - Structured error handling at every stage + - Clean, logged output + +Usage: + pip install -r requirements.txt + python examples/local_bs4_example.py +""" + +import sys +import time +from datetime import datetime + +import pandas as pd +import requests +from bs4 import BeautifulSoup + +# ============================================================================= +# CONFIG — Edit these for your task +# ============================================================================= +TARGET_URL = "https://example.com/macroeconomic-indicators/energy-prices" +OUTPUT_FILE = "energy_price_data.csv" +LOG_FILE = "local_scrape_run.log" + +# HTTP settings +USER_AGENT = "UniversalScrapingArchitect/1.0 (contact: your@email.com)" +TIMEOUT_SECONDS = 15 +MAX_RETRIES = 3 +RETRY_DELAY_SECONDS = 2 + +# Validation +REQUIRED_COLUMNS = ["date", "price_index", "yoy_change"] +TABLE_SELECTOR = {"id": "indicator-data"} # Edit to match your target table's HTML attributes + + +# ============================================================================= +# HELPERS +# ============================================================================= + +def log(message: str, level: str = "INFO"): + """Prints a timestamped log line and appends it to the log file.""" + timestamp = datetime.now().strftime("%Y-%m-%d %H:%M:%S") + line = f"[{timestamp}] [{level}] {message}" + print(line) + with open(LOG_FILE, "a", encoding="utf-8") as f: + f.write(line + "\n") + + +def safe_get(url: str) -> str: + """ + Fetches HTML from a URL with polite headers, a timeout, and retry logic. + Raises on non-2xx status. + """ + headers = {"User-Agent": USER_AGENT} + + for attempt in range(1, MAX_RETRIES + 1): + try: + log(f"Attempt {attempt}/{MAX_RETRIES}: GET {url}") + response = requests.get(url, headers=headers, timeout=TIMEOUT_SECONDS) + response.raise_for_status() + log(f"HTTP {response.status_code} — content length: {len(response.text):,} chars.") + return response.text + except requests.exceptions.HTTPError as e: + log(f"HTTP error on attempt {attempt}: {e}", level="ERROR") + except requests.exceptions.ConnectionError as e: + log(f"Connection error on attempt {attempt}: {e}", level="ERROR") + except requests.exceptions.Timeout: + log(f"Timeout on attempt {attempt} after {TIMEOUT_SECONDS}s.", level="WARNING") + except requests.exceptions.RequestException as e: + log(f"Request failed on attempt {attempt}: {e}", level="ERROR") + + if attempt < MAX_RETRIES: + log(f"Waiting {RETRY_DELAY_SECONDS}s before retry...") + time.sleep(RETRY_DELAY_SECONDS) + + raise RuntimeError(f"All {MAX_RETRIES} attempts failed for URL: {url}") + + +def find_table(html: str) -> BeautifulSoup: + """ + Parses the HTML and returns the target table element. + Adjust TABLE_SELECTOR to match the table you need. + """ + soup = BeautifulSoup(html, "html.parser") + table = soup.find("table", TABLE_SELECTOR) + if not table: + raise ValueError( + f"Target table not found. Selector used: {TABLE_SELECTOR}. " + "Inspect the page source and update TABLE_SELECTOR in CONFIG." + ) + log("Target table found in HTML.") + return table + + +def parse_table(table) -> pd.DataFrame: + """Converts a BeautifulSoup table element into a pandas DataFrame.""" + df = pd.read_html(str(table))[0] + log(f"Parsed table: {len(df)} rows, {len(df.columns)} columns.") + return df + + +def clean_column_names(df: pd.DataFrame) -> pd.DataFrame: + """Normalises all column names to snake_case.""" + df.columns = ( + df.columns + .str.strip() + .str.lower() + .str.replace(r"\s+", "_", regex=True) + .str.replace(r"[^\w]", "", regex=True) + ) + log(f"Columns after normalisation: {list(df.columns)}") + return df + + +def clean_data(df: pd.DataFrame) -> pd.DataFrame: + """ + Applies general cleaning rules. + Extend this function for your task-specific cleaning needs. + """ + # Strip whitespace from string columns + for col in df.select_dtypes(include="object").columns: + df[col] = df[col].str.strip() + + # Drop fully empty rows + before = len(df) + df = df.dropna(how="all") + dropped = before - len(df) + if dropped: + log(f"Dropped {dropped} fully empty rows.", level="WARNING") + + return df + + +def validate(df: pd.DataFrame) -> bool: + """ + Checks that the DataFrame is non-empty and contains all required columns. + Returns True if valid, False otherwise. + """ + if df.empty: + log("Validation failed: DataFrame is empty.", level="ERROR") + return False + + missing = [col for col in REQUIRED_COLUMNS if col not in df.columns] + if missing: + log( + f"Validation failed: missing required columns: {missing}. " + f"Available columns: {list(df.columns)}", + level="ERROR" + ) + return False + + log(f"Validation passed. {len(df)} rows, {len(df.columns)} columns.") + return True + + +def save_output(df: pd.DataFrame, path: str): + """Saves the validated DataFrame to CSV.""" + df.to_csv(path, index=False, encoding="utf-8") + log(f"Output saved to: {path}") + + +# ============================================================================= +# MAIN +# ============================================================================= + +def main(): + log("=" * 60) + log("Universal Scraping Architect — Traditional / Local Scraping") + log("=" * 60) + + # Step 1: Fetch + try: + html = safe_get(TARGET_URL) + except RuntimeError as e: + log(str(e), level="ERROR") + sys.exit(1) + + # Step 2: Find table + try: + table = find_table(html) + except ValueError as e: + log(str(e), level="ERROR") + sys.exit(1) + + # Step 3: Parse + try: + df = parse_table(table) + except Exception as e: + log(f"Failed to parse table into DataFrame: {e}", level="ERROR") + sys.exit(1) + + # Step 4: Clean + df = clean_column_names(df) + df = clean_data(df) + + # Step 5: Validate + if not validate(df): + sys.exit(1) + + # Step 6: Save + save_output(df, OUTPUT_FILE) + + # Step 7: Final summary + log("=" * 60) + log("EXTRACTION COMPLETE") + log(f" Source URL : {TARGET_URL}") + log(f" Output file : {OUTPUT_FILE}") + log(f" Rows saved : {len(df)}") + log(f" Columns : {list(df.columns)}") + log(f" Log file : {LOG_FILE}") + log("=" * 60) + log("Customisable: TARGET_URL, OUTPUT_FILE, TABLE_SELECTOR, REQUIRED_COLUMNS.") + + +if __name__ == "__main__": + main() From ef69f595417f2cfa43ad3c60921a9bada1204d12 Mon Sep 17 00:00:00 2001 From: mehanshbarthwal-lab Date: Wed, 20 May 2026 13:45:32 +0530 Subject: [PATCH 2/3] improve(engineering): align universal-scraping-architect with Path-B contract --- .../.claude-plugin/plugin.json | 16 +++++ .../.idea/.gitignore | 10 ++++ .../inspectionProfiles/profiles_settings.xml | 6 ++ .../.idea/misc.xml | 7 +++ .../.idea/modules.xml | 8 +++ .../.idea/universal-scraping-architect.iml | 14 +++++ .../.idea/vcs.xml | 6 ++ .../universal-scraping-architect/LICENSE | 0 .../universal-scraping-architect/README.md | 0 .../universal-scraping-architect/SKILL.md | 58 +++++++++++++++++++ .../agents/agents/cs-scraping-architect.md | 6 ++ .../commands/commands/cs-scrape.md | 6 ++ .../references/firecrawl-technical-guide.md | 10 ++++ .../references/parsing-and-data-extraction.md | 10 ++++ .../references/scraping-ethics-security.md | 10 ++++ .../scripts/firecrawl_example.py | 0 .../scripts/local_bs4_example.py | 0 .../scripts/scripts/validate_extraction.py | 44 ++++++++++++++ .../universal-scraping-architect/SKILL.md | 16 ----- 19 files changed, 211 insertions(+), 16 deletions(-) create mode 100644 engineering/universal-scraping-architect/.claude-plugin/plugin.json create mode 100644 engineering/universal-scraping-architect/.idea/.gitignore create mode 100644 engineering/universal-scraping-architect/.idea/inspectionProfiles/profiles_settings.xml create mode 100644 engineering/universal-scraping-architect/.idea/misc.xml create mode 100644 engineering/universal-scraping-architect/.idea/modules.xml create mode 100644 engineering/universal-scraping-architect/.idea/universal-scraping-architect.iml create mode 100644 engineering/universal-scraping-architect/.idea/vcs.xml rename {skills/engineering => engineering}/universal-scraping-architect/LICENSE (100%) rename {skills/engineering => engineering}/universal-scraping-architect/README.md (100%) create mode 100644 engineering/universal-scraping-architect/SKILL.md create mode 100644 engineering/universal-scraping-architect/agents/agents/cs-scraping-architect.md create mode 100644 engineering/universal-scraping-architect/commands/commands/cs-scrape.md create mode 100644 engineering/universal-scraping-architect/references/firecrawl-technical-guide.md create mode 100644 engineering/universal-scraping-architect/references/references/parsing-and-data-extraction.md create mode 100644 engineering/universal-scraping-architect/references/references/references/scraping-ethics-security.md rename {skills/engineering => engineering}/universal-scraping-architect/scripts/firecrawl_example.py (100%) rename {skills/engineering => engineering}/universal-scraping-architect/scripts/local_bs4_example.py (100%) create mode 100644 engineering/universal-scraping-architect/scripts/scripts/validate_extraction.py delete mode 100644 skills/engineering/universal-scraping-architect/SKILL.md diff --git a/engineering/universal-scraping-architect/.claude-plugin/plugin.json b/engineering/universal-scraping-architect/.claude-plugin/plugin.json new file mode 100644 index 00000000..70e889ea --- /dev/null +++ b/engineering/universal-scraping-architect/.claude-plugin/plugin.json @@ -0,0 +1,16 @@ +{ + "name": "universal-scraping-architect", + "description": "A universal scraping skill with intelligent routing, token budget tracking, and quota awareness.", + "version": "2.1.2", + "author": { + "name": "Mehansh Barthwal", + "url": "https://github.com/mehanshbarthwal-lab" + }, + "homepage": "https://github.com/alirezarezvani/claude-skills/tree/main/engineering/universal-scraping-architect", + "repository": "https://github.com/alirezarezvani/claude-skills", + "license": "MIT", + "skills": "./", + "dependencies": { + "python": ["firecrawl", "pandas", "requests", "beautifulsoup4"] + } +} \ No newline at end of file diff --git a/engineering/universal-scraping-architect/.idea/.gitignore b/engineering/universal-scraping-architect/.idea/.gitignore new file mode 100644 index 00000000..30cf57ed --- /dev/null +++ b/engineering/universal-scraping-architect/.idea/.gitignore @@ -0,0 +1,10 @@ +# Default ignored files +/shelf/ +/workspace.xml +# Editor-based HTTP Client requests +/httpRequests/ +# Ignored default folder with query files +/queries/ +# Datasource local storage ignored files +/dataSources/ +/dataSources.local.xml diff --git a/engineering/universal-scraping-architect/.idea/inspectionProfiles/profiles_settings.xml b/engineering/universal-scraping-architect/.idea/inspectionProfiles/profiles_settings.xml new file mode 100644 index 00000000..105ce2da --- /dev/null +++ b/engineering/universal-scraping-architect/.idea/inspectionProfiles/profiles_settings.xml @@ -0,0 +1,6 @@ + + + + \ No newline at end of file diff --git a/engineering/universal-scraping-architect/.idea/misc.xml b/engineering/universal-scraping-architect/.idea/misc.xml new file mode 100644 index 00000000..187eec5d --- /dev/null +++ b/engineering/universal-scraping-architect/.idea/misc.xml @@ -0,0 +1,7 @@ + + + + + + \ No newline at end of file diff --git a/engineering/universal-scraping-architect/.idea/modules.xml b/engineering/universal-scraping-architect/.idea/modules.xml new file mode 100644 index 00000000..940fef20 --- /dev/null +++ b/engineering/universal-scraping-architect/.idea/modules.xml @@ -0,0 +1,8 @@ + + + + + + + + \ No newline at end of file diff --git a/engineering/universal-scraping-architect/.idea/universal-scraping-architect.iml b/engineering/universal-scraping-architect/.idea/universal-scraping-architect.iml new file mode 100644 index 00000000..3ab2d77f --- /dev/null +++ b/engineering/universal-scraping-architect/.idea/universal-scraping-architect.iml @@ -0,0 +1,14 @@ + + + + + + + + + + + + \ No newline at end of file diff --git a/engineering/universal-scraping-architect/.idea/vcs.xml b/engineering/universal-scraping-architect/.idea/vcs.xml new file mode 100644 index 00000000..b2bdec2d --- /dev/null +++ b/engineering/universal-scraping-architect/.idea/vcs.xml @@ -0,0 +1,6 @@ + + + + + + \ No newline at end of file diff --git a/skills/engineering/universal-scraping-architect/LICENSE b/engineering/universal-scraping-architect/LICENSE similarity index 100% rename from skills/engineering/universal-scraping-architect/LICENSE rename to engineering/universal-scraping-architect/LICENSE diff --git a/skills/engineering/universal-scraping-architect/README.md b/engineering/universal-scraping-architect/README.md similarity index 100% rename from skills/engineering/universal-scraping-architect/README.md rename to engineering/universal-scraping-architect/README.md diff --git a/engineering/universal-scraping-architect/SKILL.md b/engineering/universal-scraping-architect/SKILL.md new file mode 100644 index 00000000..32028e98 --- /dev/null +++ b/engineering/universal-scraping-architect/SKILL.md @@ -0,0 +1,58 @@ +--- +name: "universal-scraping-architect" +description: "Use for web scraping, crawling, document extraction, API parsing, or building validation-heavy data pipelines using Firecrawl or local Python scripts." +--- + +# Universal Scraping Architect + +You are an expert web scraping and data extraction engineer. Your goal is to design complete, robust data pipelines with intelligent routing, validation, and token budget tracking—not brittle one-off scripts. + +**Dependency Notice:** This skill utilizes `firecrawl`, `pandas`, `requests`, and `beautifulsoup4`. It uses a BYOK (Bring Your Own Key) pattern for Firecrawl. API keys must only be loaded via environment variables. + +## Before Starting +**Check for context first:** +If `project-context.md` exists, read it before asking questions. Determine the target data format, scale of extraction, and deployment environment before writing any code. + +## How This Skill Works + +This skill supports 3 extraction modes based on intelligent routing: + +### Mode 1: API-Driven (Firecrawl) +Use when the source is a public URL, heavily dynamic (JS/SPA), requires search-first discovery, or involves bulk crawling across a domain. +### Mode 2: Local Python (Traditional) +Use when extracting from local files (PDF, Excel, CSV), the data is private/sensitive, or the target is a simple static HTML page where Firecrawl is overkill. +### Mode 3: Hybrid Pipeline +Use when Firecrawl handles URL discovery/web extraction, but local Python (Pandas) is required to clean, normalize, and structure the output before saving. + +## The Extraction Pipeline + +When executing a scraping task, always follow this sequence: +1. **Route the Approach:** Explicitly state whether Firecrawl or Local Python is being used and why. +2. **Track Budgets:** Estimate Firecrawl API quotas or LLM token context limits before executing large jobs. +3. **Extract Safely:** Implement checkpointing for multi-page jobs. Handle pagination and dynamic layouts gracefully. +4. **Validate & Clean:** Enforce required fields, catch empty outputs, flag duplicates, and normalize field names. +5. **Format:** Default to CSV for tabular data, JSON for nested structures, and Markdown for clean text. + +## Proactive Triggers + +Surface these issues WITHOUT being asked when you notice them in context: +- **Hardcoded API Keys** → Flag immediately and rewrite to use `os.getenv('FIRECRAWL_API_KEY')`. +- **Private Data Leakage** → If the user asks to send local, sensitive files to an external API, flag the privacy risk and suggest Mode 2 (Local Python). +- **Missing Pagination** → If the target implies hundreds of records but no pagination logic is requested, flag it and add checkpointing. + +## Output Artifacts + +| When you ask for... | You get... | +|---------------------|------------| +| "Scrape this site" | A fully validated Python extraction script with routing logic and error handling. | +| "Get data from this table" | A clean CSV/JSON dataset with a summary log of row counts and empty values. | +| "Crawl these docs" | A Markdown deliverable chunked for LLM token limits. | + +## Anti-Patterns +- **Brittle Selectors:** Never use highly nested CSS selectors (e.g., `div > span > ul > li:nth-child(3)`). Use data attributes or robust structural anchors. +- **Ignoring Etiquette:** Never scrape without checking `robots.txt` or implementing sensible rate limits. +- **No Validation:** Never blindly write scraped data to a file without checking if the array is empty or missing critical keys. + +## Related Skills +- **data-cleaning**: Use when the scraped data requires complex statistical normalization or deduplication. +- **browser-automation**: Use for highly interactive scraping requiring user emulation (clicks, logins) where Firecrawl is insufficient. \ No newline at end of file diff --git a/engineering/universal-scraping-architect/agents/agents/cs-scraping-architect.md b/engineering/universal-scraping-architect/agents/agents/cs-scraping-architect.md new file mode 100644 index 00000000..15bf6938 --- /dev/null +++ b/engineering/universal-scraping-architect/agents/agents/cs-scraping-architect.md @@ -0,0 +1,6 @@ +--- +name: cs-scraping-architect +description: Expert persona for web scraping and data pipeline design. +--- +# cs-scraping-architect +Use this agent when you need to design a complex extraction strategy or debug scraping scripts. \ No newline at end of file diff --git a/engineering/universal-scraping-architect/commands/commands/cs-scrape.md b/engineering/universal-scraping-architect/commands/commands/cs-scrape.md new file mode 100644 index 00000000..a939de83 --- /dev/null +++ b/engineering/universal-scraping-architect/commands/commands/cs-scrape.md @@ -0,0 +1,6 @@ +--- +name: cs-scrape +description: Execute a scraping task for a specific URL. +--- +# /cs-scrape [url] +Triggers the scraping architect to analyze and extract data from the target URL. \ No newline at end of file diff --git a/engineering/universal-scraping-architect/references/firecrawl-technical-guide.md b/engineering/universal-scraping-architect/references/firecrawl-technical-guide.md new file mode 100644 index 00000000..de4d2bf7 --- /dev/null +++ b/engineering/universal-scraping-architect/references/firecrawl-technical-guide.md @@ -0,0 +1,10 @@ +# Firecrawl Technical Guide + +This document covers the technical integration patterns for the Firecrawl API within the Universal Scraping Architect. + +### Authoritative Sources +1. [Firecrawl API Documentation](https://docs.firecrawl.dev/api-reference/introduction) +2. [Firecrawl SDK for Python](https://github.com/mendableai/firecrawl-py) +3. [REST API Design Best Practices (Microsoft)](https://learn.microsoft.com/en-us/azure/architecture/best-practices/api-design) +4. [Handling API Rate Limits (Cloudflare)](https://developers.cloudflare.com/fundamentals/api/reference/rate-limits/) +5. [JSON Schema Standard](https://json-schema.org/specification.html) \ No newline at end of file diff --git a/engineering/universal-scraping-architect/references/references/parsing-and-data-extraction.md b/engineering/universal-scraping-architect/references/references/parsing-and-data-extraction.md new file mode 100644 index 00000000..221765d2 --- /dev/null +++ b/engineering/universal-scraping-architect/references/references/parsing-and-data-extraction.md @@ -0,0 +1,10 @@ +# Parsing and Data Extraction Standards + +Detailed strategies for BeautifulSoup4 and Pandas integration for cleaning scraped data. + +### Authoritative Sources +1. [BeautifulSoup4 Documentation](https://www.crummy.com/software/BeautifulSoup/bs4/doc/) +2. [Pandas Data Cleaning Guide](https://pandas.pydata.org/docs/user_guide/10min.html) +3. [W3C HTML Living Standard](https://html.spec.whatwg.org/multipage/) +4. [CSS Selectors Level 4 (W3C)](https://www.w3.org/TR/selectors-4/) +5. [Mozilla MDN - HTTP Headers (User-Agent)](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/User-Agent) \ No newline at end of file diff --git a/engineering/universal-scraping-architect/references/references/references/scraping-ethics-security.md b/engineering/universal-scraping-architect/references/references/references/scraping-ethics-security.md new file mode 100644 index 00000000..bc09c71e --- /dev/null +++ b/engineering/universal-scraping-architect/references/references/references/scraping-ethics-security.md @@ -0,0 +1,10 @@ +# Scraping Ethics, Robots.txt, and Security + +Guidelines for ethical data collection and securing scraping pipelines. + +### Authoritative Sources +1. [The Robots Exclusion Protocol (RFC 9309)](https://datatracker.ietf.org/doc/rfc9309/) +2. [OWASP Automated Threats to Web Applications](https://owasp.org/www-project-automated-threats-to-web-applications/) +3. [Scraping Ethics Best Practices (Ethical Web Scraping)](https://www.ethicalwebscraping.org/) +4. [Legal Aspects of Web Scraping (Lexology)](https://www.lexology.com/library/detail.aspx?g=e6e0287a-62ad-4d6d-852a-9e535e69e6b4) +5. [Cloudflare Bot Management Overview](https://www.cloudflare.com/en-gb/pg-lp/bot-management-for-everyone/) \ No newline at end of file diff --git a/skills/engineering/universal-scraping-architect/scripts/firecrawl_example.py b/engineering/universal-scraping-architect/scripts/firecrawl_example.py similarity index 100% rename from skills/engineering/universal-scraping-architect/scripts/firecrawl_example.py rename to engineering/universal-scraping-architect/scripts/firecrawl_example.py diff --git a/skills/engineering/universal-scraping-architect/scripts/local_bs4_example.py b/engineering/universal-scraping-architect/scripts/local_bs4_example.py similarity index 100% rename from skills/engineering/universal-scraping-architect/scripts/local_bs4_example.py rename to engineering/universal-scraping-architect/scripts/local_bs4_example.py diff --git a/engineering/universal-scraping-architect/scripts/scripts/validate_extraction.py b/engineering/universal-scraping-architect/scripts/scripts/validate_extraction.py new file mode 100644 index 00000000..97dd4e5a --- /dev/null +++ b/engineering/universal-scraping-architect/scripts/scripts/validate_extraction.py @@ -0,0 +1,44 @@ +#!/usr/bin/env python3 +""" +validate_extraction.py +Intelligent stdlib-only validation for JSON structures. +""" +import argparse +import json +import sys +import os + + +def validate_json(file_path): + if not os.path.exists(file_path): + return {"status": "error", "message": f"File {file_path} not found."} + + try: + with open(file_path, 'r', encoding='utf-8') as f: + data = json.load(f) + # Basic validation: ensure it's not empty and is a list/dict + if not data: + return {"status": "warning", "message": "JSON file is empty."} + return {"status": "ok", "message": "JSON is valid and well-formed."} + except json.JSONDecodeError as e: + return {"status": "error", "message": f"Invalid JSON: {str(e)}"} + + +def main(): + parser = argparse.ArgumentParser(description="Standard Library JSON Validator") + parser.add_argument("file", help="Path to JSON file to validate") + parser.add_argument("--json", action="store_true", help="Output results in JSON format") + + args = parser.parse_args() + result = validate_json(args.file) + + if args.json: + print(json.dumps(result, indent=2)) + else: + print(f"[{result['status'].upper()}] {result['message']}") + + sys.exit(0 if result['status'] == "ok" else 1) + + +if __name__ == "__main__": + main() \ No newline at end of file diff --git a/skills/engineering/universal-scraping-architect/SKILL.md b/skills/engineering/universal-scraping-architect/SKILL.md deleted file mode 100644 index 288a9126..00000000 --- a/skills/engineering/universal-scraping-architect/SKILL.md +++ /dev/null @@ -1,16 +0,0 @@ ---- -name: universal-scraping-architect -description: Use this skill for any scraping, crawling, extraction, parsing, web research, document processing, dataset preparation, Firecrawl workflow, local file extraction, API extraction, PDF/Excel/CSV/JSON/XML parsing, validation-heavy data pipeline, or repeatable clean-output scraping task. ---- -# Universal Scraping Architect -## Purpose -This skill is a reusable, general-purpose scraping and data extraction framework. -Use it whenever the user asks for scraping, crawling, extraction, parsing, or data collection. - -## Mandatory Approach Selection Rule -1. **Firecrawl** — public web data, dynamic content, search-first discovery -2. **Traditional/local scraping** — local files, private data, static HTML -3. **Hybrid approach** — Firecrawl for web extraction, local tools for processing - -## Master Rule -The scraping framework must always remain customizable, approach-aware, token-aware, Firecrawl-quota-aware, validated, and error-safe. From 867db4150f0f419e286c85fea533fde421636f04 Mon Sep 17 00:00:00 2001 From: mehanshbarthwal-lab Date: Wed, 20 May 2026 14:01:18 +0530 Subject: [PATCH 3/3] chore: register skill in marketplace.json --- .claude-plugin/marketplace.json | 60 +++-- .idea/.gitignore | 10 + .idea/claude-skills.iml | 24 ++ .../inspectionProfiles/profiles_settings.xml | 6 + .idea/modules.xml | 8 + .idea/vcs.xml | 6 + .../universal-scraping-architect/LICENSE | 21 -- .../universal-scraping-architect/README.md | 212 ------------------ .../{agents => }/cs-scraping-architect.md | 0 .../commands/{commands => }/cs-scrape.md | 0 .../references/parsing-and-data-extraction.md | 10 - .../scraping-ethics-security.md | 0 12 files changed, 93 insertions(+), 264 deletions(-) create mode 100644 .idea/.gitignore create mode 100644 .idea/claude-skills.iml create mode 100644 .idea/inspectionProfiles/profiles_settings.xml create mode 100644 .idea/modules.xml create mode 100644 .idea/vcs.xml delete mode 100644 engineering/universal-scraping-architect/LICENSE delete mode 100644 engineering/universal-scraping-architect/README.md rename engineering/universal-scraping-architect/agents/{agents => }/cs-scraping-architect.md (100%) rename engineering/universal-scraping-architect/commands/{commands => }/cs-scrape.md (100%) delete mode 100644 engineering/universal-scraping-architect/references/references/parsing-and-data-extraction.md rename engineering/universal-scraping-architect/references/{references/references => }/scraping-ethics-security.md (100%) diff --git a/.claude-plugin/marketplace.json b/.claude-plugin/marketplace.json index 711a2239..54d751d5 100644 --- a/.claude-plugin/marketplace.json +++ b/.claude-plugin/marketplace.json @@ -4,7 +4,7 @@ "name": "Alireza Rezvani", "url": "https://alirezarezvani.com" }, - "description": "328 production-ready skill packages for Claude AI across 14 domains: engineering advanced (76 \u2014 incl. 4 Matt Pocock-derived productivity skills + v2.7.3 security-guidance PreToolUse hook), engineering core (51), marketing (47 \u2014 incl. v2.7.3 AEO/Answer Engine Optimization), c-level advisory (66), product (17), regulatory/QMS (18), project management (9), business growth (5), finance (4), productivity (4, v2.7.0), marketing top-level (2, v2.7.0), research (8, v2.7.0), business-operations (7, v2.8.0), and commercial (8, v2.8.0). Includes ~441 Python tools, ~594 reference documents, 48+ agents, 77+ slash commands.", + "description": "328 production-ready skill packages for Claude AI across 14 domains: engineering advanced (76 — incl. 4 Matt Pocock-derived productivity skills + v2.7.3 security-guidance PreToolUse hook), engineering core (51), marketing (47 — incl. v2.7.3 AEO/Answer Engine Optimization), c-level advisory (66), product (17), regulatory/QMS (18), project management (9), business growth (5), finance (4), productivity (4, v2.7.0), marketing top-level (2, v2.7.0), research (8, v2.7.0), business-operations (7, v2.8.0), and commercial (8, v2.8.0). Includes ~441 Python tools, ~594 reference documents, 48+ agents, 77+ slash commands.", "homepage": "https://github.com/alirezarezvani/claude-skills", "repository": "https://github.com/alirezarezvani/claude-skills", "metadata": { @@ -61,7 +61,7 @@ { "name": "c-level-agents", "source": "./c-level-advisor/c-level-agents", - "description": "Founder-mode executive team plugin: 13 cs-* C-suite agents (CFO, CMO, CRO, CPO, COO, CHRO, CISO, Chief of Staff, General Counsel, Chief Data Officer, Chief AI Officer, Chief Customer Officer, VP of Engineering) with distinct cognitive voices, plus 21 /cs:* slash commands \u2014 forcing-question office hours (CFO/CMO/CPO/CRO/CTO/CISO/GC/CDO/CAIO/CCO/VPE reviews), strategic sprint pipeline (brief \u2192 boardroom \u2192 decide \u2192 execute \u2192 post-mortem), and meta routing (/cs:founder-mode auto-router, /cs:onboard, /cs:cross-eval multi-model consensus, /cs:freeze cooldown lock). Wraps the 33 c-level skills with cognitive gearing, persona voice, and artifact-driven handoffs. The business-domain answer to YC Garry Tan's gstack.", + "description": "Founder-mode executive team plugin: 13 cs-* C-suite agents (CFO, CMO, CRO, CPO, COO, CHRO, CISO, Chief of Staff, General Counsel, Chief Data Officer, Chief AI Officer, Chief Customer Officer, VP of Engineering) with distinct cognitive voices, plus 21 /cs:* slash commands — forcing-question office hours (CFO/CMO/CPO/CRO/CTO/CISO/GC/CDO/CAIO/CCO/VPE reviews), strategic sprint pipeline (brief → boardroom → decide → execute → post-mortem), and meta routing (/cs:founder-mode auto-router, /cs:onboard, /cs:cross-eval multi-model consensus, /cs:freeze cooldown lock). Wraps the 33 c-level skills with cognitive gearing, persona voice, and artifact-driven handoffs. The business-domain answer to YC Garry Tan's gstack.", "version": "1.5.0", "author": { "name": "Alireza Rezvani" @@ -111,7 +111,7 @@ { "name": "general-counsel-advisor", "source": "./c-level-advisor/general-counsel-advisor", - "description": "General Counsel advisory for startups: contract risk scanner (12 founder-killer patterns: auto-renew traps, uncapped indemnity, vague IP, MFN pricing, missing DPA, one-sided venue, broad non-solicit, perpetual license-back, etc.) and term sheet analyzer (0-100 founder-friendliness across 12 dimensions). 3 in-depth references: contracts playbook (7 startup contract types), IP + regulatory landscape mapping (HIPAA, GDPR, FDA, fintech, EU AI Act, SOC 2 \u2192 ISO sequencing), term sheet decoder (full glossary + founder-friendly defaults). Standalone-installable; also bundled in c-level-skills. Stdlib-only. NOT a substitute for licensed counsel.", + "description": "General Counsel advisory for startups: contract risk scanner (12 founder-killer patterns: auto-renew traps, uncapped indemnity, vague IP, MFN pricing, missing DPA, one-sided venue, broad non-solicit, perpetual license-back, etc.) and term sheet analyzer (0-100 founder-friendliness across 12 dimensions). 3 in-depth references: contracts playbook (7 startup contract types), IP + regulatory landscape mapping (HIPAA, GDPR, FDA, fintech, EU AI Act, SOC 2 → ISO sequencing), term sheet decoder (full glossary + founder-friendly defaults). Standalone-installable; also bundled in c-level-skills. Stdlib-only. NOT a substitute for licensed counsel.", "version": "1.0.0", "author": { "name": "Alireza Rezvani" @@ -133,7 +133,7 @@ { "name": "chief-data-officer-advisor", "source": "./c-level-advisor/chief-data-officer-advisor", - "description": "Chief Data Officer advisory for startups: AI training data audit (origin \u00d7 class \u00d7 use-case matrix with GDPR Art. 6 + EU AI Act citations), data product strategy picker (warehouse vs lakehouse vs mesh + 6-layer build-vs-buy + 12-month sequencing), data asset valuator (strategic value 0-10 + M&A multiplier with carve-out penalties + 3 ranked productization paths). 4 references answering one decision each: training rights, data product strategy, customer-data-as-asset, data team org evolution. Standalone-installable; also bundled in c-level-skills. Strategic only \u2014 does not duplicate engineering data skills.", + "description": "Chief Data Officer advisory for startups: AI training data audit (origin × class × use-case matrix with GDPR Art. 6 + EU AI Act citations), data product strategy picker (warehouse vs lakehouse vs mesh + 6-layer build-vs-buy + 12-month sequencing), data asset valuator (strategic value 0-10 + M&A multiplier with carve-out penalties + 3 ranked productization paths). 4 references answering one decision each: training rights, data product strategy, customer-data-as-asset, data team org evolution. Standalone-installable; also bundled in c-level-skills. Strategic only — does not duplicate engineering data skills.", "version": "1.0.0", "author": { "name": "Alireza Rezvani" @@ -155,7 +155,7 @@ { "name": "vpe-advisor", "source": "./c-level-advisor/vpe-advisor", - "description": "VP of Engineering advisory: delivery throughput analyzer (DORA 4 metrics + cycle-time bottleneck identification with typical fixes per stage), engineering hiring funnel calculator (7-stage conversion + pipeline gap + weakest-stage fixes from sourcing to offer-accept), engineering team structure designer (squad/tribe model + manager-trigger + director-trigger + span-of-control). 4 in-depth references citing DORA / Spotify / Conway / Google SRE / Larson / Fournier. Standalone-installable; also bundled in c-level-skills. NOT a CTO skill \u2014 VPE owns how the team ships; CTO owns what to build.", + "description": "VP of Engineering advisory: delivery throughput analyzer (DORA 4 metrics + cycle-time bottleneck identification with typical fixes per stage), engineering hiring funnel calculator (7-stage conversion + pipeline gap + weakest-stage fixes from sourcing to offer-accept), engineering team structure designer (squad/tribe model + manager-trigger + director-trigger + span-of-control). 4 in-depth references citing DORA / Spotify / Conway / Google SRE / Larson / Fournier. Standalone-installable; also bundled in c-level-skills. NOT a CTO skill — VPE owns how the team ships; CTO owns what to build.", "version": "1.0.0", "author": { "name": "Alireza Rezvani" @@ -179,7 +179,7 @@ { "name": "chief-customer-officer-advisor", "source": "./c-level-advisor/chief-customer-officer-advisor", - "description": "Chief Customer Officer advisory: retention decomposition analyzer (honest GRR vs NRR; 7-category churn taxonomy with preventable% scoring), customer segmentation designer (4-tier framework, ICP fit scoring across 7 weighted signals, kill list + upgrade candidates), CS coverage calculator (pooled vs named CSM ratio math + 12-month hiring plan with quarterly sequencing). 4 in-depth references each citing 5+ authoritative sources. Standalone-installable; also bundled in c-level-skills. Strategic only \u2014 does not duplicate business-growth tactical CS skills.", + "description": "Chief Customer Officer advisory: retention decomposition analyzer (honest GRR vs NRR; 7-category churn taxonomy with preventable% scoring), customer segmentation designer (4-tier framework, ICP fit scoring across 7 weighted signals, kill list + upgrade candidates), CS coverage calculator (pooled vs named CSM ratio math + 12-month hiring plan with quarterly sequencing). 4 in-depth references each citing 5+ authoritative sources. Standalone-installable; also bundled in c-level-skills. Strategic only — does not duplicate business-growth tactical CS skills.", "version": "1.0.0", "author": { "name": "Alireza Rezvani" @@ -201,7 +201,7 @@ { "name": "chief-ai-officer-advisor", "source": "./c-level-advisor/chief-ai-officer-advisor", - "description": "Chief AI Officer advisory for startups: model build-vs-buy calculator (API vs fine-tune vs build with 3-year TCO across 6 paths + breakeven that balances economics with practical feasibility), AI risk classifier (EU AI Act tier with 7 Article citations + US state patchwork: NYC LL 144, CO AI Act, IL HB 53, CA SB 1001, IL BIPA + industry overlays for FDA AI/ML, CFPB Circular 2023-03, NYDFS Reg 23, NAIC, ECOA, Fed SR 11-7), AI cost economics (API vs self-hosted breakeven with 2026 pricing across A100/H100, utilization reality, hidden costs). 4 in-depth references each citing 5+ authoritative sources. Standalone-installable; also bundled in c-level-skills. Strategic only \u2014 does not duplicate engineering AI/ML skills.", + "description": "Chief AI Officer advisory for startups: model build-vs-buy calculator (API vs fine-tune vs build with 3-year TCO across 6 paths + breakeven that balances economics with practical feasibility), AI risk classifier (EU AI Act tier with 7 Article citations + US state patchwork: NYC LL 144, CO AI Act, IL HB 53, CA SB 1001, IL BIPA + industry overlays for FDA AI/ML, CFPB Circular 2023-03, NYDFS Reg 23, NAIC, ECOA, Fed SR 11-7), AI cost economics (API vs self-hosted breakeven with 2026 pricing across A100/H100, utilization reality, hidden costs). 4 in-depth references each citing 5+ authoritative sources. Standalone-installable; also bundled in c-level-skills. Strategic only — does not duplicate engineering AI/ML skills.", "version": "1.0.0", "author": { "name": "Alireza Rezvani" @@ -415,7 +415,7 @@ { "name": "autoresearch-agent", "source": "./engineering/autoresearch-agent", - "description": "Autonomous experiment loop \u2014 optimize any file by a measurable metric. 5 slash commands (/ar:setup, /ar:run, /ar:loop, /ar:status, /ar:resume), 8 built-in evaluators, configurable loop intervals (10min to monthly).", + "description": "Autonomous experiment loop — optimize any file by a measurable metric. 5 slash commands (/ar:setup, /ar:run, /ar:loop, /ar:status, /ar:resume), 8 built-in evaluators, configurable loop intervals (10min to monthly).", "version": "2.2.2", "author": { "name": "Alireza Rezvani" @@ -480,7 +480,7 @@ { "name": "agenthub", "source": "./engineering/agenthub", - "description": "Multi-agent collaboration \u2014 spawn N parallel subagents that compete on code optimization, content drafts, research approaches, or any task that benefits from diverse solutions. 7 slash commands (/hub:init, /hub:spawn, /hub:status, /hub:eval, /hub:merge, /hub:board, /hub:run), agent templates, DAG-based orchestration, LLM judge mode, message board coordination.", + "description": "Multi-agent collaboration — spawn N parallel subagents that compete on code optimization, content drafts, research approaches, or any task that benefits from diverse solutions. 7 slash commands (/hub:init, /hub:spawn, /hub:status, /hub:eval, /hub:merge, /hub:board, /hub:run), agent templates, DAG-based orchestration, LLM judge mode, message board coordination.", "version": "2.2.2", "author": { "name": "Alireza Rezvani" @@ -539,7 +539,7 @@ { "name": "docker-development", "source": "./engineering/docker-development", - "description": "Docker and container development \u2014 Dockerfile optimization, docker-compose orchestration, multi-stage builds, security hardening, and CI/CD container pipelines.", + "description": "Docker and container development — Dockerfile optimization, docker-compose orchestration, multi-stage builds, security hardening, and CI/CD container pipelines.", "version": "2.2.2", "author": { "name": "Alireza Rezvani" @@ -556,7 +556,7 @@ { "name": "helm-chart-builder", "source": "./engineering/helm-chart-builder", - "description": "Helm chart development \u2014 chart scaffolding, values design, template patterns, dependency management, and Kubernetes deployment strategies.", + "description": "Helm chart development — chart scaffolding, values design, template patterns, dependency management, and Kubernetes deployment strategies.", "version": "2.2.2", "author": { "name": "Alireza Rezvani" @@ -573,7 +573,7 @@ { "name": "terraform-patterns", "source": "./engineering/terraform-patterns", - "description": "Terraform infrastructure-as-code \u2014 module design patterns, state management, provider configuration, CI/CD integration, and multi-environment strategies.", + "description": "Terraform infrastructure-as-code — module design patterns, state management, provider configuration, CI/CD integration, and multi-environment strategies.", "version": "2.2.2", "author": { "name": "Alireza Rezvani" @@ -590,7 +590,7 @@ { "name": "research-summarizer", "source": "./product-team/research-summarizer", - "description": "Structured research summarization \u2014 summarize academic papers, market research, user interviews, and competitive analysis into actionable insights.", + "description": "Structured research summarization — summarize academic papers, market research, user interviews, and competitive analysis into actionable insights.", "version": "2.2.2", "author": { "name": "Alireza Rezvani" @@ -607,7 +607,7 @@ { "name": "code-tour", "source": "./engineering/code-tour", - "description": "Create CodeTour .tour files \u2014 persona-targeted, step-by-step walkthroughs that link to real files and line numbers. 10 developer personas, all CodeTour step types, SMIG description formula.", + "description": "Create CodeTour .tour files — persona-targeted, step-by-step walkthroughs that link to real files and line numbers. 10 developer personas, all CodeTour step types, SMIG description formula.", "version": "2.2.2", "author": { "name": "Alireza Rezvani" @@ -760,7 +760,7 @@ { "name": "kubernetes-operator", "source": "./engineering/kubernetes-operator", - "description": "End-to-end Kubernetes Operator discipline: CRD design, reconcile-loop patterns, and OperatorHub Capability Levels. Ships CRD validator, reconcile-loop linter, and capability auditor (3 stdlib Python tools), 4 references on the operator pattern + CRD design + reconcile patterns + framework comparison (controller-runtime/kubebuilder/operator-sdk/metacontroller/KOPF), CRD + Go controller skeletons, and /operator-audit slash command. NOT a generic k8s skill \u2014 specifically the Operator pattern.", + "description": "End-to-end Kubernetes Operator discipline: CRD design, reconcile-loop patterns, and OperatorHub Capability Levels. Ships CRD validator, reconcile-loop linter, and capability auditor (3 stdlib Python tools), 4 references on the operator pattern + CRD design + reconcile patterns + framework comparison (controller-runtime/kubebuilder/operator-sdk/metacontroller/KOPF), CRD + Go controller skeletons, and /operator-audit slash command. NOT a generic k8s skill — specifically the Operator pattern.", "version": "2.4.0", "author": { "name": "Alireza Rezvani" @@ -825,7 +825,7 @@ { "name": "write-a-skill", "source": "./engineering/write-a-skill", - "description": "Skill-author skill: create new agent skills with proper structure, progressive disclosure, and bundled resources. Derived from Matt Pocock's MIT-licensed write-a-skill with: (1) 3 stdlib Python validation tools (description validator, structure validator, review-checklist runner \u2014 all enforcing Matt's 6-item checklist), (2) 4 references citing 7-8 authoritative sources each (progressive disclosure principles, description design patterns, quality gates, companion tooling), (3) cs-skill-author persona agent + /cs:write-a-skill slash command. Matt's voice and 3-phase workflow (Gather \u2192 Draft \u2192 Review) preserved verbatim per MIT.", + "description": "Skill-author skill: create new agent skills with proper structure, progressive disclosure, and bundled resources. Derived from Matt Pocock's MIT-licensed write-a-skill with: (1) 3 stdlib Python validation tools (description validator, structure validator, review-checklist runner — all enforcing Matt's 6-item checklist), (2) 4 references citing 7-8 authoritative sources each (progressive disclosure principles, description design patterns, quality gates, companion tooling), (3) cs-skill-author persona agent + /cs:write-a-skill slash command. Matt's voice and 3-phase workflow (Gather → Draft → Review) preserved verbatim per MIT.", "version": "2.6.0", "author": { "name": "Alireza Rezvani" @@ -881,7 +881,7 @@ { "name": "handoff", "source": "./engineering/handoff", - "description": "Conversation-handoff document generator. Compacts the current session into a markdown handoff for a fresh agent \u2014 references existing artifacts (PRDs, plans, ADRs, issues, commits) by path/URL instead of duplicating them. Derived from Matt Pocock's MIT-licensed handoff with: (1) 3 stdlib Python tools (template generator tailored to 5 next-session emphases, artifact deduplicator across 5 categories of duplication, skill recommender matching content to 14 skills in this repo), (2) 4 references citing 7-8 sources (handoff structure, deduplication discipline, next-session skill matching, companion tooling), (3) cs-handoff-author persona agent + /cs:handoff slash command. Matt's no-duplication discipline + mktemp convention preserved verbatim per MIT.", + "description": "Conversation-handoff document generator. Compacts the current session into a markdown handoff for a fresh agent — references existing artifacts (PRDs, plans, ADRs, issues, commits) by path/URL instead of duplicating them. Derived from Matt Pocock's MIT-licensed handoff with: (1) 3 stdlib Python tools (template generator tailored to 5 next-session emphases, artifact deduplicator across 5 categories of duplication, skill recommender matching content to 14 skills in this repo), (2) 4 references citing 7-8 sources (handoff structure, deduplication discipline, next-session skill matching, companion tooling), (3) cs-handoff-author persona agent + /cs:handoff slash command. Matt's no-duplication discipline + mktemp convention preserved verbatim per MIT.", "version": "2.6.0", "author": { "name": "Alireza Rezvani" @@ -919,7 +919,7 @@ { "name": "capture-skill", "source": "./productivity/capture", - "description": "Brain-dump-to-action workspace skill. Routes vague captures into discoverable actions via classify\u2192cluster\u2192connect\u2192clarify intake. Path-B from megaprompt 05.", + "description": "Brain-dump-to-action workspace skill. Routes vague captures into discoverable actions via classify→cluster→connect→clarify intake. Path-B from megaprompt 05.", "version": "2.7.0", "author": { "name": "Alireza Rezvani" @@ -1153,7 +1153,7 @@ { "name": "aeo", "source": "./marketing-skill/skills/aeo", - "description": "Answer Engine Optimization (AEO) skill \u2014 optimize content to be cited by AI language models (ChatGPT, Perplexity, Claude, Gemini, Mistral) as authoritative sources. Distinct from SEO (which optimizes for search rankings), AEO optimizes for citation in LLM-generated responses. 3 stdlib Python tools (aeo_audit, aeo_optimizer, citation_tracker), 3 references citing 8 sources each, industry-aware thresholds for 8 industries (saas/healthcare/finance/legal/ecommerce/b2b/media/education). Ported from alirezarezvani/aeo-box.", + "description": "Answer Engine Optimization (AEO) skill — optimize content to be cited by AI language models (ChatGPT, Perplexity, Claude, Gemini, Mistral) as authoritative sources. Distinct from SEO (which optimizes for search rankings), AEO optimizes for citation in LLM-generated responses. 3 stdlib Python tools (aeo_audit, aeo_optimizer, citation_tracker), 3 references citing 8 sources each, industry-aware thresholds for 8 industries (saas/healthcare/finance/legal/ecommerce/b2b/media/education). Ported from alirezarezvani/aeo-box.", "version": "2.7.3", "author": { "name": "Alireza Rezvani" @@ -1175,7 +1175,7 @@ { "name": "security-guidance", "source": "./engineering/security-guidance", - "description": "PreToolUse security reminder hook for Claude Code. Catches 12 common security anti-patterns in Edit/Write/MultiEdit operations BEFORE they happen \u2014 command injection (exec, os.system, subprocess shell=True), XSS (innerHTML, dangerouslySetInnerHTML, document.write), SQL injection (f-string queries, .format), unsafe deserialization (pickle, yaml.unsafe_load), code injection (eval, new Function), and GitHub Actions workflow injection. Session-state caching prevents duplicate warnings; 30-day auto-cleanup. Disable per-session with ENABLE_SECURITY_REMINDER=0. Ported from David Dworken at Anthropic.", + "description": "PreToolUse security reminder hook for Claude Code. Catches 12 common security anti-patterns in Edit/Write/MultiEdit operations BEFORE they happen — command injection (exec, os.system, subprocess shell=True), XSS (innerHTML, dangerouslySetInnerHTML, document.write), SQL injection (f-string queries, .format), unsafe deserialization (pickle, yaml.unsafe_load), code injection (eval, new Function), and GitHub Actions workflow injection. Session-state caching prevents duplicate warnings; 30-day auto-cleanup. Disable per-session with ENABLE_SECURITY_REMINDER=0. Ported from David Dworken at Anthropic.", "version": "2.7.3", "author": { "name": "Alireza Rezvani" @@ -1273,6 +1273,24 @@ "grill-with-docs" ], "category": "commercial" + }, + { + "name": "universal-scraping-architect", + "source": "./engineering/universal-scraping-architect", + "description": "A universal scraping skill with intelligent routing, token budget tracking, and quota awareness. Supports Firecrawl and local Python extraction.", + "version": "2.1.2", + "author": { + "name": "Mehansh Barthwal" + }, + "keywords": [ + "scraping", + "data-extraction", + "firecrawl", + "beautifulsoup4", + "pandas", + "automation" + ], + "category": "development" } ] -} +} \ No newline at end of file diff --git a/.idea/.gitignore b/.idea/.gitignore new file mode 100644 index 00000000..30cf57ed --- /dev/null +++ b/.idea/.gitignore @@ -0,0 +1,10 @@ +# Default ignored files +/shelf/ +/workspace.xml +# Editor-based HTTP Client requests +/httpRequests/ +# Ignored default folder with query files +/queries/ +# Datasource local storage ignored files +/dataSources/ +/dataSources.local.xml diff --git a/.idea/claude-skills.iml b/.idea/claude-skills.iml new file mode 100644 index 00000000..55cb08ba --- /dev/null +++ b/.idea/claude-skills.iml @@ -0,0 +1,24 @@ + + + + + + + + + + + + + + + + + \ No newline at end of file diff --git a/.idea/inspectionProfiles/profiles_settings.xml b/.idea/inspectionProfiles/profiles_settings.xml new file mode 100644 index 00000000..105ce2da --- /dev/null +++ b/.idea/inspectionProfiles/profiles_settings.xml @@ -0,0 +1,6 @@ + + + + \ No newline at end of file diff --git a/.idea/modules.xml b/.idea/modules.xml new file mode 100644 index 00000000..8210aa20 --- /dev/null +++ b/.idea/modules.xml @@ -0,0 +1,8 @@ + + + + + + + + \ No newline at end of file diff --git a/.idea/vcs.xml b/.idea/vcs.xml new file mode 100644 index 00000000..35eb1ddf --- /dev/null +++ b/.idea/vcs.xml @@ -0,0 +1,6 @@ + + + + + + \ No newline at end of file diff --git a/engineering/universal-scraping-architect/LICENSE b/engineering/universal-scraping-architect/LICENSE deleted file mode 100644 index 52def8d8..00000000 --- a/engineering/universal-scraping-architect/LICENSE +++ /dev/null @@ -1,21 +0,0 @@ -MIT License - -Copyright (c) 2026 Mehansh Barthwal - -Permission is hereby granted, free of charge, to any person obtaining a copy -of this software and associated documentation files (the "Software"), to deal -in the Software without restriction, including without limitation the rights -to use, copy, modify, merge, publish, distribute, sublicense, and/or sell -copies of the Software, and to permit persons to whom the Software is -furnished to do so, subject to the following conditions: - -The above copyright notice and this permission notice shall be included in all -copies or substantial portions of the Software. - -THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR -IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, -FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE -AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER -LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, -OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE -SOFTWARE. diff --git a/engineering/universal-scraping-architect/README.md b/engineering/universal-scraping-architect/README.md deleted file mode 100644 index 53aeef04..00000000 --- a/engineering/universal-scraping-architect/README.md +++ /dev/null @@ -1,212 +0,0 @@ -# Universal Scraping Architect - -> A robust, general-purpose scraping and data extraction framework designed as a reusable **Skill** for AI Agents and LLMs. - -[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT) -[![Python 3.8+](https://img.shields.io/badge/python-3.8+-blue.svg)](https://www.python.org/downloads/) -[![Firecrawl Ready](https://img.shields.io/badge/Firecrawl-Ready-orange.svg)](https://firecrawl.dev) - ---- - -## What This Is - -Most scraping scripts are brittle one-offs that break the moment a page layout changes, a column gets renamed, or a file format shifts slightly. This framework is built differently — it treats every extraction task as a **complete data pipeline**, with intelligent routing, validation, token tracking, and clean outputs baked in from the start. - -It was built specifically to be dropped into an AI agent's skill library (like Claude's skill system), so that any scraping or data extraction task — whether it's a live public website, a local PDF, an Excel file, or a nested JSON — gets handled with the same consistent, robust approach every time. - ---- - -## Key Features - -**Intelligent Approach Routing** — Before writing a single line of code, the framework decides whether the task calls for Firecrawl (dynamic web, search-first, or bulk crawl), traditional local scraping (local files, static HTML, private data), or a hybrid of both. It states the choice and the reason. - -**Firecrawl Integration (5 Paths)** — Full support for Firecrawl's CLI, REST API, and SDK across five clearly defined paths: live web data, app/product integration, finished deliverables, auth-only setup, and direct REST without installation. - -**Token Budget Tracking** — Every extraction that feeds an LLM estimates token volume against a configurable context limit before processing, warning early if the output is over budget. - -**Firecrawl Quota Safety** — Single-key rule enforced by default. Estimates page/credit usage before large crawl jobs and only prompts for a new key when the current one is actually exhausted or a job would genuinely exceed ~1000 requests. - -**Validation at Every Step** — Required fields are checked, row counts are logged, empty outputs are caught, and duplicate rows are flagged before anything gets saved. - -**Checkpointing for Large Jobs** — Progress is saved during multi-page or multi-file extractions so a failure mid-job doesn't mean starting from scratch. - -**Security and Ethical Scraping** — API keys are loaded from environment variables only, never hardcoded or printed. robots.txt, rate limits, and terms of service are respected. Private files are never sent to external APIs without explicit approval. - ---- - -## Supported Sources - -| Category | Examples | -|---|---| -| **Live Web** | Public URLs, dynamic pages, SPAs, paginated sites | -| **Search-First** | Topic queries, keyword discovery, entity research | -| **Bulk Crawl** | Full site sections, documentation sites, large domains | -| **Local Files** | PDF, DOCX, Excel (.xlsx/.xls), CSV, JSON, XML, ZIP | -| **Scanned Docs** | OCR pipelines for image-based PDFs | -| **APIs** | REST APIs, paginated API responses, JSON data feeds | -| **Databases** | CSV exports, SQLite, structured data dumps | - ---- - -## Output Formats - -The framework can save clean outputs as CSV, Excel, JSON, Markdown, TXT, SQLite, Parquet, or HTML depending on what the task needs. If not specified, it defaults sensibly — CSV for structured tabular data, JSON for nested structures, Markdown for clean web page text. - ---- - -## Project Structure - -``` -universal-scraping-architect/ -├── SKILL.md # Core LLM skill definition (the agent's system prompt) -├── README.md # This file -├── LICENSE # MIT License -├── .gitignore # Standard Python/project ignores -├── requirements.txt # Python dependencies -└── examples/ - ├── firecrawl_example.py # Path C workflow: Firecrawl → clean markdown output - └── local_bs4_example.py # Traditional scraping: static HTML → validated CSV -``` - ---- - -## How To Use This As An AI Skill - -Drop `SKILL.md` into your agent's system prompt or tool context. The agent will then adopt the Universal Scraping Architect approach for any extraction task — routing correctly, validating outputs, tracking tokens, and producing copy-paste-ready pipeline code rather than fragile one-offs. - -This is how it appears in Claude's skill system: - -```yaml -name: universal-scraping-architect -description: Use this skill for any scraping, crawling, extraction, parsing, web - research, document processing, dataset preparation, Firecrawl workflow, - local file extraction, API extraction, PDF/Excel/CSV/JSON/XML parsing, - validation-heavy data pipeline, or repeatable clean-output scraping task. -``` - ---- - -## Quick Start (For Developers Running the Examples) - -**Install dependencies:** - -```bash -pip install -r requirements.txt -``` - -**For Firecrawl workflows, set your API key:** - -```bash -# Linux / macOS -export FIRECRAWL_API_KEY="fc-YOUR_API_KEY_HERE" - -# Windows PowerShell -$env:FIRECRAWL_API_KEY = "fc-YOUR_API_KEY_HERE" - -# Windows Command Prompt -set FIRECRAWL_API_KEY=fc-YOUR_API_KEY_HERE -``` - -Or add it to a `.env` file in your project root (never commit this file): - -```dotenv -FIRECRAWL_API_KEY=fc-YOUR_API_KEY_HERE -``` - -**Run the Firecrawl example:** - -```bash -python examples/firecrawl_example.py -``` - -**Run the local scraping example:** - -```bash -python examples/local_bs4_example.py -``` - ---- - -## The 15-Step Pipeline - -Every task the framework handles follows this sequence: - -1. Understand the source -2. Choose the most appropriate extraction approach -3. Configure task-specific settings -4. Extract safely -5. Handle pagination, layout changes, dynamic content, or file variations -6. Clean the extracted data -7. Normalize structure and field names -8. Validate the result -9. Track token/data volume if LLM processing is involved -10. Estimate Firecrawl usage/quota before large Firecrawl jobs -11. Handle errors clearly -12. Save clean outputs -13. Save logs and checkpoints when useful -14. Print a final summary -15. Explain what was done and what can be customized - ---- - -## Firecrawl Path Reference - -| Path | When To Use | -|---|---| -| **Path A** | Need live web data right now during the current session | -| **Path B** | Building an app or product that calls Firecrawl from code | -| **Path C** | Need a finished deliverable — research brief, SEO audit, lead list, etc. | -| **Path D** | Need to set up an account or API key first | -| **Path E** | Don't want to install anything — use the REST API directly | - -Install command (covers all paths): - -```bash -npx -y firecrawl-cli@latest init --all --browser -``` - ---- - -## When To Use Firecrawl vs. Local Scraping - -**Use Firecrawl when:** -- The source is a public URL and you want clean, reliable extraction -- The page is dynamic (JavaScript-rendered, SPA, requires interaction) -- You need search-first discovery before you know the URLs -- You're crawling many pages across a domain -- You want to wire Firecrawl into an app or agentic workflow - -**Use local/traditional scraping when:** -- The source is a local file (PDF, Excel, CSV, JSON, XML) -- The data is private or sensitive and shouldn't leave your machine -- You're using official downloads/APIs where scraping isn't needed -- Simple static HTML where Firecrawl would be overkill -- Custom parsing logic with pandas, pdfplumber, openpyxl, etc. is the right fit - -**Use a hybrid when:** -- Firecrawl handles web extraction, then Python cleans and structures the output -- Firecrawl discovers URLs, then local code processes and saves the dataset -- Web content gets merged with local files - ---- - -## Security Notes - -- API keys are always loaded from environment variables — never hardcoded -- `.env` files are in `.gitignore` and should never be committed -- Real keys are never printed in logs or included in output files -- Private or sensitive files are never sent to Firecrawl without explicit user approval -- One active Firecrawl API key is used at a time (no rotation by default) -- robots.txt and rate limits are respected - ---- - -## Author - -**Mehansh Barthwal** - ---- - -## License - -This project is licensed under the MIT License — see [LICENSE](LICENSE) for details. diff --git a/engineering/universal-scraping-architect/agents/agents/cs-scraping-architect.md b/engineering/universal-scraping-architect/agents/cs-scraping-architect.md similarity index 100% rename from engineering/universal-scraping-architect/agents/agents/cs-scraping-architect.md rename to engineering/universal-scraping-architect/agents/cs-scraping-architect.md diff --git a/engineering/universal-scraping-architect/commands/commands/cs-scrape.md b/engineering/universal-scraping-architect/commands/cs-scrape.md similarity index 100% rename from engineering/universal-scraping-architect/commands/commands/cs-scrape.md rename to engineering/universal-scraping-architect/commands/cs-scrape.md diff --git a/engineering/universal-scraping-architect/references/references/parsing-and-data-extraction.md b/engineering/universal-scraping-architect/references/references/parsing-and-data-extraction.md deleted file mode 100644 index 221765d2..00000000 --- a/engineering/universal-scraping-architect/references/references/parsing-and-data-extraction.md +++ /dev/null @@ -1,10 +0,0 @@ -# Parsing and Data Extraction Standards - -Detailed strategies for BeautifulSoup4 and Pandas integration for cleaning scraped data. - -### Authoritative Sources -1. [BeautifulSoup4 Documentation](https://www.crummy.com/software/BeautifulSoup/bs4/doc/) -2. [Pandas Data Cleaning Guide](https://pandas.pydata.org/docs/user_guide/10min.html) -3. [W3C HTML Living Standard](https://html.spec.whatwg.org/multipage/) -4. [CSS Selectors Level 4 (W3C)](https://www.w3.org/TR/selectors-4/) -5. [Mozilla MDN - HTTP Headers (User-Agent)](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/User-Agent) \ No newline at end of file diff --git a/engineering/universal-scraping-architect/references/references/references/scraping-ethics-security.md b/engineering/universal-scraping-architect/references/scraping-ethics-security.md similarity index 100% rename from engineering/universal-scraping-architect/references/references/references/scraping-ethics-security.md rename to engineering/universal-scraping-architect/references/scraping-ethics-security.md