mirror of
https://github.com/alirezarezvani/claude-skills.git
synced 2026-09-10 22:41:16 +00:00
Two small polish tasks ahead of any future Pages deploy. 1. Add /cs:* command nav entries (22 new entries) The 21 c-level-agents-* sub-skill pages now exist (since #632) but weren't surfaced in mkdocs.yml sidebar nav. Added a "Founder-Mode Commands" nested section under C-Level Advisory with: - c-level-agents index - 10 forcing-question reviews (/cs:cfo-review through /cs:vpe-review) - 5 strategic sprint pipeline commands (brief/boardroom/decide/execute/post-mortem) - 4 meta+safety commands (founder-mode/onboard/cross-eval/freeze) - /cs:office-hours 2. Clear 33 mkdocs INFO warnings mkdocs build was emitting 33 INFO-level warnings during the docs deploy. Pre-existing noise; not regressions. Three categories: a) 27 unrecognized-link warnings: relative links like `[Skills](skills/)` that mkdocs flags because the path doesn't end in .md. Fix: added explicit `index.md` suffix in 3 manual doc files. - docs/index.md: 15 links - docs/skills/index.md: 11 links - docs/custom-gpts.md: 1 link b) 2 anchor warnings in scrum-master TOC: links pointed to `#analysis-tools--usage` and `#key-metrics--targets` (double hyphen from ampersand) but mkdocs Material's slugify produces single-hyphen slugs. Fix: changed to `#analysis-tools-usage` and `#key-metrics-targets`. c) 4 anchor warnings in senior-computer-vision + senior-data-engineer TOCs: links pointed to non-existent sections. - senior-computer-vision: `#common-commands` TOC entry — no such heading anywhere; removed the entry. - senior-data-engineer: 3 sub-bullets pointing to `#workflow-1-...`, `#workflow-2-...`, `#workflow-3-...` — no such headings (only a parent `## Workflows`); removed the sub-bullets. Verification: - mkdocs build now emits 0 INFO warnings - karpathy diff_surgeon: 0 findings on staged diff - All 22 new nav entries verified to point to existing HTML pages - generate-docs.py re-run picked up the upstream SKILL.md fixes; docs/skills/ now matches sources 10 files changed, +54/-39. After the next dev->main release, the Pages deploy will have: - Cleaner build output (no INFO noise) - Fully discoverable /cs:* command pages in the sidebar nav https://claude.ai/code/session_012WtZMm5NJHqkYoRqA9fHMN
6 KiB
6 KiB
| title | description |
|---|---|
| Senior Data Engineer — Agent Skill & Codex Plugin | Data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure. Expertise in Python, SQL, Spark, Airflow, dbt. Agent skill for Claude Code, Codex CLI, Gemini CLI, OpenClaw. |
Senior Data Engineer
:material-code-braces: Engineering - Core
:material-identifier: `senior-data-engineer`
:material-github: Source
Install:
claude /plugin install engineering-skills
Production-grade data engineering skill for building scalable, reliable data systems.
Table of Contents
- Trigger Phrases
- Quick Start
- Workflows
- Architecture Decision Framework
- Tech Stack
- Reference Documentation
- Troubleshooting
Trigger Phrases
Activate this skill when you see:
Pipeline Design:
- "Design a data pipeline for..."
- "Build an ETL/ELT process..."
- "How should I ingest data from..."
- "Set up data extraction from..."
Architecture:
- "Should I use batch or streaming?"
- "Lambda vs Kappa architecture"
- "How to handle late-arriving data"
- "Design a data lakehouse"
Data Modeling:
- "Create a dimensional model..."
- "Star schema vs snowflake"
- "Implement slowly changing dimensions"
- "Design a data vault"
Data Quality:
- "Add data validation to..."
- "Set up data quality checks"
- "Monitor data freshness"
- "Implement data contracts"
Performance:
- "Optimize this Spark job"
- "Query is running slow"
- "Reduce pipeline execution time"
- "Tune Airflow DAG"
Quick Start
Core Tools
# Generate pipeline orchestration config
python scripts/pipeline_orchestrator.py generate \
--type airflow \
--source postgres \
--destination snowflake \
--schedule "0 5 * * *"
# Validate data quality
python scripts/data_quality_validator.py validate \
--input data/sales.parquet \
--schema schemas/sales.json \
--checks freshness,completeness,uniqueness
# Optimize ETL performance
python scripts/etl_performance_optimizer.py analyze \
--query queries/daily_aggregation.sql \
--engine spark \
--recommend
Workflows
→ See references/workflows.md for details
Architecture Decision Framework
Use this framework to choose the right approach for your data pipeline.
Batch vs Streaming
| Criteria | Batch | Streaming |
|---|---|---|
| Latency requirement | Hours to days | Seconds to minutes |
| Data volume | Large historical datasets | Continuous event streams |
| Processing complexity | Complex transformations, ML | Simple aggregations, filtering |
| Cost sensitivity | More cost-effective | Higher infrastructure cost |
| Error handling | Easier to reprocess | Requires careful design |
Decision Tree:
Is real-time insight required?
├── Yes → Use streaming
│ └── Is exactly-once semantics needed?
│ ├── Yes → Kafka + Flink/Spark Structured Streaming
│ └── No → Kafka + consumer groups
└── No → Use batch
└── Is data volume > 1TB daily?
├── Yes → Spark/Databricks
└── No → dbt + warehouse compute
Lambda vs Kappa Architecture
| Aspect | Lambda | Kappa |
|---|---|---|
| Complexity | Two codebases (batch + stream) | Single codebase |
| Maintenance | Higher (sync batch/stream logic) | Lower |
| Reprocessing | Native batch layer | Replay from source |
| Use case | ML training + real-time serving | Pure event-driven |
When to choose Lambda:
- Need to train ML models on historical data
- Complex batch transformations not feasible in streaming
- Existing batch infrastructure
When to choose Kappa:
- Event-sourced architecture
- All processing can be expressed as stream operations
- Starting fresh without legacy systems
Data Warehouse vs Data Lakehouse
| Feature | Warehouse (Snowflake/BigQuery) | Lakehouse (Delta/Iceberg) |
|---|---|---|
| Best for | BI, SQL analytics | ML, unstructured data |
| Storage cost | Higher (proprietary format) | Lower (open formats) |
| Flexibility | Schema-on-write | Schema-on-read |
| Performance | Excellent for SQL | Good, improving |
| Ecosystem | Mature BI tools | Growing ML tooling |
Tech Stack
| Category | Technologies |
|---|---|
| Languages | Python, SQL, Scala |
| Orchestration | Airflow, Prefect, Dagster |
| Transformation | dbt, Spark, Flink |
| Streaming | Kafka, Kinesis, Pub/Sub |
| Storage | S3, GCS, Delta Lake, Iceberg |
| Warehouses | Snowflake, BigQuery, Redshift, Databricks |
| Quality | Great Expectations, dbt tests, Monte Carlo |
| Monitoring | Prometheus, Grafana, Datadog |
Reference Documentation
1. Data Pipeline Architecture
See references/data_pipeline_architecture.md for:
- Lambda vs Kappa architecture patterns
- Batch processing with Spark and Airflow
- Stream processing with Kafka and Flink
- Exactly-once semantics implementation
- Error handling and dead letter queues
2. Data Modeling Patterns
See references/data_modeling_patterns.md for:
- Dimensional modeling (Star/Snowflake)
- Slowly Changing Dimensions (SCD Types 1-6)
- Data Vault modeling
- dbt best practices
- Partitioning and clustering
3. DataOps Best Practices
See references/dataops_best_practices.md for:
- Data testing frameworks
- Data contracts and schema validation
- CI/CD for data pipelines
- Observability and lineage
- Incident response
Troubleshooting
→ See references/troubleshooting.md for details