SkillAgentSearch skills...

senior-data-engineer

World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure. Expertise in Python, SQL, Spark, Airflow, dbt, Kafka, Flink, Kinesis, and modern data stack.

Install / Use

npx skills add benchflow-ai/skillsbench --skill senior-data-engineer

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

94/100

Category

Automation

Supported Platforms

Universal

Our assessment of senior-data-engineer

senior-data-engineer scores 94/100 on our quality scale, 366th of 3,055 Automation skills we index (top 12%).

Its SKILL.md is 23 KB long, well organised into 70 sections with 10 code examples: a thorough specification that gives an agent plenty to work with.

With 1,813 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
30/30
Structure
20/20
Description
15/15
Adoption
14/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated about 2 months ago, so senior-data-engineer is actively maintained.
  • It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

Safety scan

No issues found

Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.

Automated pattern scan on 2026-10-02. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.

senior-data-engineer compared with similar skills

All 4 of these similar skills score higher than senior-data-engineer; compare them before choosing.

SkillScoreStarsUpdatedFormat
senior-data-engineer (this skill)by benchflow-ai941.8k2mo agoSKILL.md
claude-memby thedotmack10095.2ktodayCLAUDE.md
Agent-Reachby Panniantong10087.6k16d agoCLAUDE.md
headroomby headroomlabs-ai10074.3ktodayCLAUDE.md
rufloby ruvnet10073.7ktodayCLAUDE.md

Frequently asked questions

How do I install senior-data-engineer?
Run npx skills add benchflow-ai/skillsbench --skill senior-data-engineer. The install tabs above show the steps for each supported agent.
Which AI agents does senior-data-engineer work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is senior-data-engineer safe to use?
Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is senior-data-engineer still maintained?
The repository was last updated about 2 months ago, so senior-data-engineer is actively maintained.

=== CORE IDENTITY ===

name: senior-data-engineer title: Senior Data Engineer Skill Package description: World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure. Expertise in Python, SQL, Spark, Airflow, dbt, Kafka, Flink, Kinesis, and modern data stack. Includes data modeling, pipeline orchestration, data quality, streaming quality monitoring, and DataOps. Use when designing data architectures, building batch or streaming data pipelines, optimizing data workflows, or implementing data governance. domain: engineering subdomain: data-engineering

=== WEBSITE DISPLAY ===

difficulty: advanced time-saved: "TODO: Quantify time savings" frequency: "TODO: Estimate usage frequency" use-cases:

  • Designing data pipelines for ETL/ELT processes
  • Building data warehouses and data lakes
  • Implementing data quality and governance frameworks
  • Creating analytics dashboards and reporting
  • Building real-time streaming pipelines with Kafka and Flink
  • Implementing exactly-once streaming semantics
  • Monitoring streaming quality (consumer lag, data freshness, schema drift)

=== RELATIONSHIPS ===

related-agents: [] related-skills: [] related-commands: [] orchestrated-by: []

=== TECHNICAL ===

dependencies: scripts: [] references: [] assets: [] compatibility: "Python 3.8+; platforms: macos, linux, windows" tech-stack:

  • Python
  • SQL
  • Apache Spark
  • Airflow
  • dbt
  • Apache Kafka
  • Apache Flink
  • AWS Kinesis
  • Spark Structured Streaming
  • Kafka Streams
  • PostgreSQL
  • BigQuery
  • Snowflake
  • Docker
  • Schema Registry

=== EXAMPLES ===

examples:

title: Example Usage
input: "TODO: Add example input for senior-data-engineer"
output: "TODO: Add expected output"

=== ANALYTICS ===

stats: downloads: 0 stars: 0 rating: 0.0 reviews: 0

=== VERSIONING ===

version: v2.0.0 author: Claude Skills Team contributors: [] created: 2025-10-20 updated: 2025-12-16 license: MIT

=== DISCOVERABILITY ===

tags: [architecture, data, design, engineer, engineering, senior, streaming, kafka, flink, real-time] featured: false verified: true

Senior Data Engineer

Core Capabilities

  • Batch Pipeline Orchestration - Design and implement production-ready ETL/ELT pipelines with Airflow, intelligent dependency resolution, retry logic, and comprehensive monitoring
  • Real-Time Streaming - Build event-driven streaming pipelines with Kafka, Flink, Kinesis, and Spark Streaming with exactly-once semantics and sub-second latency
  • Data Quality Management - Comprehensive batch and streaming data quality validation covering completeness, accuracy, consistency, timeliness, and validity
  • Streaming Quality Monitoring - Track consumer lag, data freshness, schema drift, throughput, and dead letter queue rates for streaming pipelines
  • Performance Optimization - Analyze and optimize pipeline performance with query optimization, Spark tuning, and cost analysis recommendations

Key Workflows

Workflow 1: Build ETL Pipeline

Time: 2-4 hours

Steps:

  1. Design pipeline architecture using Lambda, Kappa, or Medallion pattern
  2. Configure YAML pipeline definition with sources, transformations, targets
  3. Generate Airflow DAG with pipeline_orchestrator.py
  4. Define data quality validation rules
  5. Deploy and configure monitoring/alerting

Expected Output: Production-ready ETL pipeline with 99%+ success rate, automated quality checks, and comprehensive monitoring

Workflow 2: Build Real-Time Streaming Pipeline

Time: 3-5 days

Steps:

  1. Select streaming architecture (Kappa vs Lambda) based on requirements
  2. Configure streaming pipeline YAML (sources, processing, sinks, quality)
  3. Generate Kafka configurations with kafka_config_generator.py
  4. Generate Flink/Spark job scaffolding with stream_processor.py
  5. Deploy and monitor with streaming_quality_validator.py

Expected Output: Streaming pipeline processing 10K+ events/sec with P99 latency < 1s, exactly-once delivery, and real-time quality monitoring

World-class data engineering for production-grade data systems, scalable pipelines, and enterprise data platforms.

Overview

This skill provides comprehensive expertise in data engineering fundamentals through advanced production patterns. From designing medallion architectures to implementing real-time streaming pipelines, it covers the full spectrum of modern data engineering including ETL/ELT design, data quality frameworks, pipeline orchestration, and DataOps practices.

What This Skill Provides:

  • Production-ready pipeline templates (Airflow, Spark, dbt)
  • Comprehensive data quality validation framework
  • Performance optimization and cost analysis tools
  • Data architecture patterns (Lambda, Kappa, Medallion)
  • Complete DataOps CI/CD workflows

Best For:

  • Building scalable data pipelines for enterprise systems
  • Implementing data quality and governance frameworks
  • Optimizing ETL performance and cloud costs
  • Designing modern data architectures (lake, warehouse, lakehouse)
  • Production ML/AI data infrastructure

Quick Start

Pipeline Orchestration

# Generate Airflow DAG from configuration
python scripts/pipeline_orchestrator.py --config pipeline_config.yaml --output dags/

# Validate pipeline configuration
python scripts/pipeline_orchestrator.py --config pipeline_config.yaml --validate

# Use incremental load template
python scripts/pipeline_orchestrator.py --template incremental --output dags/

Data Quality Validation

# Validate CSV file with quality checks
python scripts/data_quality_validator.py --input data/sales.csv --output report.html

# Validate database table with custom rules
python scripts/data_quality_validator.py \
    --connection postgresql://user:pass@host/db \
    --table sales_transactions \
    --rules rules/sales_validation.yaml \
    --threshold 0.95

Performance Optimization

# Analyze pipeline performance and get recommendations
python scripts/etl_performance_optimizer.py \
    --airflow-db postgresql://host/airflow \
    --dag-id sales_etl_pipeline \
    --days 30 \
    --optimize

# Analyze Spark job performance
python scripts/etl_performance_optimizer.py \
    --spark-history-server http://spark-history:18080 \
    --app-id app-20250115-001

Real-Time Streaming

# Validate streaming pipeline configuration
python scripts/stream_processor.py --config streaming_config.yaml --validate

# Generate Kafka topic and client configurations
python scripts/kafka_config_generator.py \
    --topic user-events \
    --partitions 12 \
    --replication 3 \
    --output kafka/topics/

# Generate exactly-once producer configuration
python scripts/kafka_config_generator.py \
    --producer \
    --profile exactly-once \
    --output kafka/producer.properties

# Generate Flink job scaffolding
python scripts/stream_processor.py \
    --config streaming_config.yaml \
    --mode flink \
    --generate \
    --output flink-jobs/

# Monitor streaming quality
python scripts/streaming_quality_validator.py \
    --lag --consumer-group events-processor --threshold 10000 \
    --freshness --topic processed-events --max-latency-ms 5000 \
    --output streaming-health-report.html

Core Workflows

1. Building Production Data Pipelines

Steps:

  1. Design Architecture: Choose pattern (Lambda, Kappa, Medallion) based on requirements
  2. Configure Pipeline: Create YAML configuration with sources, transformations, targets
  3. Generate DAG: python scripts/pipeline_orchestrator.py --config config.yaml
  4. Add Quality Checks: Define validation rules for data quality
  5. Deploy & Monitor: Deploy to Airflow, configure alerts, track metrics

Pipeline Patterns: See frameworks.md for Lambda Architecture, Kappa Architecture, Medallion Architecture (Bronze/Silver/Gold), and Microservices Data patterns.

Templates: See templates.md for complete Airflow DAG templates, Spark job templates, dbt models, and Docker configurations.

2. Data Quality Management

Steps:

  1. Define Rules: Create validation rules covering completeness, accuracy, consistency
  2. Run Validation: python scripts/data_quality_validator.py --rules rules.yaml
  3. Review Results: Analyze quality scores and failed checks
  4. Integrate CI/CD: Add validation to pipeline deployment process
  5. Monitor Trends: Track quality scores over time

Quality Framework: See frameworks.md for complete Data Quality Framework covering all dimensions (completeness, accuracy, consistency, timeliness, validity).

Validation Templates: See templates.md for validation configuration examples and Python API usage.

3. Data Modeling & Transformation

Steps:

  1. Choose Modeling Approach: Dimensional (Kimball), Data Vault 2.0, or One Big Table
  2. Design Schema: Define fact tables, dimensions, and relationships
  3. Implement with dbt: Create staging, intermediate, and mart models
  4. Handle SCD: Implement slowly changing dimension logic (Type 1/2/3)
  5. Test & Deploy: Run dbt tests, generate documentation, deploy

Modeling Patterns: See frameworks.md for Dimensional Modeling (Kimball), Data Vault 2.0, One Big Table (OBT), and SCD implementations.

dbt Templates: See templates.md for complete dbt model templates including staging, intermediate, fact tables, and SCD Type 2 logic.

4. Performance Optimization

Steps:

  1. Profile Pipeline: Run performance analyzer on recent pipeline executions
  2. Identify Bottlenecks: Review execution time breakdown and slow tasks
  3. Apply Optimizations: Implement recommendations (partitioning, indexing, batching)
  4. Tune Spark Jobs: Optimize memory, parallelism, and shuffle settings
  5. Measure Impact: Compare before/after metrics, track cost savings

Optimization Strategies: See frameworks.md for performance best practices including partitioning strategies, query optimization, and Spark tuning.

Analysis Tools: See tools.md for complete documentation on etl_performance_optimizer.py with query analysis and Spark tuning.

5. Building Real-Time Streaming Pipelines

Steps:

  1. Architecture Selection: Choose Kappa (streaming-only) or Lambda (batch + streaming) architecture
  2. Configure Pipeline: Create YAML config with sources, processing engine, sinks, quality thresholds
  3. Generate Kafka Configs: python scripts/kafka_config_generator.py --topic events --partitions 12
  4. Generate Job Scaffolding: python scripts/stream_processor.py --mode flink --generate
  5. Deploy Infrastructure: Use Docker Compose for local dev, Kubernetes for production
  6. Monitor Quality: python scripts/streaming_quality_validator.py --lag --freshness --throughput

Streaming Patterns: See frameworks.md for stateful processing, stream joins, windowing, exactly-once semantics, and CDC patterns.

Templates: See templates.md for Flink DataStream jobs, Kafka Streams applications, PyFlink templates, and Docker Compose configurations.

Python Tools

pipeline_orchestrator.py

Automated Airflow DAG generation with intelligent dependency resolution and monitoring.

Key Features:

  • Generate production-ready DAGs from YAML configuration
  • Automatic task dependency resolution
  • Built-in retry logic and error handling
  • Multi-source support (PostgreSQL, S3, BigQuery, Snowflake)
  • Integrated quality checks and alerting

Usage:

# Basic DAG generation
python scripts/pipeline_orchestrator.py --config pipeline_config.yaml --output dags/

# With validation
python scripts/

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars1.8k
CategoryAutomation
Updated2mo ago
Forks368

Languages

PDDL

Trust signals

100/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

No cautions