e2e-medallion-architecture
Plans and builds end-to-end Fabric data platforms on the Bronze, Silver, and Gold medallion pattern: ingestion to analytics, PySpark and Delta layers, pipeline orchestration, per-layer Spark tuning, Materialized Lake View versus notebook choices for Silver and Gold, and Direct Lake handoff.
Install / Use
npx skills add microsoft/skills-for-fabric --skill e2e-medallion-architectureInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AutomationSupported Platforms
Our assessment of e2e-medallion-architecture
e2e-medallion-architecture scores 91/100 on our quality scale, 1033rd of 2,881 Automation skills we index (top 36%).
Its SKILL.md is 26 KB long, well organised into 31 sections with 3 code examples: a thorough specification that gives an agent plenty to work with.
With 1,181 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 15 days ago, so e2e-medallion-architecture is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
e2e-medallion-architecture compared with similar skills
All 4 of these similar skills score higher than e2e-medallion-architecture; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| e2e-medallion-architecture (this skill)by microsoft | 91 | 1.2k | 15d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 89.0k | 17d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.3k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 85.3k | 2d ago | MCP Server |
| crawl4aiby unclecode | 100 | 84.7k | 8d ago | MCP Server |
Frequently asked questions
- How do I install e2e-medallion-architecture?
- Run
npx skills add microsoft/skills-for-fabric --skill e2e-medallion-architecture. The install tabs above show the steps for each supported agent. - Which AI agents does e2e-medallion-architecture work with?
- It is written for Zed, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is e2e-medallion-architecture safe to use?
- It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is e2e-medallion-architecture still maintained?
- The repository was last updated 15 days ago, so e2e-medallion-architecture is actively maintained.
Skill content
View source on GitHubname: e2e-medallion-architecture description: "Plans and builds end-to-end Fabric data platforms on the Bronze, Silver, and Gold medallion pattern: ingestion to analytics, PySpark and Delta layers, pipeline orchestration, per-layer Spark tuning, Materialized Lake View versus notebook choices for Silver and Gold, and Direct Lake handoff. Use for layered architecture design, planning and recommendations. For questions over an existing Power BI report use fabriciq."
Telemetry — MANDATORY. Every
api.fabric.microsoft.comcall must carryx-ms-fabric-skill: e2e-medallion-architecture(az rest:--headers "x-ms-fabric-skill=e2e-medallion-architecture"), including every LRO poll,fabric_lroand retry. Snippets omit it — add it anyway.
CRITICAL NOTES
- To find the workspace details (including its ID) from workspace name: list all workspaces and, then, use JMESPath filtering
- To find the item details (including its ID) from workspace ID, item type, and item name: list all items of that type in that workspace and, then, use JMESPath filtering
End-to-End Medallion Architecture
Prerequisite Knowledge
Read these companion documents — they contain the foundational context this skill depends on:
- COMMON-CORE.md — Fabric REST API patterns, authentication, token audiences, item discovery
- COMMON-CLI.md —
az rest,az login, token acquisition, Fabric REST via CLI - SPARK-AUTHORING-CORE.md — Notebook deployment, lakehouse creation, job execution
- notebook-api-operations.md — Required for notebook creation —
.ipynbstructure requirements, cell format,getDefinition/updateDefinitionworkflow
For Spark-specific optimization details, see data-engineering-patterns.md.
Architecture Overview
Medallion Architecture is a data lakehouse pattern with three progressive layers:
| Layer | Purpose | Optimization Profile | Use Case | |-------|---------|---------------------|----------| | Bronze (Raw) | Land raw data exactly as received | Write-optimized, append-only, partitioned by ingestion date | Audit trail, reprocessing, lineage | | Silver (Cleaned) | Deduplicated, validated, conformed data | Balanced read/write, partitioned by business date | Feature engineering, operational reporting | | Gold (Aggregated) | Pre-calculated metrics for analytics | Read-optimized (ZORDER, compaction), partitioned by month/year | Power BI reports, dashboards, ad-hoc analytics via SQL endpoint |
- Bronze: Schema-on-read — flexible schema, Delta time travel supports audit and rollback
- Silver: Schema enforcement — reject non-conforming writes; handle schema evolution with
mergeSchemawhen sources change - Gold: Strict schema governance — curated, business-approved datasets only
Must/Prefer/Avoid
MUST DO
- Choose lakehouse architecture based on schema-enabled availability (see infrastructure-orchestration.md):
- Preferred: Schema-enabled lakehouse → create ONE workspace + ONE lakehouse with
bronze,silver,goldschemas - Legacy: Non-schema-enabled → create separate workspaces per layer (Bronze, Silver, Gold) for governance and access control
- Preferred: Schema-enabled lakehouse → create ONE workspace + ONE lakehouse with
- Use Livy API for schema and table creation — to create schemas and tables in a schema-enabled lakehouse, submit Spark SQL statements via Livy sessions (
POST /livyApi/versions/2023-12-01/sessions→POST .../statements). This is the only programmatic REST path for DDL operations (CREATE SCHEMA, CREATE TABLE) in Fabric lakehouses. - Add metadata columns in Bronze: ingestion timestamp, source file, batch ID
- Apply data quality rules in the Bronze-to-Silver transformation (deduplication, null handling, range validation)
- Use Delta Lake format for all medallion layer tables
- Use partition-aware overwrite in Silver/Gold writes to avoid reprocessing unchanged data
- Include validation steps after each layer (row counts, schema checks, anomaly detection)
- Follow the
.ipynbvalidation + Fabric nuances in notebook-api-operations.md when creating notebooks via REST API — every code cell must include"outputs": []and"execution_count": null - Complete the full end-to-end flow — do not stop after creating notebooks; always bind lakehouses, execute notebooks sequentially (Bronze → Silver → Gold), verify results, and connect Power BI to the Gold layer unless the user explicitly requests a partial setup
- In every MLV-versus-notebook recommendation, state that MLVs require a schema-enabled lakehouse. Hand off MLV definition and incremental-refresh review to
spark-cliauthoring mode; hand off scheduling, refresh, monitoring, and failure diagnosis tospark-climlv mode. - For recurring MLV refresh, give the exact interactive path Lakehouse → Materialized lake views → Manage → Schedules. For automation, use
POST /workspaces/{workspaceId}/lakehouses/{lakehouseId}/jobs/refreshMaterializedLakeViews/schedules; never invent/mlvRefreshSchedulesor route recurring refresh through notebook scheduling.
PREFER
- Incremental processing (watermark pattern) over full refresh
- Separate notebooks per layer for independent testing and debugging
- ZORDER on frequently filtered columns in Gold tables
- Running OPTIMIZE after writes in Silver and Gold layers
- Environment-specific Spark configs (write-heavy for Bronze, balanced for Silver, read-heavy for Gold)
- OneLake shortcuts to expose Gold data to consumer workspaces without duplication
- Clear layer ownership: engineers own Bronze/Silver, analysts own Gold
- Fabric Variable Libraries to centralize paths and configuration across layers
- Multi-workspace deployment patterns for medium/high governance requirements (Bronze/Silver/Gold in separate workspaces)
- Use Materialized Lake Views (MLVs) for Silver/Gold tables when the transformation is expressible in Spark SQL and benefits from declarative refresh semantics. See spark-cli — Materialized Lake View patterns and MLV incremental refresh patterns.
- Treat "materialized view", "spark materialized view", and "MLV" as the same Fabric feature.
AVOID
- Storing all layers in a single lakehouse WITHOUT schemas — non-schema lakehouses require notebook init cells or Environment configuration to enable OneLake Spark Catalog for RLS/CLS and MLVs. Use separate lakehouses for isolation if schemas aren't available.
- Creating 3 separate lakehouses when schema-enabled lakehouse is available — use schemas within one lakehouse instead (cleaner, no boilerplate init cells, more efficient for MLV cross-schema transformations)
- Skipping the Silver layer and going directly from Bronze to Gold
- Hardcoded workspace IDs, lakehouse IDs, or FQDNs — discover via REST API
- SELECT * without LIMIT on Bronze tables (they grow unboundedly)
- Running VACUUM without checking downstream dependencies
- Chaining OneLake shortcuts between medallion layers (Bronze→Silver→Gold) — each layer must be physically materialized for lineage and governance
- Copying complete implementation code into skills — guide the LLM to generate instead
- Reading from external HTTP/HTTPS URLs directly in Spark — Fabric Spark cannot access arbitrary external URLs; land data in lakehouse
Files/first (viacurl, OneLake API, or Fabric pipeline Copy activity), then read from the lakehouse path - Creating notebooks via REST API without validating
.ipynbstructure — missingexecution_count: nulloroutputs: []on code cells causes silent failures or "Job instance failed without detail error"
Workspace Setup Guidance
When setting up a medallion workspace, choose your architecture pattern first (see infrastructure-orchestration.md for detailed guidance):
Option A: Schema-Enabled Lakehouse (Preferred)
- Create single workspace:
{project}-{env} - Create one lakehouse with schemas:
{project}_lakehouse - Create schemas within the lakehouse:
bronzeschema for raw ingestionsilverschema for cleaned/validated datagoldschema for aggregated analytics
- Choose transformation approach:
- Option 4a: Use notebooks for each layer (PySpark or Spark SQL transformations)
- Option 4b: Use Materialized Lake Views (Spark SQL) for declarative transformations with incremental refresh (when query is IR-eligible) — see materialized-lake-view-patterns.md and mlv-incremental-refresh-patterns.md
- Note: PySpark MLVs exist but use full refresh only (no incremental) — use when you need UDFs/complex Python logic
- MLV benefit: OneLake Spark Catalog is automatically enabled for schema-enabled lakehouses — MLVs work out-of-box with no notebook init cells or Environment configuration required
- RBAC (optional): Use row-level security and column masking within schemas for fine-grained access control (also requires OneLake Spark Catalog)
Option B: Separate Lakehouses (Legacy)
- Create three workspaces:
{project}-bronze-{env}{project}-silver-{env}{project}-gold-{env}
- Create one lakehouse per workspace:
- Bronze workspace →
{project}_bronzelakehouse - Silver workspace →
{project}_silverlakehouse - Gold workspace →
{project}_goldlakehouse
- Bronze workspace →
- Assign RBAC per layer workspace:
- Bronze: ingestion/engineering write permissions
- Silver: engineering/data quality permissions
- Gold: analytics/BI consumer access with stricter curation controls
- Enable OneLake Spark Catalog for non-schema lakehouses (required for RLS/CLS and catalog-backed access patterns):
- Primary: Set
spark.sql.fabric.catalog.enable-schemaless-lakehouses=truein an Environment and attach it to notebooks. - Alternative: Omit default lakehouse binding from notebooks. Use four-part fully-qualified references (
workspace.lakehouse.schema.table). OneLake Spark Catalog auto-enables when no default lakehouse is set. - Alternative (internal/unsupported): Add this as the first cell in every notebook:
%%pyspark !echo "spark.sql.fabric.catalog.enable-schemaless-lakehouses=true" >> /home/trusted-service-user/.trident-context- ⚠️ Note: This workaround uses an internal runtime configuration path that may change in future Fabric releases. Prefer schema-enabled lakehouses for stable, documented OneLake Spark Catalog support.
- With this configuration, non-schema lakehouses support:
- ✅ Row-level security (RLS) and column-level security (CLS)
- Note: MLVs require schema-enabled lakehouses (Option A). For non-schema lakehouses, use notebooks with Delta tables.
- Primary: Set
Common Steps (Both Options)
After completing Option A or Option B above, perform these steps:
- Create notebooks for each layer (one per transformation stage) — follow
.ipynbvalidation + Fabric nuances - Bind each notebook to its lakehouse — set
metadata.dependencies.lakehousewith the correct lakehouse ID (see [notebook-api-operations.md § Default Lakehouse Binding](../spark-cli/references/authoring/resources/notebook-api-operations.md#default-lakehouse-binding
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
89.0kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.3kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Scrapling
85.3k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
crawl4ai
84.7kOpen-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
