matlab-use-duckdb
Use DuckDB from MATLAB via Database Toolbox (R2026a+) as a non-math operations engine on large tabular files (CSV/Parquet/JSON) and as a zero-config embedded database
Install / Use
npx skills add matlab/matlab-agentic-toolkit --skill matlab-use-duckdbInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Data & AnalyticsSupported Platforms
Our assessment of matlab-use-duckdb
matlab-use-duckdb scores 93/100 on our quality scale, 138th of 510 Data & Analytics skills we index (top 28%).
Its SKILL.md is 15 KB long, well organised into 27 sections with 9 code examples: a thorough specification that gives an agent plenty to work with.
With 1,098 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 18 days ago, so matlab-use-duckdb is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
matlab-use-duckdb compared with similar skills
All 4 of these similar skills score higher than matlab-use-duckdb; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| matlab-use-duckdb (this skill)by matlab | 93 | 1.1k | 18d ago | SKILL.md |
| claude-memby thedotmack | 100 | 95.5k | today | CLAUDE.md |
| algorithmic-artby anthropics | 100 | 177.9k | 11d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 11d ago | SKILL.md |
| designby nextlevelbuilder | 100 | 130.2k | 12d ago | SKILL.md |
Frequently asked questions
- How do I install matlab-use-duckdb?
- Run
npx skills add matlab/matlab-agentic-toolkit --skill matlab-use-duckdb. The install tabs above show the steps for each supported agent. - Which AI agents does matlab-use-duckdb work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is matlab-use-duckdb safe to use?
- It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is matlab-use-duckdb still maintained?
- The repository was last updated 18 days ago, so matlab-use-duckdb is actively maintained.
Skill content
View source on GitHubname: matlab-use-duckdb description: "Use DuckDB from MATLAB via Database Toolbox (R2026a+) as a non-math operations engine on large tabular files (CSV/Parquet/JSON) and as a zero-config embedded database. Use when connecting to DuckDB, querying CSV, Parquet, and JSON files directly with SQL, reducing or profiling large data before MATLAB analysis, creating portable development databases, or installing DuckDB extensions. Triggers on: DuckDB, duckdb(), large CSV/Parquet/JSON, file too large for readtable, filter/aggregate at source, deduplicate, reduce before analysis, profile large file, persistent file import, analytical engine, SQL on CSV, SQL on Parquet, SQL on JSON, query CSV with SQL, query Parquet with SQL, run SQL on files, SQL queries on files, query files directly, SQL without database, in-process SQL." license: https://www.mathworks.com/content/dam/mathworks/license/pmrl/license.md metadata: author: MathWorks version: "1.2"
MATLAB Database Toolbox Interface to DuckDB
Use when working with DuckDB databases from MATLAB using Database Toolbox. DuckDB is an embedded analytical database engine that ships with Database Toolbox starting in R2026a. It enables SQL-based analytics on files, out-of-memory data preprocessing, and portable development databases — all without external database server configuration.
When to Use This Skill
- Connecting to a DuckDB database (in-memory or file-based)
- Creating a new DuckDB database file for development workflows
- Querying CSV, Parquet, or JSON files directly with SQL
- Reducing or profiling large data before analysis
- Using DuckDB as a non-math operations engine (filter, aggregate, deduplicate, string cleanup, type cast, sort, sample)
- Installing and using DuckDB extensions
- Persisting file data into a
.duckdbfile for repeated queries - User mentions keywords: DuckDB, duckdb(), analytical engine, embedded database, large CSV, large Parquet/JSON, file too large for readtable, filter/aggregate at source, deduplicate, reduce before analysis, profile large file, persistent file import, SQL on CSV, SQL on Parquet, SQL on JSON, query CSV with SQL, query Parquet with SQL, run SQL on files, SQL queries on files, query files directly, SQL without database, in-process SQL
When NOT to Use
- Single file that fits in memory — use native MATLAB I/O (
readtable/readmatrix) - Multi-file workflows — use
datastore/tallfirst; DuckDB glob is the fallback, not the default - Data exceeds memory and no SQL reduction will make it fit — use
datastore/tallfor streaming - Math-heavy or numerically sensitive operations (normalize, windowed stats, outlier detection) — hand off to MATLAB after reduction
- Connecting to MySQL, PostgreSQL, SQLite, or other external databases — use their native interfaces or JDBC/ODBC
- Object-relational mapping — use ORM (
ormread/ormwritewithMappableclasses) - MongoDB, Cassandra, or Neo4j — use their dedicated Database Toolbox interfaces
On-Load Protocol
When this skill is loaded into a session where code already exists:
-
Audit the data pipeline — Identify how data enters the workflow:
- Is data read via MATLAB I/O (
readtable,readmatrix,readcell,readtimetable,parquetread,xlsread,csvread)? - Is data then written to DuckDB with
sqlwritebefore querying? - If yes: this is the load-then-query anti-pattern. DuckDB can likely read the source file directly via
read_csv,read_parquet, orread_xlsx(excel extension).
- Is data read via MATLAB I/O (
-
Evaluate each data source against the decision framework:
- Can DuckDB read this file type directly? (CSV, Parquet, JSON, Excel via extension)
- Can filtering/aggregation be pushed into the SQL read?
- Is the MATLAB I/O step a performance bottleneck?
-
Recommend architectural changes — Do not limit review to API correctness. Propose replacing
readtable/xlsread+sqlwrite+ query chains withfetch(conn, "SELECT ... FROM read_csv/read_parquet/read_xlsx(...)").
The highest-value patterns in this skill are architectural: file-analytics pushdown eliminates entire pipeline stages and can yield 10x+ speedups.
Pre-Flight Check
Before committing to the DuckDB path, estimate the file-size-to-available-RAM ratio (accounting for in-memory expansion of the file format) and route accordingly:
| Condition | Route | Rationale |
|-----------|-------|-----------|
| Single file, ratio < 0.5 | Native MATLAB I/O | Fits in memory; DuckDB adds unnecessary complexity |
| Multi-file OR per-file processing | datastore/tall (fallback: DuckDB glob) | Streaming is the primary multi-file pattern |
| Single file, ratio ≥ 0.5, SQL-expressible reduction | DuckDB: profile → operate → close | DuckDB is warranted |
| Single file, ratio ≥ 0.5, no SQL reduction applies | datastore/tall | No reduction = no DuckDB advantage |
Operations Menu
DO in DuckDB SQL
| Operation | Example SQL |
|-----------|-------------|
| Filter rows | WHERE status = 'active' |
| Aggregate | GROUP BY region, SUM(revenue), COUNT(*), AVG(price) |
| Deduplicate | SELECT DISTINCT ... or ROW_NUMBER() OVER (PARTITION BY ...) |
| String cleanup | TRIM(), LOWER(), REPLACE(), REGEXP_REPLACE() |
| Type casting | TRY_CAST(col AS DATE), CAST(col AS DOUBLE) |
| Sort | ORDER BY timestamp |
| Sample | USING SAMPLE 10000 or TABLESAMPLE RESERVOIR(10%) |
| Profile | COUNT(*), MIN/MAX, APPROX_COUNT_DISTINCT |
| Limit | LIMIT 50000 |
DO NOT in DuckDB SQL
| Operation | Why Not | |-----------|---------| | Normalize / z-score | Requires population context | | Windowed statistics (moving avg, cumsum) | Fragile semantics, NULL handling differs | | Outlier detection | Statistical judgment call | | Interpolation / resampling | Domain-specific logic | | Signal processing | Not SQL's domain | | Machine learning features | Model-dependent |
SQL operations do not always match MATLAB semantics (NULL handling, sort order, deduplication). See reference/cards/reduce-large-data.md for known mismatches.
Reduction Workflow
If the pre-flight check routes to DuckDB, follow the profile → operate → close pattern in reference/cards/reduce-large-data.md.
DuckDB selects, reshapes, cleans, and reduces data. Once data fits in memory and the connection is closed, DuckDB's role is complete.
What Is DuckDB and Why Does Database Toolbox Ship It?
DuckDB is an embedded, serverless analytical database engine. Unlike MySQL or PostgreSQL, it requires no server, no configuration, and runs in-process within MATLAB.
Why it ships with Database Toolbox (R2026a+):
- Zero-config database —
conn = duckdb()gives you a full SQL engine instantly. - Analytical engine for files — Query CSV, Parquet, and JSON files directly with SQL without loading them into memory.
- Out-of-memory preprocessing — Filter, aggregate, join, and sort datasets larger than memory, then bring only results into MATLAB.
- Portable development databases —
.duckdbor.dbfiles work on any machine with Database Toolbox. No database setup needed. - AI agent advantage — An agent's SQL knowledge directly translates to powerful analytical queries.
DuckDB does NOT replace MATLAB's file I/O (readtable, etc.) or datastore/tall. It is a performant alternative when data exceeds memory and the task reduces to a SQL-expressible operation. When no reduction applies, datastore/tall remains the correct path.
Critical Rules
Pre-Flight Gate
- ALWAYS run the pre-flight size-to-RAM check BEFORE writing any DuckDB code. If ratio < 0.5, STOP and use native MATLAB I/O instead. Do not proceed with DuckDB patterns.
Connection
- ALWAYS use
duckdb()to connect — notdatabase(), not JDBC, not ODBC. - ALWAYS verify with
isopen(conn)and close withclose(conn).
API Surface
- All standard functions work:
sqlread,fetch,execute,sqlwrite,sqlfind,sqlinnerjoin,sqlouterjoin,commit,rollback. - DuckDB does NOT support
databasePreparedStatement. Useexecuteorsqlwriteinstead. - Use
ExcludeDuplicatesviadatabaseImportOptionswhen reading from database tables (withsqlread). For direct file queries (read_csv/read_parquetviafetch), useSELECT DISTINCTin SQL.
Named Tables vs. File Queries
- ALWAYS use
sqlreadwithRowFilterfor simple row filtering on named database tables — notfetchwith WHERE. Create arowfilterobject and pass it via theRowFiltername-value argument. - Use
fetchwith SQL for aggregation, joins, or complex queries on named tables. - ALWAYS use
fetch(notsqlread) for file queries — they require SQL syntax likeSELECT * FROM read_csv('file.csv'). - ALWAYS use single quotes for file paths inside SQL:
read_csv('data.csv').
Connection Modes
| Goal | Connection | Why |
|------|-----------|-----|
| Analytical queries on files | duckdb() | No persistence needed; query files directly |
| Temporary workspace | duckdb() | Fast, discarded on close |
| Portable development database | duckdb("mydata.duckdb") | Creates a .duckdb or .db file; works on any machine |
| Open existing database | duckdb("existing.db") | Read/write access to pre-existing .db or .duckdb file |
| Read-only shared database | duckdb("shared.duckdb", ReadOnly=true) | Prevents accidental writes |
Common Patterns
Pattern 1: Analytical Engine on Files
conn = duckdb();
result = fetch(conn, "SELECT region, SUM(revenue) as total " + ...
"FROM read_parquet('sales.parquet') " + ...
"GROUP BY region ORDER BY total DESC");
close(conn);
Pattern 2: Out-of-Memory Preprocessing
conn = duckdb();
summary = fetch(conn, "SELECT date, AVG(value) as avg_val " + ...
"FROM read_csv('huge_dataset.csv') " + ...
"WHERE status = 'valid' " + ...
"GROUP BY date ORDER BY date");
close(conn);
% summary fits in memory — continue with MATLAB analysis
Pattern 3: Read Named Table with RowFilter
conn = duckdb("inventory.duckdb", ReadOnly=true);
rf = rowfilter("quantity");
data = sqlread(conn, "products", RowFilter=rf.quantity < 10);
close(conn);
Pattern 4: Development Database
conn = duckdb("dev.duckdb");
sqlwrite(conn, "experiments", experimentData);
rf = rowfilter("score");
results = sqlread(conn, "experiments", RowFilter=rf.score > 0.8);
close(conn);
Pattern 5: Multi-File Query with Glob
Use when datastore/tall is unsuitable (e.g., SQL aggregation across files). For sequential per-file processing, prefer datastore.
conn = duckdb();
data = fetch(conn, "SELECT * FROM read_parquet('data/year=2024/*.parquet') " + ...
"WHERE category = 'A'");
close(conn);
Pattern 6: Extensions
conn = duckdb();
execute(conn, "INSTALL httpfs");
execute(conn, "LOAD httpfs");
data = fetch(conn, "SELECT * FROM read_parquet('https://example.com/data.parquet') LIMIT 1000");
close(conn);
For Excel files, use the excel extension with read_xlsx (NOT st_read from spatial):
conn = duckdb();
execute(conn, "INSTALL excel");
execute(conn, "LOAD excel");
data = fetch(conn, "SELECT * FROM read_xlsx('report.xlsx')");
close(conn);
Pattern 7: Persistent Import from File
conn = duckdb("analytics.duckdb");
execute(conn, "CREATE TABLE events AS SELECT * FROM read_parquet('raw_events.parquet')");
% Future sessions: query by table name (no file re-read)
result = sqlread(conn, "events");
close(conn);
For detailed examples, see:
- File analytics and out-of-memory preprocessing:
reference/cards/file-analytics.md - Development database workflows:
reference/cards/development-database.md - DuckDB extensions:
reference/cards/extensions.md
Common Mistakes
% WRONG — using database() or JDBC to connect to DuckDB
conn = database("", "", "", "org.duckdb.DuckDBDri
Truncated for display — read the full file on GitHub.
Related Skills
claude-mem
95.5kPersistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
design
130.2kComprehensive design skill: brand identity, design tokens, UI styling, logo generation (55 styles, Gemini, Atlas Cloud, or MuAPI AI), corporate identity program (50 deliverables, CIP mockups), HTML presentations (Chart.js), banner design (22 styles, social/ads/web/print), icon design (15 styles, SVG…
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
