matlab-choose-big-data-solution
Guide users or agents to the correct MATLAB tool for processing large tabular data in file-based formats (CSV, Parquet, delimited text, spreadsheets, MDF) that may not fit in memory
Install / Use
npx skills add matlab/matlab-agentic-toolkit --skill matlab-choose-big-data-solutionInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Data & AnalyticsSupported Platforms
Tags
Our assessment of matlab-choose-big-data-solution
matlab-choose-big-data-solution scores 93/100 on our quality scale, 135th of 513 Data & Analytics skills we index (top 27%).
Its SKILL.md is 18 KB long, well organised into 14 sections with 9 code examples: a thorough specification that gives an agent plenty to work with.
With 1,098 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 18 days ago, so matlab-choose-big-data-solution is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
matlab-choose-big-data-solution compared with similar skills
All 4 of these similar skills score higher than matlab-choose-big-data-solution; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| matlab-choose-big-data-solution (this skill)by matlab | 93 | 1.1k | 18d ago | SKILL.md |
| algorithmic-artby anthropics | 100 | 177.9k | 11d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 11d ago | SKILL.md |
| designby nextlevelbuilder | 100 | 130.2k | 12d ago | SKILL.md |
| ui-ux-pro-maxby nextlevelbuilder | 100 | 130.2k | 12d ago | SKILL.md |
Frequently asked questions
- How do I install matlab-choose-big-data-solution?
- Run
npx skills add matlab/matlab-agentic-toolkit --skill matlab-choose-big-data-solution. The install tabs above show the steps for each supported agent. - Which AI agents does matlab-choose-big-data-solution work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is matlab-choose-big-data-solution safe to use?
- It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is matlab-choose-big-data-solution still maintained?
- The repository was last updated 18 days ago, so matlab-choose-big-data-solution is actively maintained.
Skill content
View source on GitHubname: matlab-choose-big-data-solution description: > Guide users or agents to the correct MATLAB tool for processing large tabular data in file-based formats (CSV, Parquet, delimited text, spreadsheets, MDF) that may not fit in memory. Use when a user or agent mentions large files, big data, out-of-memory errors, OOM, scaling up, tall arrays, datastores, or needs to process multiple tabular files. Covers the decision between datastore + tall, datastore + transform, and parallel execution. Also use when a user or agent has working in-memory code (readtable, parquetread) that runs out of memory and needs a migration path. Also covers building custom datastore classes for proprietary or non-standard formats — use when the task requires subclassing matlab.io.Datastore, implementing a custom reader, building an extensible datastore, or integrating a new file format with tall arrays or parallel computing. Do NOT use for MAT files (use matfile instead). license: https://www.mathworks.com/content/dam/mathworks/license/pmrl/license.md metadata: author: MathWorks version: "2.0"
Choose Big Data Solution
Help users and agents select the right MATLAB tool for large tabular data in file-based formats (CSV, Parquet, delimited text, spreadsheets, MDF). The skill encodes a decision flowchart — recommend one clear path, not a menu of options.
When to Use
- User or agent has large tabular data in file-based formats (CSV, Parquet, delimited text, spreadsheets, MDF)
- User or agent says data is "large", "big", or "huge" (even without a specific size)
- User or agent hits an out-of-memory error with
readtable,parquetread, or similar - User or agent has multiple tabular files to process as a single dataset
- User or agent has multiple tabular files to process independently (per-file)
- User or agent asks how to scale up an existing workflow on tabular file data
- User or agent asks about datastores, tall arrays, or mapreduce
- User or agent needs to build a custom datastore class for a proprietary or non-standard format
- User or agent has large XML, JSON, HTML, or Word document files —
readtablesupports these formats but the built-in datastores do not. Scaling these requires a custom datastore (see "Formats Without a Built-in Datastore" and "Custom Datastores" sections below)
When NOT to Use
- Single file that fits in memory (file-size-to-RAM ratio < 0.5) — see Pre-Flight Check and
matlab-import-export-dataskill - User or agent is working with MAT files —
matfileprovides partial I/O for large.matfiles, different workflow - User or agent is working with databases (SQL, ODBC) — use Database Toolbox. See the
matlab-use-databaseandmatlab-use-duckdbskills in thereporting-and-database-accesscategory - User or agent needs GPU acceleration — different domain
- Deep learning training workflows — datastore is correct for data loading, but use
minibatchqueuefor batching (not tall or transform)
Pre-Flight Check
ALWAYS run this check BEFORE recommending datastore or tall array patterns. Estimate the file-size-to-available-RAM ratio (accounting for in-memory expansion of the file format).
| Condition | Route | Rationale |
|-----------|-------|-----------|
| Single file, ratio < 0.5 | Native MATLAB I/O (readtable, parquetread) | Fits in memory; datastore/tall adds unnecessary complexity |
| Single file, ratio ≥ 0.5, OR user/agent reports OOM | Continue to Decision Flowchart below | Data may not fit in memory |
| Multiple files | Continue to Decision Flowchart below | Datastore patterns provide unified multi-file access |
| Size unknown and user/agent describes data as "large", "big", or "huge" | Continue to Decision Flowchart below | Assume large until proven otherwise |
If the data fits in memory (single file, ratio < 0.5): recommend native I/O and stop. Mention that datastore and tall array patterns exist if the data grows beyond memory in the future, but do not implement them now.
Decision Flowchart
Follow this flowchart strictly. Present ONE recommended path, not multiple alternatives. Only mention alternatives if the situation is ambiguous.
Is the data described as "large" or causing OOM?
│
├── YES
│ │
│ ├── Is the goal to process all data as ONE continuous dataset?
│ │ │
│ │ └── YES → Datastore + Tall Arrays
│ │ Choose datastore by format:
│ │ CSV/delimited text → tabularTextDatastore
│ │ Parquet → parquetDatastore
│ │ Excel (.xlsx/.xls) → spreadsheetDatastore
│ │ MDF (.mf4/.mdf) → mdfDatastore (requires Vehicle Network Toolbox)
│ │ Other formats → Custom Datastore
│ │
│ └── Is the goal to process each unit INDEPENDENTLY?
│ │
│ ├── One read = one FILE
│ │ CSV/delimited text → tabularTextDatastore
│ │ Excel (.xlsx/.xls) → spreadsheetDatastore
│ │ Other formats → Custom Datastore or fileDatastore
│ │
│ └── One read = one ROW GROUP
│ Parquet → parquetDatastore
│
└── NO / UNCLEAR
└── STOP. Ask: (1) what is the file format? (2) do you need to process
all data as one dataset, or each file/unit independently?
Do NOT show code until these are answered.
Optional (requires Parallel Computing Toolbox):
├── Speed up tall arrays locally → Open a parallel pool
├── Speed up readall on transforms → UseParallel
└── Speed up tall arrays on Hadoop/Spark → mapreducer (also requires MATLAB Parallel Server)
Critical rules:
- When the processing goal is unclear (continuous vs per-file), STOP and ask before writing any code. Do not show code examples, do not generate scripts, do not demonstrate both approaches. Ask which applies and wait for the answer. The correct response to ambiguity is a short clarifying question, not a menu of options with code for each.
- When data is described as "large", NEVER lead with
readtableorparquetread - For per-file or per-row-group processing, NEVER suggest tall arrays (tall merges all data into one dataset and has no concept of file or row group boundaries)
- For Parquet, the natural independent unit is the row group, not the file —
parquetDatastoredefaults toReadSize = "rowgroup". Do NOT forceReadSize = "file"on a Parquet datastore unless the workflow truly needs whole-file granularity - For parallelism, recommend parallel pool FIRST; mapreducer is only for Hadoop/Spark
- Do NOT present manual chunking (while/read loops) as the first option — tall arrays handle chunking automatically
- Do NOT present multiple options — follow the flowchart and recommend ONE clear path
Datastore + Tall Arrays (continuous dataset)
Use when: processing one large file OR multiple files as a single dataset. Use the datastore matching the format (see Decision Flowchart above).
% Single file
ds = tabularTextDatastore("largedata.csv");
tt = tall(ds);
% Multiple files
ds = tabularTextDatastore("data/*.csv");
tt = tall(ds);
% For example, compute statistics with the tall array and gather results.
% Tall handles chunking automatically. Specify DataVars to avoid errors
% on non-numeric columns.
result = groupsummary(tt, "GroupVar", {"mean", "std"}, "NumericVar");
result = gather(result);
Key points:
- Most common functions work on tall:
mean,std,min,max,groupsummary,sortrows,topkrows - Combine multiple gathers:
[a, b] = gather(tallA, tallB)
DuckDB alternative (R2026a+): When the goal is to filter, aggregate, deduplicate, or sample a large file down to a small in-memory result, a single DuckDB query may be faster than datastore + tall. DuckDB queries CSV/Parquet/JSON files directly without loading them into memory. Requires Database Toolbox.
conn = duckdb();
result = fetch(conn, "SELECT * FROM read_csv_auto('large.csv') WHERE Region = 'West'");
close(conn);
For complex DuckDB workflows, see the matlab-use-duckdb skill in the
reporting-and-database-access category.
Migrating from readtable (OOM on large files):
% Before:
T = readtable("large.csv");
stats = groupsummary(T, "Category", "mean");
% After:
ds = tabularTextDatastore("large.csv");
tt = tall(ds);
stats = gather(groupsummary(tt, "Category", "mean"));
tabularTextDatastore uses a stricter textscan-based parser. Key gotchas:
TextTypemust be set at creation time — read-only after constructionTrimNonNumericunsupported — read as%q, strip on the tall array%ferrors on non-numeric content — read as%q, convert withstr2doubleExtraColumnsRuledoes not exist — append"%*[^\r\n]"toTextscanFormatsdetectImportOptionsdoes not apply — configure via datastore properties directly
See references/readtable-to-datastore-migration.md for the full property mapping.
Migrating from parquetread: Replace with parquetDatastore + tall. Types
are preserved exactly — no format specifier issues. All parquetread options
map 1:1 to datastore properties. See references/parquetread-to-datastore-migration.md.
Datastore + Transform (per-file)
Use when: each file in a folder should be processed independently (per-file statistics, per-file transformations, file-level aggregation).
% Set up datastore to read one file at a time
ds = tabularTextDatastore("data/*.csv");
ds.ReadSize = "file";
% For example, compute statistics by transforming the datastore with a custom function.
tds = transform(ds, @computeFileStats);
results = readall(tds);
function out = computeFileStats(data)
numVars = vartype("numeric");
out = table( ...
min(data{:, numVars}, [], 1, "omitmissing"), ...
max(data{:, numVars}, [], 1, "omitmissing"), ...
mean(data{:, numVars}, 1, "omitmissing"), ...
VariableNames=["Min", "Max", "Mean"]);
end
For Excel files, use spreadsheetDatastore instead of tabularTextDatastore.
spreadsheetDatastore defaults to ReadSize = "file" so no override is needed.
For per-sheet processing (e.g., statistics per worksheet), set ReadSize = "sheet"
so each read returns exactly one sheet's data.
Key points:
- For
tabularTextDatastore, setReadSize = "file"so eachreadreturns exactly one file's data. The defaultReadSizeis a row count, which can split a single file across multiple reads but it will not go across file boundaries transformapplies a function to each read — the datastore output contains only the transformed results, not the original data. The transform function does not need to return the same number of rows as its input (e.g., computing the mean of each variable produces a single row per read)readallon the transformed datastore collects all per-file results- Use
"omitmissing"for missing-data flags (R2023a+). On R2022b and earlier, use"omitnan"instead.
Do NOT use tall arrays for per-file processing. Tall arrays treat all files as one continuous dataset — they have no concept of file boundaries.
Datastore + Transform (per-row-group, Parquet)
Use when: a Parquet dataset is partitioned by row group (e.g., one row group per item, sensor, region, or time bucket) and each row group must be processed independently. Row groups are the natural unit of independence in Parquet — data from different row groups should not be mixed when the partitioning encodes a meaningful grouping.
% parquetDatastore defaults ReadSize to "rowgroup" — one read = one row group.
% Do NOT set ReadSize = "file": that would mix row groups within the same file.
pds = parquetDatastore("data/");
% For example, compute statistics by transforming the datastore with a custom function.
tds = transform(pds, @computeRowGroupStats);
results = readall(tds);
function out = computeRowGroupStats(data)
out = table( ...
unique(data.item), ...
mean(data.val, "omitmissi
Truncated for display — read the full file on GitHub.
Related Skills
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
design
130.2kComprehensive design skill: brand identity, design tokens, UI styling, logo generation (55 styles, Gemini, Atlas Cloud, or MuAPI AI), corporate identity program (50 deliverables, CIP mockups), HTML presentations (Chart.js), banner design (22 styles, social/ads/web/print), icon design (15 styles, SVG…
ui-ux-pro-max
130.2kUI/UX design intelligence for web, mobile, and desktop. This skill should be used when designing, building, reviewing, or fixing interfaces, including pages, components, design systems, accessibility, interaction, responsive layout, typography, color, charts, and stack-specific UI implementation.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
