matlab-classify-tabular-data
Use this skill to classify tabular data end-to-end in MATLAB — load a dataset, prepare and clean it, select promising classifiers, train them, and compare accuracies with cross-validation, holdout, or hyperparameter optimization plus statistical tests.
Install / Use
npx skills add matlab/matlab-agentic-toolkit --skill matlab-classify-tabular-dataInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Development & EngineeringSupported Platforms
Our assessment of matlab-classify-tabular-data
matlab-classify-tabular-data scores 93/100 on our quality scale, 803rd of 4,646 Development & Engineering skills we index (top 18%).
Its SKILL.md is 36 KB long, well organised into 34 sections with 14 code examples: a thorough specification that gives an agent plenty to work with.
With 1,098 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 18 days ago, so matlab-classify-tabular-data is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
matlab-classify-tabular-data compared with similar skills
All 4 of these similar skills score higher than matlab-classify-tabular-data; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| matlab-classify-tabular-data (this skill)by matlab | 93 | 1.1k | 18d ago | SKILL.md |
| ai-job-searchby MadsLorentzen | 100 | 44.9k | today | CLAUDE.md |
| claude-howtoby luongnv89 | 100 | 41.7k | 3d ago | CLAUDE.md |
| LocalAIby mudler | 100 | 49.4k | today | MCP Server |
| algorithmic-artby anthropics | 100 | 177.9k | 11d ago | SKILL.md |
Frequently asked questions
- How do I install matlab-classify-tabular-data?
- Run
npx skills add matlab/matlab-agentic-toolkit --skill matlab-classify-tabular-data. The install tabs above show the steps for each supported agent. - Which AI agents does matlab-classify-tabular-data work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is matlab-classify-tabular-data safe to use?
- It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is matlab-classify-tabular-data still maintained?
- The repository was last updated 18 days ago, so matlab-classify-tabular-data is actively maintained.
Skill content
View source on GitHubname: matlab-classify-tabular-data description: > Use this skill to classify tabular data end-to-end in MATLAB — load a dataset, prepare and clean it, select promising classifiers, train them, and compare accuracies with cross-validation, holdout, or hyperparameter optimization plus statistical tests. TRIGGER when: user asks to classify tabular data, pick classifiers for a dataset, compare classifier accuracy, run cross-validation or a holdout evaluation, or find the best model with statistical uncertainty. DO NOT TRIGGER when: user has non-tabular inputs (images, sequences, time series), wants a regression model, is training a specific neural network architecture (use matlab-train-network), or wants cost-sensitive learning or an arbitrary class-prior vector (this skill only supports the built-in uniform-prior toggle for imbalanced data). license: https://www.mathworks.com/content/dam/mathworks/license/pmrl/license.md metadata: author: MathWorks version: "1.1"
Compare Classification Models with Statistical Uncertainty
Compare classifiers on the user's dataset and identify the top tier of models that are statistically equivalent in accuracy.
This skill bundles the workflow in references/ (step-by-step markdown instructions read on demand) and scripts/ (MATLAB code). Under scripts/, helpers/ holds the reusable computation and plotting functions the agent calls directly, and model_catalog/ holds the declarative branch-table .m files invoked internally by build_model_definitions via run(...). Do not invoke files in references/ as separate skills — they are only loaded via the Read tool when the corresponding step runs. See references/README.md for the layout.
When to Use
- User wants to classify tabular data (matrix or table of predictors + categorical response).
- User asks to compare multiple classifiers, pick the best model, or evaluate classifier accuracy.
- User needs cross-validation, a holdout evaluation, or hyperparameter optimization for classifiers.
- User needs statistical tests (McNemar, 5×2 cv, Friedman) to know which accuracy differences are significant.
When NOT to Use
- Response is continuous — use a regression skill instead.
- Predictors are images, sequences, or time series — this skill assumes a numeric matrix or a table of scalar predictors.
- User wants to design or train a specific neural network architecture — use
matlab-train-network. This skill does includefitcnetas one candidate on regular-branch data, but does not tune network layers or hyperparameters. - User only wants to score a pretrained model on new data — this skill trains and compares; scoring an existing model does not need it.
- User wants cost-sensitive learning or a custom class-prior vector. Refuse plainly and stop — do not attempt a workaround. This skill's only supported prior control is the built-in uniform-prior toggle for imbalanced data (
'Prior','uniform', offered via theUseUniformPriorflag in Step 5). Arbitrary'Prior',[...]vectors and'Cost',Cmatrices are not supported: neither the model-definition helper nor the CV/holdout scoring helpers thread these through, the branch tables and ECOC-expansion logic assume the built-in prior/cost defaults, and pairwise statistical tests (testckfold,testcholdout) score misclassification rate rather than expected cost. Attempting to bypass the helpers to inject a custom prior or cost is a hard refusal, not a judgment call — tell the user: "This skill does not support custom class priors or cost matrices. If you need cost-sensitive learning or a specific prior vector, usefitc*directly with the'Prior'or'Cost'name-value pair; this skill's statistical-comparison workflow will not give correct results in that setting."
Running MATLAB
Run all MATLAB code via the MATLAB MCP server (mcp__matlab__evaluate_matlab_code, or mcp__matlab__run_matlab_file for scripts). Set project_path to this skill's scripts/helpers/ directory so the workflow helpers resolve on the current working folder without any addpath calls.
Communication style while running this skill
Talk to the user about their data and results, not about the skill's plumbing. Everything under references/ and scripts/ (both helpers/ and model_catalog/) is internal — a user watching the run should never have to ask what a filename means.
Concretely, while executing this skill:
- Do not name internal files or helpers in progress messages.
references/select-classifiers.md,build_model_definitions,compute_pairwise_pvalues_cv,classifier_branches,resolve_recipe,modelDefs,cvFitFcn, etc., are all internal. If you must reference them (e.g., surfacing a bug the user can act on), name them once and explain what they are. - Do not narrate branch dispatch or filter decisions by name. "Dispatching to the wide branch", "applying the imbalanced overlay", "
isSparseis false so we skip the sparse branch" — all internal. The user only needs to hear the outcome: "This dataset is wide (200 features, 40 samples), so I'm using linear models." - Do not read reference files out loud. When a step says STOP and read
references/foo.md, that is a directive to you, not a status update to broadcast. Read it silently and continue. - Do announce what the user chose to run, and roughly how long it will take, before a long training loop or HPO run. One sentence.
- Do surface user-facing questions verbatim where the step specifies them (evaluation strategy, interpretability, uniform prior, HPO selection). Those are the user's interface to the skill.
- Do surface a real problem if one appears — a failed check gate, a branch that couldn't match, a fit call that errored. Name it plainly; then, and only then, is it fine to reference the internal file where the fix belongs.
Illustrative contrast:
Don't: "Reading
references/dataprep.md... computingflagsviacompute_data_flags...isImbalanced=true, so readingreferences/select-classifiers-imbalanced.mdand applyingIMBALANCED_BOOSTING_MODELSoverlay tobuild_model_definitions.modelDefsnow has 8 entries including RUSBoost-OVO and RUSBoost-OVA."Do: "Class ratio is 9:1, so I'll use imbalance-aware boosting models. Before I train, I need to ask you about the class prior — [uniform-prior question verbatim]."
Rule of thumb: if a sentence would only make sense to someone who has read this skill's source files, don't say it.
Before Writing Code
Do not reuse any variables from previous analysis runs. Always execute the full prescription and set all variables from scratch for each analyzed dataset. Every step must define its own variables — never assume anything remains in the workspace from a prior run.
Define skillPath up front. Several helpers take skillPath as an argument (the skill's root directory, two levels above scripts/helpers/). When you invoke MATLAB via evaluate_matlab_code with project_path set to this skill's scripts/helpers/ folder, MATLAB's working directory is scripts/helpers/, so skillPath = fileparts(fileparts(pwd)); gives the correct value. Set it at the top of the first code block that needs it (Step 2 or Step 3) and rely on the same value thereafter:
skillPath = fileparts(fileparts(pwd)); % skill root; used by build_model_definitions, THRESHOLDS load, imbalanced overlay, export_workflow_script
For all rules about how to write the MATLAB code itself (use built-ins, do not inspect template objects, training-time rules, figure rules), see scripts/README.md.
Step 1: Load data
- Load the dataset
- Determine whether a separate test set is provided (e.g., separate training and test tables/matrices).
- If yes: set
hasHoldout = trueand assignXTrain,YTrain,XTest,YTest. - If no: assign
X,Yfrom the loaded data. Do NOT split or ask about splitting yet.
- If yes: set
Assumption on input shape. If any predictor is categorical, X must be a table with the categorical column(s) stored as MATLAB categorical (or string/cellstr, which compute_data_flags also treats as categorical). Matrix inputs are assumed fully numeric. If the user hands you predictors as a set of separate variables of mixed types (some numeric, some categorical/string), assemble them into a table with table(...) before proceeding — do not concatenate them into a numeric matrix, which would silently coerce categoricals to numeric codes. CategoricalPredictors NV pairs on fitc* are not used by this skill; categorical detection is entirely through the table column dtype.
Step 2: Analyze and clean data
STOP. Use the Read tool on references/dataprep.md (relative to this skill's directory). Do NOT write any MATLAB code until you have read that file. Follow its instructions exactly as written.
Run the dataprep instructions on XTrain/YTrain if hasHoldout = true (dataset came with a separate test set), or on X/Y otherwise. On return, the workspace must contain a flags struct produced by compute_data_flags with fields: N, D, nClasses, classSize, smallestClassSize, classRatio, isBinary, isImbalanced, isWide, isHighD, isBig, hasManyMissing, isSparse, hasCategorical.
The workspace must also contain a preproc struct array recording every mutation to X or Y in the order applied. Its op catalog — the only four ops the exported Step 14 workflow script can replay via apply_preproc — is:
| .op | When emitted | .payload fields |
|---|---|---|
| omitColumns | User chose to drop high-missingness or high-cardinality features | columnNames (table X) or columnIndices (matrix X) |
| dropZeroVariance | Automatic drop of globally constant features | columnNames (table X) or columnIndices (matrix X) |
| removeClasses | User chose to remove minority classes | classes (cell array of class labels) |
| mergeClasses | User chose to merge minority classes into one label | map.from (source labels), map.to (target label) |
Any op name outside this catalog will cause apply_preproc to error at replay time — inventing new op names is a bug, not an extension point.
Step 3: Choose evaluation strategy
If hasHoldout = true (dataset came with a separate test set): skip this step.
Otherwise, load the thresholds and use flags.smallestClassSize (computed in Step 2) to present a recommendation:
run(fullfile(skillPath, 'scripts', 'model_catalog', 'classifier_thresholds.m')); % populates THRESHOLDS
If flags.smallestClassSize > THRESHOLDS.holdout_smallest_class, recommend a 70/30 holdout split. Interpolate the actual threshold into the prompt via sprintf — do not hardcode the number:
This dataset has [flags.N] observations and the smallest class has [flags.smallestClassSize] observations — larger than the [THRESHOLDS.holdout_smallest_class]-observation threshold for holdout evaluation. I recommend a 70/30 stratified train/test split. This is faster and evaluates on unseen data.
Alternatively, I can use 5-fold cross-validation. CV uses all data for both training and evaluation but is slower.
Which do you prefer? (holdout / cv)
Otherwise (flags.smallestClassSize <= THRESHOLDS.holdout_smallest_class), recommend cross-validation:
This dataset has [flags.N] observations and the smallest class has [flags.smallestClassSize] observations — at or below the [THRESHOLDS.holdout_smallest_class]-observation threshold for a reliable holdout split. I recommend 5-fold cross-validation.
Alternatively, I can use a 70/30 stratified holdout split, though the test set may be small.
Which do you prefer? (cv / holdout)
Wait for the user's response. If the user chooses holdout, create the split via make_holdout_split and set hasHoldout = true. If the user chooses cv, set hasHoldout = false.
All subsequent ste
Truncated for display — read the full file on GitHub.
Related Skills
ai-job-search
44.9kThe job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.
claude-howto
41.7kA visual, example-driven guide to Claude Code — from basic concepts to advanced agents, with copy-paste templates that bring immediate value.
LocalAI
49.4kLocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
