SkillAgentSearch skills...

matlab-classify-tabular-data

Use this skill to classify tabular data end-to-end in MATLAB — load a dataset, prepare and clean it, select promising classifiers, train them, and compare accuracies with cross-validation, holdout, or hyperparameter optimization plus statistical tests.

Install / Use

npx skills add matlab/matlab-agentic-toolkit --skill matlab-classify-tabular-data

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

93/100

Supported Platforms

Universal

Our assessment of matlab-classify-tabular-data

matlab-classify-tabular-data scores 93/100 on our quality scale, 803rd of 4,646 Development & Engineering skills we index (top 18%).

Its SKILL.md is 36 KB long, well organised into 34 sections with 14 code examples: a thorough specification that gives an agent plenty to work with.

With 1,098 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
30/30
Structure
20/20
Description
15/15
Adoption
13/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 18 days ago, so matlab-classify-tabular-data is actively maintained.
  • No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
  • Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

matlab-classify-tabular-data compared with similar skills

All 4 of these similar skills score higher than matlab-classify-tabular-data; compare them before choosing.

SkillScoreStarsUpdatedFormat
matlab-classify-tabular-data (this skill)by matlab931.1k18d agoSKILL.md
ai-job-searchby MadsLorentzen10044.9ktodayCLAUDE.md
claude-howtoby luongnv8910041.7k3d agoCLAUDE.md
LocalAIby mudler10049.4ktodayMCP Server
algorithmic-artby anthropics100177.9k11d agoSKILL.md

Frequently asked questions

How do I install matlab-classify-tabular-data?
Run npx skills add matlab/matlab-agentic-toolkit --skill matlab-classify-tabular-data. The install tabs above show the steps for each supported agent.
Which AI agents does matlab-classify-tabular-data work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is matlab-classify-tabular-data safe to use?
It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is matlab-classify-tabular-data still maintained?
The repository was last updated 18 days ago, so matlab-classify-tabular-data is actively maintained.

name: matlab-classify-tabular-data description: > Use this skill to classify tabular data end-to-end in MATLAB — load a dataset, prepare and clean it, select promising classifiers, train them, and compare accuracies with cross-validation, holdout, or hyperparameter optimization plus statistical tests. TRIGGER when: user asks to classify tabular data, pick classifiers for a dataset, compare classifier accuracy, run cross-validation or a holdout evaluation, or find the best model with statistical uncertainty. DO NOT TRIGGER when: user has non-tabular inputs (images, sequences, time series), wants a regression model, is training a specific neural network architecture (use matlab-train-network), or wants cost-sensitive learning or an arbitrary class-prior vector (this skill only supports the built-in uniform-prior toggle for imbalanced data). license: https://www.mathworks.com/content/dam/mathworks/license/pmrl/license.md metadata: author: MathWorks version: "1.1"

Compare Classification Models with Statistical Uncertainty

Compare classifiers on the user's dataset and identify the top tier of models that are statistically equivalent in accuracy.

This skill bundles the workflow in references/ (step-by-step markdown instructions read on demand) and scripts/ (MATLAB code). Under scripts/, helpers/ holds the reusable computation and plotting functions the agent calls directly, and model_catalog/ holds the declarative branch-table .m files invoked internally by build_model_definitions via run(...). Do not invoke files in references/ as separate skills — they are only loaded via the Read tool when the corresponding step runs. See references/README.md for the layout.

When to Use

  • User wants to classify tabular data (matrix or table of predictors + categorical response).
  • User asks to compare multiple classifiers, pick the best model, or evaluate classifier accuracy.
  • User needs cross-validation, a holdout evaluation, or hyperparameter optimization for classifiers.
  • User needs statistical tests (McNemar, 5×2 cv, Friedman) to know which accuracy differences are significant.

When NOT to Use

  • Response is continuous — use a regression skill instead.
  • Predictors are images, sequences, or time series — this skill assumes a numeric matrix or a table of scalar predictors.
  • User wants to design or train a specific neural network architecture — use matlab-train-network. This skill does include fitcnet as one candidate on regular-branch data, but does not tune network layers or hyperparameters.
  • User only wants to score a pretrained model on new data — this skill trains and compares; scoring an existing model does not need it.
  • User wants cost-sensitive learning or a custom class-prior vector. Refuse plainly and stop — do not attempt a workaround. This skill's only supported prior control is the built-in uniform-prior toggle for imbalanced data ('Prior','uniform', offered via the UseUniformPrior flag in Step 5). Arbitrary 'Prior',[...] vectors and 'Cost',C matrices are not supported: neither the model-definition helper nor the CV/holdout scoring helpers thread these through, the branch tables and ECOC-expansion logic assume the built-in prior/cost defaults, and pairwise statistical tests (testckfold, testcholdout) score misclassification rate rather than expected cost. Attempting to bypass the helpers to inject a custom prior or cost is a hard refusal, not a judgment call — tell the user: "This skill does not support custom class priors or cost matrices. If you need cost-sensitive learning or a specific prior vector, use fitc* directly with the 'Prior' or 'Cost' name-value pair; this skill's statistical-comparison workflow will not give correct results in that setting."

Running MATLAB

Run all MATLAB code via the MATLAB MCP server (mcp__matlab__evaluate_matlab_code, or mcp__matlab__run_matlab_file for scripts). Set project_path to this skill's scripts/helpers/ directory so the workflow helpers resolve on the current working folder without any addpath calls.

Communication style while running this skill

Talk to the user about their data and results, not about the skill's plumbing. Everything under references/ and scripts/ (both helpers/ and model_catalog/) is internal — a user watching the run should never have to ask what a filename means.

Concretely, while executing this skill:

  • Do not name internal files or helpers in progress messages. references/select-classifiers.md, build_model_definitions, compute_pairwise_pvalues_cv, classifier_branches, resolve_recipe, modelDefs, cvFitFcn, etc., are all internal. If you must reference them (e.g., surfacing a bug the user can act on), name them once and explain what they are.
  • Do not narrate branch dispatch or filter decisions by name. "Dispatching to the wide branch", "applying the imbalanced overlay", "isSparse is false so we skip the sparse branch" — all internal. The user only needs to hear the outcome: "This dataset is wide (200 features, 40 samples), so I'm using linear models."
  • Do not read reference files out loud. When a step says STOP and read references/foo.md, that is a directive to you, not a status update to broadcast. Read it silently and continue.
  • Do announce what the user chose to run, and roughly how long it will take, before a long training loop or HPO run. One sentence.
  • Do surface user-facing questions verbatim where the step specifies them (evaluation strategy, interpretability, uniform prior, HPO selection). Those are the user's interface to the skill.
  • Do surface a real problem if one appears — a failed check gate, a branch that couldn't match, a fit call that errored. Name it plainly; then, and only then, is it fine to reference the internal file where the fix belongs.

Illustrative contrast:

Don't: "Reading references/dataprep.md... computing flags via compute_data_flags... isImbalanced=true, so reading references/select-classifiers-imbalanced.md and applying IMBALANCED_BOOSTING_MODELS overlay to build_model_definitions. modelDefs now has 8 entries including RUSBoost-OVO and RUSBoost-OVA."

Do: "Class ratio is 9:1, so I'll use imbalance-aware boosting models. Before I train, I need to ask you about the class prior — [uniform-prior question verbatim]."

Rule of thumb: if a sentence would only make sense to someone who has read this skill's source files, don't say it.

Before Writing Code

Do not reuse any variables from previous analysis runs. Always execute the full prescription and set all variables from scratch for each analyzed dataset. Every step must define its own variables — never assume anything remains in the workspace from a prior run.

Define skillPath up front. Several helpers take skillPath as an argument (the skill's root directory, two levels above scripts/helpers/). When you invoke MATLAB via evaluate_matlab_code with project_path set to this skill's scripts/helpers/ folder, MATLAB's working directory is scripts/helpers/, so skillPath = fileparts(fileparts(pwd)); gives the correct value. Set it at the top of the first code block that needs it (Step 2 or Step 3) and rely on the same value thereafter:

skillPath = fileparts(fileparts(pwd));  % skill root; used by build_model_definitions, THRESHOLDS load, imbalanced overlay, export_workflow_script

For all rules about how to write the MATLAB code itself (use built-ins, do not inspect template objects, training-time rules, figure rules), see scripts/README.md.

Step 1: Load data

  • Load the dataset
  • Determine whether a separate test set is provided (e.g., separate training and test tables/matrices).
    • If yes: set hasHoldout = true and assign XTrain, YTrain, XTest, YTest.
    • If no: assign X, Y from the loaded data. Do NOT split or ask about splitting yet.

Assumption on input shape. If any predictor is categorical, X must be a table with the categorical column(s) stored as MATLAB categorical (or string/cellstr, which compute_data_flags also treats as categorical). Matrix inputs are assumed fully numeric. If the user hands you predictors as a set of separate variables of mixed types (some numeric, some categorical/string), assemble them into a table with table(...) before proceeding — do not concatenate them into a numeric matrix, which would silently coerce categoricals to numeric codes. CategoricalPredictors NV pairs on fitc* are not used by this skill; categorical detection is entirely through the table column dtype.

Step 2: Analyze and clean data

STOP. Use the Read tool on references/dataprep.md (relative to this skill's directory). Do NOT write any MATLAB code until you have read that file. Follow its instructions exactly as written.

Run the dataprep instructions on XTrain/YTrain if hasHoldout = true (dataset came with a separate test set), or on X/Y otherwise. On return, the workspace must contain a flags struct produced by compute_data_flags with fields: N, D, nClasses, classSize, smallestClassSize, classRatio, isBinary, isImbalanced, isWide, isHighD, isBig, hasManyMissing, isSparse, hasCategorical.

The workspace must also contain a preproc struct array recording every mutation to X or Y in the order applied. Its op catalog — the only four ops the exported Step 14 workflow script can replay via apply_preproc — is:

| .op | When emitted | .payload fields | |---|---|---| | omitColumns | User chose to drop high-missingness or high-cardinality features | columnNames (table X) or columnIndices (matrix X) | | dropZeroVariance | Automatic drop of globally constant features | columnNames (table X) or columnIndices (matrix X) | | removeClasses | User chose to remove minority classes | classes (cell array of class labels) | | mergeClasses | User chose to merge minority classes into one label | map.from (source labels), map.to (target label) |

Any op name outside this catalog will cause apply_preproc to error at replay time — inventing new op names is a bug, not an extension point.

Step 3: Choose evaluation strategy

If hasHoldout = true (dataset came with a separate test set): skip this step.

Otherwise, load the thresholds and use flags.smallestClassSize (computed in Step 2) to present a recommendation:

run(fullfile(skillPath, 'scripts', 'model_catalog', 'classifier_thresholds.m'));  % populates THRESHOLDS

If flags.smallestClassSize > THRESHOLDS.holdout_smallest_class, recommend a 70/30 holdout split. Interpolate the actual threshold into the prompt via sprintf — do not hardcode the number:

This dataset has [flags.N] observations and the smallest class has [flags.smallestClassSize] observations — larger than the [THRESHOLDS.holdout_smallest_class]-observation threshold for holdout evaluation. I recommend a 70/30 stratified train/test split. This is faster and evaluates on unseen data.

Alternatively, I can use 5-fold cross-validation. CV uses all data for both training and evaluation but is slower.

Which do you prefer? (holdout / cv)

Otherwise (flags.smallestClassSize <= THRESHOLDS.holdout_smallest_class), recommend cross-validation:

This dataset has [flags.N] observations and the smallest class has [flags.smallestClassSize] observations — at or below the [THRESHOLDS.holdout_smallest_class]-observation threshold for a reliable holdout split. I recommend 5-fold cross-validation.

Alternatively, I can use a 70/30 stratified holdout split, though the test set may be small.

Which do you prefer? (cv / holdout)

Wait for the user's response. If the user chooses holdout, create the split via make_holdout_split and set hasHoldout = true. If the user chooses cv, set hasHoldout = false.

All subsequent ste

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars1.1k
CategoryDevelopment
Updated18d ago
Forks134

Languages

MATLAB

Trust signals

88/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

1 medium