matlab-engineer-tabular-features
Use when engineering or selecting the best features for single-response classification or regression in MATLAB, whatever the data's modality — for non-tabular data it routes extraction to a domain skill, then selects, assesses, and delivers on the resulting table.
Install / Use
npx skills add matlab/matlab-agentic-toolkit --skill matlab-engineer-tabular-featuresInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AutomationSupported Platforms
Tags
Our assessment of matlab-engineer-tabular-features
matlab-engineer-tabular-features scores 93/100 on our quality scale, 741st of 2,848 Automation skills we index (top 27%).
Its SKILL.md is 19 KB long, well organised into 15 sections with 8 code examples: a thorough specification that gives an agent plenty to work with.
With 1,098 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 18 days ago, so matlab-engineer-tabular-features is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
matlab-engineer-tabular-features compared with similar skills
All 4 of these similar skills score higher than matlab-engineer-tabular-features; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| matlab-engineer-tabular-features (this skill)by matlab | 93 | 1.1k | 18d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 89.8k | 18d ago | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 85.4k | today | MCP Server |
| rufloby ruvnet | 100 | 73.8k | today | MCP Server |
| algorithmic-artby anthropics | 100 | 177.9k | 11d ago | SKILL.md |
Frequently asked questions
- How do I install matlab-engineer-tabular-features?
- Run
npx skills add matlab/matlab-agentic-toolkit --skill matlab-engineer-tabular-features. The install tabs above show the steps for each supported agent. - Which AI agents does matlab-engineer-tabular-features work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is matlab-engineer-tabular-features safe to use?
- It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is matlab-engineer-tabular-features still maintained?
- The repository was last updated 18 days ago, so matlab-engineer-tabular-features is actively maintained.
Skill content
View source on GitHubname: matlab-engineer-tabular-features description: > Use when engineering or selecting the best features for single-response classification or regression in MATLAB, whatever the data's modality — for non-tabular data it routes extraction to a domain skill, then selects, assesses, and delivers on the resulting table. Not for multi-response problems, model training, or raw data acquisition. license: https://www.mathworks.com/content/dam/mathworks/license/pmrl/license.md metadata: author: MathWorks version: "1.0"
Engineer Tabular Features
A lean, functional pipeline: intake → feature pool → select → assess →
deliver → report. There is no shared context object — each phase is a direct
call to leaf utilities in scripts/, and the reference file for each phase
carries the detail. Your value is the structured, data-driven process and,
above all, the consensus selection at its center — not an ad-hoc answer.
This skill bundles the workflow in references/ (per-phase detail read on demand)
and scripts/ (the computation and plotting utilities). Do not invoke files in
references/ as separate skills — they are loaded only via the Read tool when
the phase that needs them runs. The per-phase files call the pipeline's leaf helpers
for you; if you ever need a helper's signature,
references/internal-helpers.md documents each one's
inputs, outputs, and an example call — so you never open a helper's source.
When to Use
- Engineering or selecting the best predictors for single-response supervised
classification or regression on a plain in-memory
table. - You want a structured, data-driven selection — a ranker panel, a consensus vote, and an elbow cut — rather than an ad-hoc hand-picked feature set.
- The data is non-tabular (signals, images, battery/machinery telemetry): this skill routes extraction to the matching domain skill, then engineers, selects, assesses, and delivers on the resulting table (see references/domain-routing.md).
When NOT to Use
- Multi-response problems — this skill is single-response only.
- Model training, tuning, or deployment — it prepares features and stops. Hand the delivered table to a model-training/classification workflow to fit and compare models.
- Raw data acquisition. And for the extraction step on non-tabular data, the actual feature computation belongs to the matching domain extraction skill — this skill orchestrates that handoff (see references/domain-routing.md), it does not re-implement it.
Requires the Statistics and Machine Learning Toolbox (SMLT) —
gencfeatures/genrfeatures build the pool and the ranker/assessment utilities
are SMLT-based. MATLAB Report Generator is optional (enables the PDF report;
markdown is always produced).
Running MATLAB
Run all MATLAB through the MATLAB MCP server (mcp__matlab__evaluate_matlab_code,
or mcp__matlab__run_matlab_file for scripts). Set project_path to this skill's
scripts/ directory so the utilities resolve on the current working folder
without any addpath calls. Every utility is a leaf function called directly —
there is no initialization step and no context object to construct. Validate any
code you author with mcp__matlab__check_matlab_code before running it.
Start each dataset from scratch — but use the live workspace within a run. The
MCP session is stateful, so a run's intermediates should live in the workspace: set
RawTbl, Splits, FullEng, SelectedNames, Baseline, etc. once and pass them
phase-to-phase. Do not round-trip them through save/load .mat files (noise,
risks stale reads) and do not addpath. Across different datasets/runs, carry
nothing — begin each analysis by setting every variable afresh.
Communication style while running this skill
Talk to the user about their data and results, not the skill's plumbing.
Everything under references/ and scripts/ is internal. Rule of thumb: if a
sentence would only make sense to someone who has read this skill's source files,
don't say it.
- Never name internal files, helpers, or phase/gating mechanics.
runConsensusSelection,GenInfo.BinaryReliant, "the redundancy dimension", etc. are internal — give the outcome ("these features duplicate each other, so I'm keeping the strongest"), not the mechanism. Name an internal only when it is a problem the user can act on. Read reference files silently. - Use plain words for each check. The three assessment reads: performance → whether the new features improve predictions (a held-out estimate, or a cross-validated mean ± std); fixed-pool stability → whether the same features get picked when rows are resampled; generation stability → whether the same features get built and picked when the whole pipeline re-runs on resampled rows. Say "the ranking step" not "the borda voter"; "reliably re-selected" not "consensus core".
- Don't narrate uncertainty or mid-flight course-corrections — settle how a function is called silently, then report only the outcome. Surface a difficulty only when the user must decide on it.
- Announce cost before long work, one sentence — pool size before selection, expected time before a K-fold, and before either stability gate (both re-run selection many times; the generation gate also re-builds the pool each time). And surface user-facing questions verbatim where a phase specifies one (output directory, wide-input, domain routing).
- Report what was dropped at every phase (screened predictors, excluded WoE columns, skipped rankers) — a silent shrink reads as data loss. But selection evaluates the pool, it doesn't necessarily shrink it — never call it a reduction.
Output directory — REQUIRED, HARD HALT
Deliverables are written to disk. Always confirm the output directory with the user before writing anything. Do not assume the working directory, do not create one silently.
The pipeline
Follow the phases in order. Each links to its reference; read the reference before executing the phase.
1. Intake — references/intake.md
Ask before running any code. Intake is a required conversation, not a
default-fill. Confirm every run parameter with the user before proceeding past
the screen — data source, response, dataset name, output directory (hard-halt),
domain description, separate-test-set, model family (+ lens if agnostic),
evaluation strategy, report opt-out — asked one at a time, in the order pinned in
intake.md §1 (never dump the whole list in one message).
Every item must be asked; offer a default where one exists, but confirm rather than
assume — when the opening request implies an answer, state what you inferred and
have the user confirm it. Two are
non-negotiable — do not proceed without an explicit answer:
- Output directory — the disk-write hard-halt (see above).
- Domain description — the sole input to domain routing. Ask what the data is and where it came from; route to a domain extractor if one fits, else the generic path. Never infer the domain from column names or fall through to generic generation on silence. Verbatim prompt in intake.md.
Then assemble the data into one plain table, briefly confirm what was
loaded (shape, response, problem type), and screen degenerate predictors. A
timetable/tall/gpuArray/datastore isn't a dead end — it's a signal to run a
tabularizing step first (often a domain skill, see
domain-routing.md) and then re-enter intake with
the resulting table; only halt if no tabular path exists.
[ScreenedTbl, ScreenInfo] = screenPredictors(RawTbl, Response); % or (X, y)
screenPredictors handles the polymorphic response (a name already in the table,
or a separately-supplied vector/table it concatenates) and drops constant and
near-empty predictors. Then profile, reserve the user's untouched slice, and split:
Profile = profileForSplit(ScreenedTbl, ScreenInfo.ResponseVar);
[WorkingIdx, UserHeldOutIdx, ReserveInfo] = reserveHoldoutForUser( ...
ScreenedTbl, Profile.ProblemType, ScreenInfo.ResponseVar, ReserveForUser = HasNoSeparateTest);
[Splits, SplitDecision] = splitStrategy(ScreenedTbl, Profile.ProblemType, ...
ScreenInfo.ResponseVar, Subset = WorkingIdx, EvaluationStrategy = EvaluationStrategy);
reserveHoldoutForUser sets aside an untouched slice for the user's own testing
when they have no separate test set (default 20%, user-settable via HoldoutFraction;
else nothing); the rest is the working data all phases run on. When a carve
happens, materialize the slice as the table variable's name + _test from the original
rows and narrate the split in plain words (fraction, row counts, and method —
stratified/random — from ReserveInfo); see intake.md. splitStrategy then sets
Splits.TrainIdx/.TestIdx over the working rows — a train/test split under
holdout, or all working rows with empty TestIdx under cross_validated.
Generation and selection run on TrainIdx only.
2. Feature pool — references/feature-pool.md
Produce the candidate pool. First check whether a domain skill fits the data (references/domain-routing.md); otherwise use the default SMLT path:
Opts.TargetModel = TargetModel; Opts.Standardization = "auto";
OptArgs = namedargs2cell(Opts);
TrainIdx = Splits.TrainIdx; TestIdx = Splits.TestIdx;
TrainTbl = ScreenedTbl(TrainIdx, :); % fit generation on train rows only
[~, Transformer, GenInfo] = generateFeatures(TrainTbl, Response, ProblemType, OptArgs{:});
Recipe = Transformer; % SMLT recipe (domain path: the captured struct)
FullEng = transformFeatures(Recipe, ScreenedTbl); % engineered pool over ALL rows
FullEng.(Response) = ScreenedTbl.(Response); % transformFeatures returns predictors only
Generate-only (external consensus does the cutting). Mind the wide-input guard
(generateFeatures:tooManyPredictors) — hold the wide-input conversation and
re-call with Opts.NumFeatures set. FullEng is the canonical pool: engineered
over all rows with the response re-attached, train-fit so the held-out rows
stay leakage-clean. On the domain path the captured table already spans all
rows — use it as FullEng directly. Whichever path runs, downstream reads only
the pool contract (Recipe, describeFeatures, transformFeatures) — never
the producer. Announce GenInfo.PoolSize.
Set OriginalData/OriginalPredVars here — the baseline's "original" reference is
path-dependent (raw columns when they exist, else the full pool). assess.md §1 pins the rule.
3. Select — references/select.md
The heart of the skill. One call runs the ranker panel, the consensus vote, and the elbow cut:
[SelectedNames, VoteTable, PanelInfo] = runConsensusSelection( ...
FullEng(TrainIdx, :), Response, ProblemType, ExcludeFeatures = GenInfo.BinaryReliant, ...
TargetModel = TargetModel);
Selection runs on the training rows only — slice FullEng(TrainIdx,:); the
held-out rows never enter ranking. TargetModel gates the ranker panel: a declared
family runs its own embedded probe (linear→lasso, tree_ensemble→oob,
kernel_distance→nca) plus the two model-agnostic rankers; agnostic (the default)
keeps the full five-ranker panel. Report PanelInfo.Reasoning. Build a
plotSelectionConsensus figure only when the report is on (GenerateReport) —
every figure is a report input, so a report opt-out skips all plot call
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
89.8kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
Scrapling
85.4k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
ruflo
73.8k🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
