matlab-connect-databricks
Connect MATLAB to Databricks via Spark (Databricks Connect) or JDBC (Database Toolbox)
Install / Use
npx skills add matlab/matlab-agentic-toolkit --skill matlab-connect-databricksInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Data & AnalyticsSupported Platforms
Our assessment of matlab-connect-databricks
matlab-connect-databricks scores 90/100 on our quality scale, 324th of 591 Data & Analytics skills we index.
Its SKILL.md is 14 KB long, well organised into 23 sections with 14 code examples: a thorough specification that gives an agent plenty to work with.
With 1,098 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 21 days ago, so matlab-connect-databricks is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
matlab-connect-databricks compared with similar skills
All 4 of these similar skills score higher than matlab-connect-databricks; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| matlab-connect-databricks (this skill)by matlab | 90 | 1.1k | 21d ago | SKILL.md |
| claude-memby thedotmack | 100 | 97.1k | today | CLAUDE.md |
| algorithmic-artby anthropics | 100 | 177.9k | 14d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 14d ago | SKILL.md |
| designby nextlevelbuilder | 100 | 133.6k | 3d ago | SKILL.md |
Frequently asked questions
- How do I install matlab-connect-databricks?
- Run
npx skills add matlab/matlab-agentic-toolkit --skill matlab-connect-databricks. The install tabs above show the steps for each supported agent. - Which AI agents does matlab-connect-databricks work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is matlab-connect-databricks safe to use?
- It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is matlab-connect-databricks still maintained?
- The repository was last updated 21 days ago, so matlab-connect-databricks is actively maintained.
Skill content
View source on GitHubname: matlab-connect-databricks description: > Connect MATLAB to Databricks via Spark (Databricks Connect) or JDBC (Database Toolbox). Use when setting up the MATLAB Interface for Databricks, configuring authentication (OauthU2M, OauthM2M, PAT), creating Spark sessions with getDatabricksSession(), reading Unity Catalog tables with server-side filtering, creating JDBC connections with databricks.JDBCConnection or StandaloneJDBCConnection, connecting to SQL Warehouses, selecting JDBC drivers (Simba/OSS), or writing data back to Databricks. Triggers on: Databricks Connect, Spark from MATLAB, getDatabricksSession, .databrickscfg, databricks.JDBCConnection, StandaloneJDBCConnection, SQLWarehouse, Databricks JDBC, Databricks cluster, large table server-side filtering. license: https://www.mathworks.com/content/dam/mathworks/license/pmrl/license.md metadata: author: MathWorks version: "1.0"
Connect MATLAB to Databricks
Connect MATLAB to Databricks via Spark (Databricks Connect) or JDBC (Database Toolbox). This skill covers first-time setup, authentication, path selection, session/connection creation, and data operations through both paths.
When to Use
- First-time setup of the MATLAB Interface for Databricks package
- Configuring
.databrickscfgand authentication (OauthU2M, OauthM2M, PAT) - Choosing between Spark and JDBC for a Databricks workflow
- Reading large tables via Spark with server-side filtering (
getDatabricksSession) - Creating JDBC connections to clusters or SQL Warehouses (
databricks.JDBCConnection,SQLWarehouse.connect()) - Standalone JDBC connectivity without the full package (
StandaloneJDBCConnection) - Writing data back to Databricks (via Spark DataFrames or JDBC
sqlwrite) - Selecting and configuring JDBC drivers (Simba vs. OSS)
When NOT to Use
- Generic Database Toolbox operations after connection is established — use
matlab-use-database - ODBC connections (
databricks.ODBCConnection) - Databricks REST APIs (Clusters, Jobs, DBFS, Unity Catalog admin)
- MLflow from MATLAB
- Statement Execution REST API
- Deploying compiled MATLAB code to Databricks clusters (Job workflow)
- DuckDB — use
matlab-use-duckdb - Databricks notebooks, Databricks CLI, or
databricks-sdkPython workflows - PySpark without MATLAB context (pure Python Spark usage)
Decision Framework: Spark vs. JDBC
| Scenario | Path | Why |
|----------|------|-----|
| Large table (millions of rows), need server-side filtering before pulling locally | Spark | DataFrame operations run on cluster; only filtered results transfer |
| SQL queries on small-to-medium datasets | JDBC | Direct SQL via Database Toolbox; simpler setup |
| Need DataFrame transformations (withColumn, select, filter chains) | Spark | Native DataFrame API; operations stay on cluster |
| Writing large data with performance optimization | JDBC | Simba driver's UseNativeQuery optimizes sqlwrite |
| No MATLAB Interface for Databricks package installed | JDBC | StandaloneJDBCConnection works with just Database Toolbox + driver jar |
| Need to use Database Explorer app | JDBC | saveSource() + copyToken() integration |
| Reading files from Unity Catalog Volumes (CSV, Parquet, JSON) | Spark | spark.read.format().load() with Volumes paths |
| Interactive exploration with sqlread/fetch | JDBC | Standard Database Toolbox workflow on j.Connection |
Default: Use Spark when the user mentions large data, server-side filtering, or DataFrames. Use JDBC when the user mentions SQL queries, Database Toolbox, or small/medium datasets. If unclear, ask the user about their data size and preferred workflow.
Workflow
- Obtain the package — Ask if the user has the MATLAB Interface for Databricks. If not, direct them to https://www.mathworks.com/solutions/partners/databricks.html (or use
StandaloneJDBCConnectionfor JDBC-only without the package) - Setup — Run
setupfrom the package'sSoftware/MATLABdirectory. It configures settings,.databrickscfg, and installs the Databricks Connect library via pip into a venv - Startup — Run
startupto add package paths (required once per MATLAB session) - Choose path — Use the Decision Framework above to select Spark or JDBC
- Connect — Create a session (
getDatabricksSession) or connection (databricks.JDBCConnection) - Verify — Spark:
table(spark.range(1)). JDBC: checkj.Connection.Messageis empty - Operate — Read, filter, write data using the appropriate path's API
- Close — JDBC:
close(j). Spark: sessions are managed automatically
Key Functions
Spark Path
| Function | Purpose |
|----------|---------|
| getDatabricksSession() | Creates a Spark session (classic compute) |
| getDatabricksSession(serverless=true) | Creates a serverless session (no cluster, 10-min timeout) |
| spark.read().table("catalog.schema.table") | Reads a Unity Catalog table as a DataFrame |
| spark.read.format(fmt).load(path) | Reads files from Volumes (csv, parquet, json) |
| DF.filter(expr) | Server-side row filtering |
| DF.select(cols) | Server-side column selection |
| DF.limit(n) | Server-side row limiting |
| DF.withColumn(name, col) | Adds/transforms a column (requires Column objects) |
| table(DF) | Converts DataFrame to MATLAB table (pulls data locally) |
| matlab.sparkutils.table2dataset(T, spark) | Converts MATLAB table back to Spark DataFrame |
| DF.write.mode(m).format(f).saveAsTable(name) | Writes DataFrame to Unity Catalog table |
| matlab.pyspark.sql.functions.col(name) | Creates a Column reference |
| matlab.pyspark.sql.functions.lit(value) | Creates a literal Column constant |
JDBC Path
| Function | Purpose |
|----------|---------|
| databricks.JDBCConnection() | Creates a JDBC connection (full package) |
| StandaloneJDBCConnection() | Creates a JDBC connection (no package dependencies) |
| databricks.SQLWarehouse.connect() | Connects to a SQL Warehouse by ID |
| j.Connection | The database.jdbc.connection object for Database Toolbox functions |
| j.testConnection() | Verifies connection is working |
| j.saveSource() | Saves connection as a Database Toolbox data source |
| close(j) | Closes connection and releases resources |
Patterns
First-Time Setup
Run setup from the package's Software/MATLAB directory. It is interactive — follow prompts for host URL, auth method, cluster ID, and Databricks Connect library installation.
The Databricks Connect library is downloaded via pip into Software/MATLAB/Connect/<version>/venv/. Requires Python 3.10-3.12. If the download fails repeatedly, ask the user to check with their IT team — do not retry with modified arguments.
cd('/path/to/databricks-package/Software/MATLAB')
setup
startup
After setup, MATLAB's pyenv must point to the venv Python. If getDatabricksSession fails with "databricks.connect package is not installed":
terminate(pyenv);
pyenv(Version="/path/to/databricks-package/Software/MATLAB/Connect/17.3/venv/bin/python");
Warning: If Python is already loaded InProcess, terminate(pyenv) fails. A full MATLAB restart is required. Do not attempt to switch pyenv mid-session after Python has been used.
For authentication configuration details (.databrickscfg format, profiles, environment variables, token caching), see references/authentication.md — consult when configuring auth methods or troubleshooting credential issues.
Spark: Read and Filter a Table
Data stays on the cluster until explicitly collected. Filter server-side first, then collect.
spark = getDatabricksSession(authMethod="PAT");
DF = spark.read().table("catalog.schema.sensor_readings");
filtered = DF.filter("temperature > 100 AND event_date > '2024-01-01'");
T = table(filtered);
Spark: Serverless Session
No cluster needed. Starts instantly with 10-minute inactivity timeout. Requires Python 3.11-3.12 and Databricks Connect >= 15.4.
spark = getDatabricksSession(serverless=true);
DF = spark.read().table("catalog.schema.events");
T = table(DF.limit(50));
Spark: Add Computed Columns
withColumn requires Column objects — not raw scalars. Use col() for references and lit() for constants.
import matlab.pyspark.sql.functions.col
import matlab.pyspark.sql.functions.lit
DF = spark.read().table("catalog.schema.measurements");
DF2 = DF.withColumn("temp_fahrenheit", col("temp_celsius") * lit(9/5) + lit(32));
Spark: Write Back to Databricks
Always specify .mode() — without it, writes fail if the target exists.
DF_new = matlab.sparkutils.table2dataset(T, spark);
DF_new.write.mode("overwrite").format("delta").saveAsTable("catalog.schema.output_table");
Spark: Read Files from Volumes
DF = spark.read.format("csv").option("header", "true").load("/Volumes/catalog/schema/volume/data.csv");
T = table(DF.limit(1000));
JDBC: Cluster Connection
j = databricks.JDBCConnection(catalog="mycatalog", schema="myschema");
data = sqlread(j.Connection, "mytable");
close(j);
JDBC: SQL Warehouse Connection
warehouse.connect() returns a database.jdbc.connection directly — pass it to sqlread/fetch without .Connection.
warehouse = databricks.SQLWarehouse;
warehouse.id = "abc123def456";
conn = warehouse.connect();
data = fetch(conn, "SELECT * FROM mycatalog.myschema.mytable LIMIT 10");
close(conn);
JDBC: Standalone (No Package)
When the user does NOT have the MATLAB Interface for Databricks, use StandaloneJDBCConnection. Requires Database Toolbox and the Simba driver jar only. See references/standalone-jdbc.md for setup and JSON template.
j = StandaloneJDBCConnection(schema="myschema", catalog="mycatalog");
data = fetch(j.Connection, "SELECT * FROM mytable LIMIT 10");
close(j);
JDBC: On-Databricks (Browser MATLAB)
The JDBC driver's OAuth flow cannot open a browser when MATLAB runs on a Databricks cluster. Use package-managed auth instead.
if databricks.internal.isOnDatabricks()
j = databricks.JDBCConnection(authMethod="OauthU2M", useDriverAuth=false);
else
j = databricks.JDBCConnection();
end
data = fetch(j.Connection, "SELECT * FROM mycatalog.myschema.mytable LIMIT 10");
close(j);
JDBC: Write-Optimized Connection
Simba driver write performance improves with native query mode (enabled by default).
j = databricks.JDBCConnection(catalog="main", schema="telemetry");
sqlwrite(j.Connection, "measurements", data);
close(j);
Connection Cleanup
Use onCleanup to guarantee closure even when operations fail.
j = databricks.JDBCConnection(catalog="main", schema="analytics");
cleanup = onCleanup(@() close(j));
data = fetch(j.Connection, "SELECT * FROM large_table WHERE id > 1000");
For Spark stale sessions:
clear spark
spark = getDatabricksSession(forceNewSession=true);
Conventions
- Always filter DataFrames server-side before calling
table(DF)— pulling millions of unfiltered rows wastes bandwidth and memory - Use three-level names for Unity Catalog tables:
"catalog.schema.table" - Use
getDatabricksSession()for Spark — never construct sessions manually or mimic PySpark builder patterns - Use
databricks.JDBCConnectionorStandaloneJDBCConnectionfor JDBC — never manually construct JDBC URLs withdatabase() - Always pass
authMethodexplicitly — the default chain may trigger unexpected browser prompts - Never hardcode tokens or secrets in MATLAB code — use
.databrickscfgor environment variables - Always call
close(j)when done with JDBC connections - Use
forceNewSession=truefor stale Spark sessions, notclear all - For JDBC driver selection details, see [
references/driver-selection.md](references/driver-selection.
Truncated for display — read the full file on GitHub.
Related Skills
claude-mem
97.1kPersistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
design
133.6kComprehensive design skill: brand identity, design tokens, UI styling, logo generation (55 styles, Gemini, Atlas Cloud, or MuAPI AI), corporate identity program (50 deliverables, CIP mockups), HTML presentations (Chart.js), banner design (22 styles, social/ads/web/print), icon design (15 styles, SVG…
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
