SkillAgentSearch skills...

Rapid-MLX

The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.

Install / Use

npx skills add raullenchai/Rapid-MLX

Installs into whichever agent you are using.

About this skill
🤖

CLAUDE.md

Claude Code project instructions

Quality Score

95/100

Supported Platforms

Claude Code
Cursor
Aider
<img width="1600" height="800" alt="banner" src="https://github.com/user-attachments/assets/f3743bb7-7287-4b24-ac97-a7037974396f" /> <p align="center"> <strong>The fastest local AI engine for Apple Silicon.</strong> <br> <em>Drop-in OpenAI / Anthropic API · 2–4× faster than Ollama · Runs on any M-series Mac.</em> </p> <p align="center"> <a href="https://pypi.org/project/rapid-mlx/"><img src="https://img.shields.io/pypi/v/rapid-mlx?color=blue&label=PyPI" alt="PyPI"></a> <a href="https://formulae.brew.sh/formula/rapid-mlx"><img src="https://img.shields.io/badge/Homebrew-core-orange?logo=homebrew" alt="Homebrew core"></a> <a href="https://www.python.org/downloads/"><img src="https://img.shields.io/badge/python-3.10+-blue.svg" alt="Python 3.10+"></a> <a href="https://support.apple.com/en-us/HT211814"><img src="https://img.shields.io/badge/Apple_Silicon-M1%20|%20M2%20|%20M3%20|%20M4-black.svg?logo=apple" alt="Apple Silicon"></a> <a href="LICENSE"><img src="https://img.shields.io/badge/License-Apache_2.0-blue.svg" alt="License"></a> </p> <p align="center"> <a href="https://github.com/raullenchai/Rapid-MLX/actions/workflows/ci.yml"><img src="https://github.com/raullenchai/Rapid-MLX/actions/workflows/ci.yml/badge.svg" alt="CI"></a> <a href="https://github.com/raullenchai/Rapid-MLX/stargazers"><img src="https://img.shields.io/github/stars/raullenchai/Rapid-MLX?style=social" alt="GitHub stars"></a> <a href="https://github.com/raullenchai/Rapid-MLX/graphs/contributors"><img src="https://img.shields.io/github/contributors/raullenchai/Rapid-MLX?color=orange" alt="Contributors"></a> <a href="https://github.com/raullenchai/Rapid-MLX/commits/main"><img src="https://img.shields.io/github/last-commit/raullenchai/Rapid-MLX?color=orange" alt="Last commit"></a> <a href="https://deepwiki.com/raullenchai/Rapid-MLX"><img src="https://deepwiki.com/badge.svg" alt="Ask DeepWiki"></a> </p> <p align="center"> <sub> <a href="https://rapidmlx.com"><b>rapidmlx.com</b></a> · <a href="https://rapidmlx.com/docs/">Docs</a> · <a href="https://models.rapidmlx.com/">Model mirror</a> · <a href="https://rapidmlx.com/desktop">Desktop app</a> </sub> </p>

Quick Start (60 seconds)

1. Install — pick one path (run only one of these):

One-liner — detects your RAM, picks a starter model (recommended):

curl -fsSL https://rapidmlx.com/install.sh | bash

or Homebrew — prebuilt bottle straight from homebrew-core:

brew install rapid-mlx

Both land the same rapid-mlx CLI. The curl installer additionally installs Python 3.10+ if missing, creates an isolated venv at ~/.rapid-mlx/, symlinks the rapid-mlx CLI into ~/.local/bin/, and prints a serve command sized to your Mac (8–15 GB → lfm2.5-2.6b-4bit; 16–17 GB → qwen3.5-4b-4bit; 18–23 GB → qwen3.5-9b-4bit; 24–31 GB → bonsai-27b-2bit; 32–63 GB → gemma-4-26b-4bit; 64–95 GB → qwen3.6-35b-8bit; 96 GB+ → qwen3.5-122b-mxfp4).

Install security. install.sh is served over HTTPS (HSTS-preload) from rapidmlx.com and is a byte-identical mirror of install.sh at the release commit — read it before running if you like. If you want a cryptographically verified installer rather than trusting the website pipe, don't curl | bash the URL above: instead download the release's install.sh asset, verify it against the cosign-signed SHA256SUMS.txt shipped alongside it, and run that verified copy — full recipe in SECURITY.md. PyPI artifacts additionally carry Sigstore attestations (PEP 740). Two more low-trust paths:

  • Pin to a commit hashcurl -fsSL https://raw.githubusercontent.com/raullenchai/Rapid-MLX/<commit>/install.sh -o install.sh && shasum -a 256 install.sh && bash install.sh
  • Skip the shell script entirely — use Homebrew, uv, or pip below.

See Alternative install methods for the non-curl paths.

2. Chat with a model right now:

rapid-mlx chat

Defaults to qwen3.5-4b-4bit. First run downloads the weights (~2.5 GB) with a progress bar and drops you into a REPL. Type /help for slash commands, /exit to quit.

3. Or serve it for use from other apps:

rapid-mlx serve qwen3.5-4b-4bit

Starts an OpenAI-compatible HTTP server bound to http://localhost:8000. Point any client that supports a local custom endpoint (Aider, LangChain, OpenCode, PydanticAI, your own scripts) at http://localhost:8000/v1; Claude Code / Anthropic SDK uses http://localhost:8000 (the Anthropic messages route lives at /v1/messages under the same host).

curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"default","messages":[{"role":"user","content":"Say hello"}]}'
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")
print(client.chat.completions.create(
    model="default",
    messages=[{"role": "user", "content": "Say hello"}],
).choices[0].message.content)

4. Or wire up your coding agent — one command:

rapid-mlx launch claude-code

With a server running (step 3), this patches Claude Code's local config (~/.config/claude/settings.json) to route at http://localhost:8000 — no manual env vars, no editing JSON by hand. You get a fully local Claude Code: $0 per token, nothing leaves your Mac. Swap in cline or continue-dev for the other IDE clients, or run rapid-mlx launch list to see what's detected on this machine.

Cursor: Cursor currently routes BYOK requests through its own servers, so its servers cannot reach a Rapid-MLX endpoint on localhost. Rapid-MLX therefore does not generate a Cursor localhost config. If you intentionally expose the server through a public HTTPS tunnel, set RAPID_MLX_API_KEY=your-secret for both rapid-mlx serve ... and rapid-mlx launch cursor --server-url https://your-public-host. This is no longer a fully local connection; never expose an unauthenticated server. Rapid-MLX rejects explicit local/private addresses but cannot verify reachability from Cursor's network, whose DNS view may differ from your Mac.

Vision / audio / video / diffusion models? Base install is text-only (~460 MB). Vision, audio (TTS, STT, voice cloning), video generation, embeddings, and DFlash speculative decoding ship as opt-in extras. → Optional extras

Not into the terminal? Rapid-MLX Desktop bundles the same engine inside a one-click Mac app.


Video generation

Run text-to-video or image-to-video locally through the OpenAI-compatible Videos API. Three backends ship — Wan 2.1 / 2.2, CogVideoX-Fun and LTX-2.3 — across 8 registered checkpoints. wan2.2-ti2v-5b-q8 is the recommended starting point: smallest of the Wan set, and TI2V means one checkpoint does both text-to-video and image-to-video.

Requires Python 3.11+ (the video runtime does not support 3.10; core text and audio still do) and ffmpeg for the final MP4 mux.

pip install 'rapid-mlx[video]'
brew install ffmpeg
rapid-mlx serve wan2.2-ti2v-5b-q8

Create and download a clip:

curl http://localhost:8000/v1/videos \
  -F model=wan2.2-ti2v-5b-q8 \
  -F 'prompt=A fox running through fresh snow, cinematic tracking shot' \
  -F seconds=1 \
  -F size=832x512

# Poll until GET /v1/videos/VIDEO_ID reports "status": "completed", then:
curl http://localhost:8000/v1/videos/VIDEO_ID/content -o output.mp4

The create call returns a job immediately. Poll GET /v1/videos/VIDEO_ID until status is completed. Add -F input_reference=@start.png for image-to-video.

Generation is serialized — one clip at a time — because two diffusion pipelines resident at once will exhaust unified memory. Expect minutes of compute per second of footage, not real time.

Every checkpoint, RAM requirement and tuning knob


Audio: speech, transcription, voice cloning

41 audio aliases behind the OpenAI-compatible /v1/audio/* endpoints — any OpenAI SDK works unchanged.

pip install 'rapid-mlx[audio]'

# Text to speech
rapid-mlx serve kokoro
curl http://localhost:8000/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{"model":"kokoro","input":"hello from rapid-mlx"}' --output hello.wav

# Transcription (Whisper / Parakeet / SenseVoice)
rapid-mlx serve whisper-large-v3-turbo
curl http://localhost:8000/v1/audio/transcriptions \
  -F file=@hello.wav -F model=whisper-large-v3-turbo

Beyond the basics, three things you may not expect to run locally:

  • Zero-shot voice cloning from a reference clip. indextts is the only one that takes the clip alone; qwen3-tts-clone, f5-tts-zh and chatterbox all require ref_text (the clip's exact transcript) paired with ref_audio, and the request is rejected before generation if it is missing.
  • Voice designqwen3-tts-voicedesign has no named speakers at all. Describe the voice you want in natural language via instructions (timbre, gender, age, accent, emotion, prosody) and it synthesises it.
  • Forced alignmentqwen3-aligner takes audio plus the transcript you already have and returns per-character timings. It never guesses at the words, so it cannot mis-hear them; that is what karaoke captions and beat-synced editing need.

Also: word-level timestamps on transcription, and local text-to-music at /v1/audio/music.

All 41 aliases across 12 families


Why Rapid-MLX

| | | |---|---| | Apple-Silicon-native | Pure MLX kernels — no llama.cpp fallback, no Metal shim. Continuous batching, prompt cache (radix + DeltaNet RNN snapshots), and a quantized live KV cache (int4/int8 on the continuous-batching cache + TurboQuant K8V4 codec) run at native MLX bandwidth on M1 → M4. | | Drop-in OpenAI / Anthropic API | /v1/chat/completions, /v1/responses (Codex CLI), /v1/messages (Anthropic SDK / Claude Code), /v1/embeddings, /v1/audio/*, /v1/videos — same wire as ChatGPT / Claude, no client adapter. | | First-class ecosystem coverage | 11 agent CLIs and 3 Python frameworks are wire-verified against real weights every release (4 are Tier-1, re-verified on current binaries) — Codex CLI, Claude Code, OpenCode, Qwen Code, OpenHands, Hermes Agent, Aider, Kilo Code, GitHub Copilot, Factory Droid, Moonshot Kimi Code + LangChain, PydanticAI, smolagents. |

Full feature breakdown


Use Cases

| | | | |---|---|---| | Chat in the terminal | rapid-mlx chat qwen3.5-9b-4bit | Streaming REPL, /help for slash commands, --think / --no-think to control CoT. | | OpenAI server for your apps | rapid-mlx serve qwen3.5-9b-4bit | Point Aider, LibreChat, Open WebUI, or LangChain at http://localhost:8000/v1. | | Agent backends | rapid-mlx serve qwen3.6-35b-8bit &<br>rapid-mlx agents codex --setup && codex | 8 agents auto-configure via agents <name> --setup once the server is up (11 wire-verified total, 4 Tier-1) — see Agent support. | | Benchmark your Mac | rapid-mlx bench qwen3.5-9b-4bit --submit | Standardized B=1 bench, opens a PR to publish your row on rapidmlx.com. |

One-shot IDE setup with rapid-mlx launch <claude-code|cline|continue-dev>


Agent Support

All 11 agents below are wire-verified against real weights every release via their own integration-test cell. Of these, four are Tier-1Claude Code, Codex CLI, Hermes, and Aider — re-verified end-to-end against the current client binary every release, with one guardian per API wire (Anthropic /v1/messages, OpenAI /v1/responses, and /v1/chat/completions covered for both to

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars3.8k
CategoryAI
Updated4h ago
Forks417

Languages

Python

Security Score

88/100

Audited on Sep 21, 2026

1 medium