lemonade
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk
Install / Use
claude mcp add lemonade-sdk -- npx -y github:lemonade-sdk/lemonadeIf the server publishes to npm under a different name, use that package instead — check the repo README.
MCP Server
Model Context Protocol server
Quality Score
Category
CommunicationSupported Platforms
Skill content
View source on GitHub🍋 Lemonade: Refreshingly fast local AI
<p align="center"> <a href="https://discord.gg/5xXzkMu8Zk"> <img src="https://img.shields.io/badge/Discord-7289DA?logo=discord&logoColor=white" alt="Discord" /></a> <a href="https://github.com/lemonade-sdk/lemonade/blob/main/docs/dev/contribute.md" title="Contribution Guide"> <img src="https://img.shields.io/badge/PRs-welcome-brightgreen.svg" alt="PRs Welcome" /></a> <a href="https://github.com/lemonade-sdk/lemonade/releases/latest" title="Download the latest release"> <img src="https://img.shields.io/github/v/release/lemonade-sdk/lemonade?include_prereleases" alt="Latest Release" /></a> <a href="https://tooomm.github.io/github-release-stats/?username=lemonade-sdk&repository=lemonade"> <img src="https://img.shields.io/github/downloads/lemonade-sdk/lemonade/total.svg" alt="GitHub downloads" /></a> <a href="https://github.com/lemonade-sdk/lemonade/issues"> <img src="https://img.shields.io/github/issues/lemonade-sdk/lemonade" alt="GitHub issues" /></a> <a href="https://github.com/lemonade-sdk/lemonade/blob/main/LICENSE"> <img src="https://img.shields.io/badge/License-Apache-yellow.svg" alt="License: Apache" /></a> <a href="https://star-history.com/#lemonade-sdk/lemonade"> <img src="https://img.shields.io/badge/Star%20History-View-brightgreen" alt="Star History Chart" /></a> </p> <p align="center"> <img src="https://github.com/lemonade-sdk/assets/blob/main/docs/banner_02.png?raw=true" alt="Lemonade Banner" /> </p> <h3 align="center"> <a href="https://lemonade-server.ai/docs/guide/install/">Download</a> | <a href="https://lemonade-server.ai/docs/">Documentation</a> | <a href="https://discord.gg/5xXzkMu8Zk">Discord</a> </h3>Lemonade is the local AI server that gives you the same capabilities as cloud APIs, except 100% free and private. Use the latest models for chat, coding, speech, and image generation on your own NPU and GPU.
Lemonade comes in two flavors:
- Lemonade Server installs a service you can connect to hundreds of great apps using standard OpenAI, Anthropic, and Ollama APIs.
- Embeddable Lemonade is a portable binary you can package into your own application to give it multi-modal local AI that auto-optimizes for your user’s PC.
This project is built by the community for every PC, with optimizations by AMD engineers to get the most from Ryzen AI, Radeon, and Strix Halo PCs.
Getting Started
- Install: Windows · Linux · macOS · Docker · Source
- Get Models: Browse and download with the Model Manager
- Generate: Try models with the built-in interfaces for chat, image gen, speech gen, and more
- Mobile: Take your lemonade to go: iOS · Android · Source
- Connect: Use Lemonade with your favorite apps:
Supported Platforms
| Platform | Build |
|----------|-------|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
Using the CLI
To run and chat with Gemma:
lemonade run Gemma-4-E2B-it-GGUF
To code with Lemonade models:
lemonade launch claude
Multi-modality:
# image gen
lemonade run SDXL-Turbo
# speech gen
lemonade run kokoro-v1
# transcription
lemonade run Whisper-Large-v3-Turbo
To see available models and download them:
lemonade list
lemonade pull Gemma-4-E2B-it-GGUF
To see the backends available on your PC:
lemonade backends
For hybrid setups, Lemonade can also route to any OpenAI-compatible cloud provider (Fireworks, OpenAI, OpenRouter, Together, …) alongside local models — see Cloud Offload. (Experimental.)
Model Library
<img align="right" src="https://github.com/lemonade-sdk/assets/blob/main/docs/model_manager_02.png?raw=true" alt="Model Manager" width="280" />Lemonade supports a wide variety of LLMs (GGUF, FLM, and ONNX), whisper, stable diffusion, etc. models across CPU, GPU, and NPU.
Use lemonade pull or the built-in Model Manager to download models. Custom GGUF/ONNX models can be pulled from Hugging Face or ModelScope, with their source retained for future updates.
Supported Configurations
Lemonade supports multiple inference engines for LLM, speech, TTS, and image generation, and each has its own backend and hardware requirements.
<!-- BEGIN GENERATED: backends-matrix --> <table> <thead> <tr> <th>Modality</th> <th>Engine</th> <th>Backend</th> <th>Device</th> <th>OS</th> </tr> </thead> <tbody> <tr> <td rowspan="9"><strong>Text generation</strong></td> <td rowspan="6"><code>llamacpp</code></td> <td><code>system</code></td> <td><code>x86_64</code>/ARM64 CPU, GPU</td> <td>Linux</td> </tr> <tr> <td><code>metal</code></td> <td>Apple Silicon GPU</td> <td>macOS</td> </tr> <tr> <td><code>cuda</code></td> <td>NVIDIA GPUs (Turing or newer)**</td> <td>Windows, Linux</td> </tr> <tr> <td><code>vulkan</code></td> <td><code>x86_64</code> CPU, AMD iGPU, AMD dGPU; ARM64 CPU/GPU (Linux)</td> <td>Windows, Linux</td> </tr> <tr> <td><code>rocm</code></td> <td>Supported AMD ROCm iGPU/dGPU families, incl. AMD Instinct MI300X (gfx942) and MI350X (gfx950, Linux + stable only)*</td> <td>Windows, Linux</td> </tr> <tr> <td><code>cpu</code></td> <td><code>x86_64</code> CPU; ARM64 CPU (Linux)</td> <td>Windows, Linux</td> </tr> <tr> <td rowspan="1"><code>flm</code></td> <td><code>npu</code></td> <td>XDNA2 NPU</td> <td>Windows, Linux</td> </tr> <tr> <td rowspan="1"><code>ryTruncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
84.2kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
73.4kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
ruflo
73.0k🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
CowAgent
47.1kOpen-source super AI assistant & Agent Harness. Plans tasks, runs tools and skills, self-evolves with memory and knowledge. Multi-agent, multi-model, multi-channel. Lightweight, extensible, one-line install.
