633 skills found · Page 1 of 22
AlexsJones / LlmfitHundreds of models & providers. One command to find what runs on your hardware.
mozilla-ai / LlamafileDistribute and run LLMs with a single file.
Andyyyy64 / WhichllmFind the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.
Michael-A-Kuykendall / Shimmy⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
KoboldAI / KoboldAI ClientFor GGUF support, see KoboldCPP: https://github.com/LostRuins/koboldcpp
city96 / ComfyUI GGUFGGUF Quantization support for native ComfyUI models
off-grid-ai / OGAMThe Swiss Army Knife of Offline AI. Chat, see, speak, and generate images on your phone or Mac — GGUF LLMs, vision, Whisper speech-to-text, Stable Diffusion, tool calling, and local-network servers. Runs on your CPU, GPU, or NPU. No account, no API key, zero data leaves your device.
janhq / Cortex.cppLocal AI API Platform
Mobile-Artificial-Intelligence / MaidMaid is a free and open source application for interfacing with llama.cpp models locally, and with Anthropic, DeepSeek, Ollama, Mistral and OpenAI models remotely.
datawhalechina / Handy Ollama动手学Ollama,CPU玩转大模型部署,在线阅读地址:https://datawhalechina.github.io/handy-ollama/
heshengtao / Comfyui LLM PartyLLM Agent Framework in ComfyUI includes MCP sever, Omost,GPT-sovits, ChatTTS,GOT-OCR2.0, and FLUX prompt nodes,access to Feishu,discord,and adapts to all llms with similar openai / aisuite interfaces, such as o1,ollama, gemini, grok, qwen, GLM, deepseek, kimi,doubao. Adapted to local llms, vlm, gguf such as llama-3.3 Janus-Pro, Linkage graphRAG
withcatai / Node Llama CppRun AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation level
sammcj / GollamaGo manage your Ollama models
handy-computer / Transcribe.cppggml speech-to-text inference for 16+ model families
AtomicBot-ai / Atomic AgentLocal First Ai Agent. Optimized for Local Ai models. Long context window. Proper tools callings. Runs privately on your device.
intel / Auto RoundA SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
QwenAudio / Fun ASROpen-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR, vLLM, streaming, and llama.cpp runtimes.
edwko / OuteTTSInterface for OuteTTS models.
kitops-ml / KitopsAn open source DevOps tool from the CNCF for packaging and versioning AI/ML models, datasets, code, and configuration into an OCI Artifact.
AtomicBot-ai / Atomic ChatLocal AI app and inference engine for agents. Run open-weight LLMs locally — private, 100% offline on your computer. Join our Discord: https://discord.com/8wGSsvmg4V