211 skills found · Page 1 of 8
QwenLM / Qwen3 OmniQwen3-omni is a natively end-to-end, omni-modal LLM developed by the Qwen team at Alibaba Cloud, capable of understanding text, audio, images, and video, as well as generating speech in real time.
lyuchenyang / Macaw LLMMacaw-LLM: Multi-Modal Language Modeling with Image, Video, Audio, and Text Integration
DAMO-NLP-SG / VideoLLaMA2VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Yuan-ManX / AI Game DevtoolsYour AI Game Dev Hub. The ultimate resource hub for AI-powered game development tools. Discover cutting-edge LLMs, World Model, Agent, Code, Image, Texture, Shader, 3D Model, Animation, Video, Audio, Music, Singing Voice and Analytics. 🔥
ga642381 / Speech TridentAwesome speech/audio LLMs, representation learning, and codec models
stepfun-ai / Step Audio EditXA powerful 3B-parameter, LLM-based Reinforcement Learning audio edit model excels at editing emotion, speaking style, and paralinguistics, and features robust zero-shot text-to-speech
BinWang28 / Audio AI HubThe hub for audio AI research: papers, open models, benchmarks & datasets across audio LLMs, speech recognition, TTS, music & audio generation.
kardolus / Chatgpt CliChatGPT CLI is a powerful, multi-provider command-line interface for working with modern LLMs. It supports OpenAI, Azure, Perplexity, LLaMA, and more, with features like streaming, interactive chat, prompt files, image/audio I/O, MCP tool calls, and an experimental agent mode for safe, multi-step automation.
EmulationAI / Awesome Large Audio ModelsCollection of resources on the applications of Large Language Models (LLMs) in Audio AI.
NVlabs / OmniVinciOmniVinci is an omni-modal LLM for joint understanding of vision, audio, and language.
YingqingHe / Awesome LLMs Meet Multimodal Generation🔥🔥🔥 A curated list of papers on LLMs-based multimodal generation (image, video, 3D and audio).
Audio-AGI / WavJourneyWavJourney: Compositional Audio Creation with LLMs
artokun / Comfyui MCPLocal-first, agent-native control plane for ComfyUI — MCP server + sidebar agent that generates images, video & audio, authors and runs workflows, and edits your live graph in natural language on ANY LLM (Claude, ChatGPT, Gemini, offline Ollama, or any hosted model). 178 tools, 36 AI skills, 55 installer packs. Local, LAN, VPS, or Comfy Cloud.
apocas / RestaiRESTai is an AIaaS (AI as a Service) open-source platform. Supports many public and local LLM suported by Ollama/vLLM/etc. Precise embeddings usage, tuning, analytics etc. Built-in image/audio generation with dynamic loading generators. Live chat deployment. Built-in block based graphical language. Prompt versioning and much more...
TheDeathDragon / LiveTranslateReal-time audio translation, captures system audio + mic, runs ASR (Whisper/SenseVoice), translates via LLM API with streaming display. Perfect for VTubers, livestreamers, and watching foreign content. Windows 实时音频翻译,ASR 语音识别后 LLM 流式翻译显示,适合 VTuber、主播和外语视频观看。
alibaba / Unified AudioAn Open-Source Project to Unify Audio Processing and Generation
infiniV / VoiceFlowLocal voice dictation and meeting recorder for Windows + Linux. Hold a hotkey to dictate, or record long-form meetings with system audio. Whisper transcription, bring-your-own-LLM summaries. Open source.
QwenAudio / FunAudioLLM APPNo description available
ptnghia-j / ChordMiniAppMusic Analysis, Chord Recognition, Beat Tracking, Guitar Diagrams, Piano Visualizer, Lyrics Transcription Application, context-aware LLM inference for analysis from uploaded audio and YouTube video
LingyiChen-AI / Comfyui Workflow SkillNatural language → ComfyUI workflow JSON. 34 built-in templates, 360+ node definitions, auto model download. Supports txt2img, img2img, txt2vid, img2vid, audio, 3D generation across SD1.5/SDXL/SD3/FLUX/Wan2.2/HunyuanVideo/LTXV/Mochi/Cosmos + LLM integration. Works as a skill for Claude Code, Cursor, and other AI coding agents.