18 skills found
SamurAIGPT / Generative-Media-SkillsMulti-modal Generative Media Skills for AI Agents (Claude Code, Cursor, Gemini CLI). High-quality image, video, and audio generation powered by muapi.ai.
jordanrendric / claude-video-visionGive Claude the ability to watch and understand videos — Claude Code plugin with frame extraction and multimodal audio analysis
youichi-uda / godot-mcp-pro162 MCP tools for AI-powered Godot 4 development. Scene, animation, 3D, physics, particles, audio, shader, input simulation, runtime analysis, navigation, testing & more. $15 one-time.
artokun / comfyui-mcpLocal-first, agent-native control plane for ComfyUI — MCP server + sidebar agent that generates images, video & audio, authors and runs workflows, and edits your live graph in natural language on ANY LLM (Claude, ChatGPT, Gemini, offline Ollama, or any hosted model).
AgriciDaniel / claude-shortsInteractive longform-to-shortform video creator — Claude Code skill with Remotion-rendered animated captions, AI segment scoring, cursor tracking, and audio-aware boundary snapping
xDarkzx / Reaper-MCPAI-powered music production and post-production in REAPER via MCP — 172 tools for composition, mixing, mastering, batch editing, audio QC, and ReaScript automation.
jftuga / transcript-criticClaude Code skill that transcribes audio/video with whisper.cpp to get structured critical analysis including timestamped summaries, evidence notes, logical fallacies, and underdeveloped areas
wells1137 / media-skillsA collection of open-source Agent Skills for content creation — images, audio, and video.
Anil-matcha / Veo-4-APIPython wrapper for Veo 4 API by Google DeepMind — native 4K AI video with integrated audio, character consistency & advanced camera controls.
aezizhu / mcp-captcha-solverCaptcha solver MCP server with 99% success rate. Supports reCAPTCHA v2/v3, hCaptcha, FunCaptcha, GeeTest, Cloudflare Turnstile, slider, rotate, audio captcha. Multi-service fallback.
HarperZ9 / gatherResearch intake that reaches the hard places: web, video, papers, scanned PDFs, browser, OCR, and audio into structured research packets. DOM extraction and change tracking built in; provenance rides along on every item.
second-state / audio-ttsGenerate speech audio from text using Qwen3 TTS, or clone a voice from reference audio. Triggered when the user wants to convert text to speech, generate audio, read text aloud, or clone/mimic a voice. Supports multiple speakers, English and Chinese, and emotion/style control.
alessandro9110 / Speech-To-Text-With-DatabricksAn end-to-end, scalable STT solution on Databricks that transcribes audio into structured text in Delta Lake, ready for analytics, search, and GenAI/RAG.
0xquqi / sun-skillJustin Sun Cognitive Framework - 孙宇晨认知操作系统 | 21,829 tweets + 288-page autobiography + 155-episode audio course + 269+ sources → 14 mental models + 19 decision heuristics
KingJing1 / podcast-transcript-txtDeterministic workflow to find and export full podcast transcripts as cleaned TXT files from YouTube URLs, episode webpages (including Xiaoyuzhou), Apple Podcasts title search, X/Twitter links, direct audio URLs, or plain episode titles
ychoi-kr / ffmpeg-usageffmpeg recipes and best practices: convert, concatenate, merge, resize, compress, GIF creation, audio extraction, subtitles, optimize for social platforms.
xmg2024 / sun-skillJustin Sun Cognitive Framework - 孙宇晨认知操作系统 | 21,829 tweets + 288-page autobiography + 155-episode audio course + 269+ sources → 14 mental models + 19 decision heuristics
sakamoto-family-smile / videodbSee, Understand, Act on video and audio. See- ingest from local files, URLs, RTSP/live feeds, or live record desktop; return realtime context and playable stream links. Understand- extract frames, build visual/semantic/temporal indexes, and search moments with timestamps and auto-clips.