10 skills found
AI272 / speakerSpeaker is a Codex skill project for academic presentations: read real.pptx, combine text extraction, PPTX structure parsing, page-by-page rendering, OCR, and visual review to generate page-by-page speaker notes, and write a clean version of the lecture into the PowerPoint comment area.
second-state / audio-ttsGenerate speech audio from text using Qwen3 TTS, or clone a voice from reference audio. Triggered when the user wants to convert text to speech, generate audio, read text aloud, or clone/mimic a voice. Supports multiple speakers, English and Chinese, and emotion/style control.
NarratorAI-Studio / narrator-ai-cli-skillAI 解说大师 — Agent skill;封装 narrator-ai-cli 供 Claude/Codex 等工具调用
alessandro9110 / Speech-To-Text-With-DatabricksAn end-to-end, scalable STT solution on Databricks that transcribes audio into structured text in Delta Lake, ready for analytics, search, and GenAI/RAG.
howdoiusekeyboard / TrueVoice-MCPA Model Context Protocol server that helps AI generate human-like text without AI slop
waxberry-dev / live-translate-mcpMCP server for local speech translation (EN ↔ 中文) via Whisper + Claude + Piper
modelscope / FunASROpen-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
aws-samples / recorded-voice-insight-extraction-webappA generative AI tool to boost productivity by transcribing and analyzing audio or video recordings containing speech
julilaoshi / flowmotion-skillFlowMotion Skill / 头脑风暴动态流程图 Skill: turn voice notes, SRT transcripts, and messy ideas into structured flow specs, connection specs, and motion-ready diagram briefs.
giannisanni / kokoro-tts-mcpMCP server for Kokoro text-to-speech, with adjustable voices/speed and an optional OpenAI-compatible (kokoro-fastapi) backend.