oma-voice
Generate speech or transcribe audio locally with Voicebox. Use for
Install / Use
npx skills add first-fluke/oh-my-agent --skill oma-voiceInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
OtherSupported Platforms
Our assessment of oma-voice
oma-voice scores 89/100 on our quality scale, 44th of 158 Other skills we index (top 28%).
Its SKILL.md is 13 KB long, well organised into 36 sections with 2 code examples: a thorough specification that gives an agent plenty to work with.
With 1,324 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 7 days ago, so oma-voice is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
oma-voice compared with similar skills
All 4 of these similar skills score higher than oma-voice; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| oma-voice (this skill)by first-fluke | 89 | 1.3k | 7d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 86.6k | 15d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.2k | today | CLAUDE.md |
| rufloby ruvnet | 100 | 73.6k | today | CLAUDE.md |
| CowAgentby zhayujie | 100 | 47.2k | today | CLAUDE.md |
Frequently asked questions
- How do I install oma-voice?
- Run
npx skills add first-fluke/oh-my-agent --skill oma-voice. The install tabs above show the steps for each supported agent. - Which AI agents does oma-voice work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is oma-voice safe to use?
- It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is oma-voice still maintained?
- The repository was last updated 7 days ago, so oma-voice is actively maintained.
Skill content
View source on GitHubname: oma-voice description: Generate speech or transcribe audio locally with Voicebox. Use for narration, voice assets, dictation, and meeting transcription.
Voice Skill - Local TTS and STT via Voicebox
Scheduling
Goal
Drive the Voicebox local app through its MCP server so any MCP-aware agent can speak (TTS) or listen (STT) without invoking cloud vendors. The skill standardizes intent routing, voice profile resolution, output layout, and guardrails while voicebox itself owns the engines, voice cloning UI, captures archive, and stories editor.
Intent signature
- User asks to generate speech, narrate text, produce a voiceover, create an mp3 or wav from text.
- User wants an audio file transcribed into text, meeting notes, or a transcript.
- User asks for a voice notification when a long task completes or a workflow step is blocked.
- Another skill needs local audio generation infrastructure.
When to use
- Generating short notification audio for agent task completion or blockers.
- Producing voiceover, narration, or audio assets (mp3 or wav) for apps and content.
- Transcribing local audio files (mp3, wav, m4a, webm, flac) to Markdown.
- Comparing voice profiles by re-running the same text against different profile ids.
When NOT to use
- Cloud TTS or high-fidelity multilingual cloud voices -> out of scope; future multi-vendor extension.
- Real-time microphone dictation loop in the terminal -> use Voicebox app's built-in hotkey dictation.
- Voice cloning sample upload and profile creation -> done in the Voicebox desktop app UI.
- Video synthesis, music, sound design -> out of scope.
- Stories Editor multi-voice timeline composition -> use the Voicebox app UI.
Expected inputs
- TTS: text (<= 5000 chars per call), optional profile id, optional engine, optional language, optional output path.
- STT: audio file path (absolute or relative to
$CWD), optional language hint. - Notification: short message (<= 240 chars), profile id resolved from config.
Expected outputs
- TTS: audio file (
wav— Voicebox's native TTS output format;mp3optional via a local ffmpeg transcode) at.agents/results/voice/{timestamp}-{shortid}/output.{wav|mp3}plusmanifest.json. - STT:
transcript.mdat.agents/results/voice/transcripts/{timestamp}-{shortid}/plusmanifest.json. - Notification: ephemeral playback through Voicebox; no disk write by default.
Dependencies
- Voicebox desktop app installed and running locally.
- Voicebox MCP registered (
claude mcp add --transport http voicebox http://127.0.0.1:17493/mcp). - TTS only: at least one voice profile created in the Voicebox app UI.
- TTS only: optionally pre-downloaded engine models for the selected profile.
Control-flow features
- Branches by mode (notify, asset, transcribe), language, and profile availability.
- Calls voicebox via MCP tools, with REST
GET /healthas the handshake probe. - Reads input audio files and writes generated audio plus manifests.
- Caches discovered MCP tool names after the first successful
tools/list.
Structural Flow
Entry
- Detect the requested mode: notification, asset TTS, or transcription.
- Verify Voicebox is reachable via MCP handshake or
GET /health. - On the first run only, call MCP
tools/listand cache the resolved tool names. - For notification or asset TTS, resolve the target voice profile id. For transcription, validate the audio input and continue without a profile.
Scenes
- PREPARE: Validate text length, audio duration, language, output path, and profile id.
- ACQUIRE: If a required signal is missing, run the clarification protocol once.
- ACT: Invoke the appropriate MCP tool (TTS or STT) with the resolved parameters.
- VERIFY: Confirm the response carries audio output or transcript content. Validate manifest fields.
- FINALIZE: Write
manifest.jsonalongside the output. Report the path or transcript to the user.
Transitions
- If voicebox is unreachable, surface the install or launch hint and exit. Do not attempt auto-relaunch.
- If a TTS request has no usable profile, point the user at the Voicebox app UI to create a profile, then exit. A transcription request never needs a profile.
- If a TTS request exceeds 5000 chars, ask whether to truncate or split. Do not auto-chunk in v1.
- If an STT input exceeds 30 minutes, ask whether to proceed. Do not auto-split.
- If the selected engine model is not loaded, ask the user before triggering a download.
Failure and recovery
| Failure | Recovery |
|---------|----------|
| Voicebox app not running | Print install/launch hint, exit code 5 |
| No voice profile for TTS | Print "create a profile in Voicebox" hint, exit code 3 |
| Engine model missing | Ask before triggering download |
| Output path outside $PWD | Use an explicitly requested path; ask only if the destination is ambiguous or overwrites unrelated data |
| TTS over 5000 chars | Ask the user to split or truncate |
| STT over 30 minutes | Confirm only if the requested duration or resource cost is unresolved |
| MCP tool name drift | Re-run tools/list and update the cache |
| SIGINT | Abort the MCP call, write no partial output |
Exit
- Success: audio file or transcript exists with a complete manifest, and the path is reported.
- Partial success: output exists but a guardrail warning is surfaced (length, disk, model fallback).
- Failure: no output, the blocker (auth, profile, engine, network) is explicit.
Logical Operations
Actions
| Action | SSL primitive | Evidence |
|--------|---------------|----------|
| Validate mode and inputs | VALIDATE | Clarification protocol in execution-protocol.md |
| Resolve TTS voice profile | SELECT | voicebox_list_profiles + config defaults |
| Health check | READ | MCP handshake or GET /health |
| Generate speech | CALL_TOOL | MCP voicebox_speak |
| Transcribe audio | CALL_TOOL | MCP voicebox_transcribe |
| Write output and manifest | WRITE | Audio or transcript plus manifest.json |
| Inspect result | VALIDATE | Output presence, duration, manifest fields |
| Report result | NOTIFY | Final user-facing summary |
Tools and instruments
- Voicebox MCP server at
http://127.0.0.1:17493/mcp. - REST surface for health and audio retrieval (
GET /health,GET /audio/{generation_id}). - Resource references: voice matrix, prompt tips, execution protocol, checklist.
Canonical command path
# 1. MCP handshake or REST health
GET http://127.0.0.1:17493/health -> 200 OK
# 2. Discover tool names on first run
MCP tools/list -> cache real names
# 3. TTS only: resolve profile
MCP voicebox_list_profiles -> pick profile by name or config default
# 4. Generate or transcribe (STT skips profile lookup and model-status check)
MCP voicebox_speak { text, profile, language?, engine?, personality? }
MCP voicebox_transcribe { audio_path | audio_base64, language?, model? }
# 5. Fetch the generated audio (MCP has no save-to-disk; TTS output is wav)
GET http://127.0.0.1:17493/audio/{generation_id} -> wav bytes
# 6. Persist output + manifest
.agents/results/voice/<timestamp>-<shortid>/output.wav + manifest.json
.agents/results/voice/transcripts/<timestamp>-<shortid>/transcript.md + manifest.json
MCP tool mapping (verified against Voicebox 0.5.0)
| Use case | MCP tool | REST backing |
|---|---|---|
| TTS generation | voicebox_speak | POST /speak |
| STT transcription | voicebox_transcribe | POST /transcribe |
| Profile listing | voicebox_list_profiles | GET /profiles |
| Captures listing | voicebox_list_captures | GET /history (captures view) |
Tools not exposed via MCP (REST only): model status (GET /models/status), audio file serving (GET /audio/{generation_id}), per-version audio (GET /audio/version/{version_id}). The skill calls those over loopback HTTP when needed.
Notes on voicebox_speak:
- Required:
text. Optional:profile,engine,language,personality(bool). - Audio plays on the user speakers and is saved to the Captures / History panel automatically. There is no
save_to_disktoggle on the MCP tool itself; to persist a local copy, fetchGET /audio/{generation_id}(Voicebox TTS always storeswav). - Without a default profile set in Voicebox Settings,
profile=is required.
Notes on voicebox_transcribe:
- Accepts exactly one of
audio_base64oraudio_path(loopback only). Optionallanguage,model.
Resource scope
| Scope | Resource target |
|-------|-----------------|
| LOCAL_FS | Input audio, generated audio, transcripts, manifests |
| PROCESS | Local Voicebox app subprocess (managed by the user) |
| NETWORK | Loopback HTTP to 127.0.0.1:17493 only |
| MEMORY | Cached MCP tool names, resolved profile metadata |
| CREDENTIALS | None. Voicebox is local and key-free. |
Preconditions
- Voicebox app is running and the MCP handshake succeeds.
- TTS only: at least one voice profile exists.
- TTS only: the selected engine model is loaded or the user approves a download.
- Output directory is inside
$PWDunless explicitly allowed.
Effects and side effects
- Creates audio files, transcripts, and manifests under
.agents/results/voice/. - Triggers local Voicebox generation, which consumes CPU or GPU.
- May trigger an engine model download when the user approves.
- Does not call any cloud service. No external network traffic.
Guardrails
- Voicebox required: if the MCP handshake or
GET /healthfails, exit with a one-shot install or launch hint. Do not retry, do not auto-relaunch. - Profile required for TTS only: resolve a profile and check the TTS engine only for notification or asset mode. Transcription proceeds with a valid audio input and optional STT model even when no profile exists.
- Tool-name discovery: on first invocation, call MCP
tools/listand cache the resolved names. Reuse the cache for subsequent calls in the same session. - Length limits: TTS calls cap at 5000 chars per call; warn at 2000. STT inputs cap at 30 minutes. v1 does not auto-chunk or auto-split.
- Auto-invocation transparency: notifications fire automatically only when the active task exceeds
auto_notify_after_sec(default 60s). This threshold is agent-enforced guidance — no hook measures task duration — so apply it by judgment when a long task completes or blocks. Always announce intent in one short line before generating audio. - Path safety: an explicitly requested output path authorizes writing there. Resolve ambiguity or unrelated-data replacement before the dependent write; preserve required CLI path flags.
- Cancellation: SIGINT aborts the MCP call and writes no partial output.
- Manifest required for persisted output: asset TTS and transcription modes write
manifest.jsonwith at minimum:skill,mode,voicebox_generation_id,text(ortranscript_preview),profile,engine,language,format(TTS only),created_at. Notification mode is exempt because Voicebox Captures is its system of record and no disk output is written by default. - Out of scope: voice cloning UI, captures archive, stories editor, microphone dictation loop, and cloud vendors are intentionally not exposed.
- No cost guard: Voicebox is free. The cost guardrail from
oma-imagedoes not apply.
Clarification protocol
Before invoking a TTS or STT call, the agent checks the following. If any required signal is missing, clarify with the user first.
TTS (asset mode) required:
- [ ] Text content provided?
- [ ] Voice profile id or tone description provided?
TTS strongly recommended:
- [ ] Language explicit or detectable from the text?
- [ ] Output format (wav default; mp3 requires a local ffmpeg transcode)?
STT required:
- [ ] Audio path provided and the file exists?
- [ ] Duration within 30 minutes, or user approves splitting?
**Noti
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
86.6kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.2kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
ruflo
73.6k🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
CowAgent
47.2kOpen-source personal AI assistant & Agent Harness. Plans tasks, runs tools and skills, self-evolves with memory and knowledge. Multi-agent, multi-model, multi-channel. Lightweight, extensible, one-line install.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
