SkillAgentSearch skills...

Voicehub

VoiceHub: A Unified Inference Interface for TTS Models

Install / Use

npx skills add kadirnar/voicehub

Installs into whichever agent you are using.

About this skill

Quality Score

0/100

Supported Platforms

Universal

README

<h2 align="center">Unified Inference, Training, and Optimization for TTS, ASR, and VAD</h2> <div align="center"> <img width="100%" alt="Abstract sound waves representing VoiceHub's unified speech toolkit" src="https://raw.githubusercontent.com/kadirnar/voicehub/main/assets/readme-hero.png"> </div>

Install

VoiceHub supports Python 3.10 through 3.12.

python -m pip install voicehub

For fine-tuning:

python -m pip install "voicehub[training]"

GPU users should install the correct PyTorch build for their machine first. See the installation guide.

Verify the package without downloading a checkpoint:

python -c "import voicehub; print(voicehub.__version__, len(voicehub.list_model_specs()))"

TTS

Use a long sample and verify the generated duration. Speaking rate varies by model, so duration must be checked from the waveform rather than assumed from word count.

from voicehub import AutoModelForTextToSpeech, TTSGenerationConfig

text = (
    "Welcome to VoiceHub. This sample is intentionally long enough for a "
    "meaningful speech test. It checks pronunciation, pacing, sentence "
    "transitions, and sustained audio quality while the speaker explains a "
    "simple workflow for reliable text to speech inference. During this "
    "longer passage, listen for stable volume, natural pauses, clear word "
    "endings, and consistent tone from the opening sentence through the "
    "final measurement."
)

tts_model = AutoModelForTextToSpeech.from_pretrained(
    "parler-tts/parler-tts-mini-v1",
    model_type="parlertts",
    device="cuda",
)
output = tts_model.generate(
    text,
    description="A clear speaker talks at a natural, relaxed pace.",
    generation_config=TTSGenerationConfig(
        seed=42,
        output_file="tts-sample.wav",
    ),
)

samples = (
    output.audio.shape[-1] if hasattr(output.audio, "shape") else len(output.audio)
)
duration = samples / output.sample_rate
if duration < 10:
    raise RuntimeError(f"Expected at least 10 seconds, generated {duration:.2f}")
print(output.file_path, f"{duration:.2f}s")

Model-specific conditioning fields such as speaker references, voices, or descriptions are listed in the TTS model matrix.

ASR

from voicehub import AutoModelForSpeechRecognition

asr_model = AutoModelForSpeechRecognition.from_pretrained(
    "Qwen/Qwen3-ASR-0.6B",
    model_type="asr_qwen3",
    device="cuda",
)
output = asr_model.transcribe("speech.wav", language="English")
print(output.text)

VAD

from voicehub import AutoModelForVoiceActivityDetection

vad_model = AutoModelForVoiceActivityDetection.from_pretrained(
    model_type="vad_silero",
)
output = vad_model.detect("speech.wav", threshold=0.55)
for segment in output.segments:
    print(segment.start, segment.end)

See the ASR and VAD matrix for checkpoints, inputs, outputs, and training boundaries.

Optimize TTS

Start with eager inference, then benchmark one change at a time on the same text, seed, warm-up count, and device.

from voicehub import TTSOptimizationConfig

result = tts_model.optimize(
    TTSOptimizationConfig(
        attn_implementation="auto",
        kernel_backend="auto",
        compile="auto",
    )
)
print(result.manifest())

The optimization result records what was applied and what stayed on the quality-preserving fallback. Do not publish speed or memory percentages from configuration alone; measure them on the target hardware. Use the optimization guide, TTS model benchmarks, and current RTX 4090 speech results for reproducible comparisons.

Fine-tune

Every trainable integration advertises its exact objective and data contract. Check support before loading weights:

from voicehub import get_training_spec

spec = get_training_spec("dia")
print(spec.support.value, spec.family_name)

Then begin with a one-step smoke run. The training guide and training matrix show the required dataset fields, frozen components, checkpoint type, and export path. The data guide and ASR/VAD data guide cover manifests and leakage-safe splits.

Notebooks

The notebooks use a short, top-to-bottom workflow: install, configure, run, and inspect.

| Notebook | GitHub | Colab | | --------------------------- | --------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------ | | TTS, ASR, and VAD inference | View | Run | | Data preparation | View | Run | | Fine-tuning | View | Run | | Dia end-to-end workflow | View | Run |

For a dedicated inference page for each Hub-backed model, open the Hugging Face model notebook gallery.

Read the notebook guide for expected hardware and opt-in execution flags.

Documentation

Development

git clone https://github.com/kadirnar/voicehub.git
cd voicehub
python -m pip install -e ".[test,training]"
python -m pytest
python scripts/check_distribution.py

check_distribution.py builds the wheel and source distribution, installs the wheel, sdist, and editable checkout in separate environments, and checks lazy import plus required package data. It skips PyTorch downloads by default; pass --with-dependencies on a release machine for full dependency installs.

License

VoiceHub is licensed under Apache-2.0. Vendored components retain their own license notices in their package directories. Checkpoint licenses are separate from source-code licenses and must be reviewed before use.

Related Skills

View on GitHub
GitHub Stars104
CategoryDevelopment
Updated1d ago
Forks9

Languages

Python

Security Score

100/100

Audited on Aug 7, 2026

No findings