Voice Activity Detection (VAD)
Detect speech segments in audio using VAD tools like Silero VAD, SpeechBrain VAD, or WebRTC VAD
Install / Use
npx skills add benchflow-ai/skillsbench --skill voice-activity-detectionInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
OtherSupported Platforms
Our assessment of Voice Activity Detection (VAD)
Voice Activity Detection (VAD) scores 86/100 on our quality scale, 69th of 154 Other skills we index (top 45%).
Its SKILL.md is 4.5 KB long, well organised into 17 sections with 5 code examples: a solid amount of guidance for an agent.
With 1,813 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated about 2 months ago, so Voice Activity Detection (VAD) is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Voice Activity Detection (VAD) compared with similar skills
All 4 of these similar skills score higher than Voice Activity Detection (VAD); compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| Voice Activity Detection (VAD) (this skill)by benchflow-ai | 86 | 1.8k | 2mo ago | SKILL.md |
| LocalAIby mudler | 100 | 49.3k | today | MCP Server |
| algorithmic-artby anthropics | 100 | 177.9k | 7d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 7d ago | SKILL.md |
| designby nextlevelbuilder | 100 | 130.2k | 9d ago | SKILL.md |
Frequently asked questions
- How do I install Voice Activity Detection (VAD)?
- Run
npx skills add benchflow-ai/skillsbench --skill "Voice Activity Detection (VAD)". The install tabs above show the steps for each supported agent. - Which AI agents does Voice Activity Detection (VAD) work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is Voice Activity Detection (VAD) safe to use?
- It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is Voice Activity Detection (VAD) still maintained?
- The repository was last updated about 2 months ago, so Voice Activity Detection (VAD) is actively maintained.
Skill content
View source on GitHubname: Voice Activity Detection (VAD) description: Detect speech segments in audio using VAD tools like Silero VAD, SpeechBrain VAD, or WebRTC VAD. Use when preprocessing audio for speaker diarization, filtering silence, or segmenting audio into speech chunks. Choose Silero VAD for short segments, SpeechBrain VAD for general purpose, or WebRTC VAD for lightweight applications.
Voice Activity Detection (VAD)
Overview
Voice Activity Detection identifies which parts of an audio signal contain speech versus silence or background noise. This is a critical first step in speaker diarization pipelines.
When to Use
- Preprocessing audio before speaker diarization
- Filtering out silence and noise
- Segmenting audio into speech chunks
- Improving diarization accuracy by focusing on speech regions
Available VAD Tools
1. Silero VAD (Recommended for Short Segments)
Best for: Short audio segments, real-time applications, better detection of brief speech
import torch
# Load Silero VAD model
model, utils = torch.hub.load(
repo_or_dir='snakers4/silero-vad',
model='silero_vad',
force_reload=False,
onnx=False
)
get_speech_timestamps = utils[0]
# Run VAD
speech_timestamps = get_speech_timestamps(
waveform[0], # mono audio waveform
model,
threshold=0.6, # speech probability threshold
min_speech_duration_ms=350, # minimum speech segment length
min_silence_duration_ms=400, # minimum silence between segments
sampling_rate=sample_rate
)
# Convert to boundaries format
boundaries = [[ts['start'] / sample_rate, ts['end'] / sample_rate]
for ts in speech_timestamps]
Advantages:
- Better at detecting short speech segments
- Lower false alarm rate
- Optimized for real-time processing
2. SpeechBrain VAD
Best for: General-purpose VAD, longer audio files
from speechbrain.inference.VAD import VAD
VAD_model = VAD.from_hparams(
source="speechbrain/vad-crdnn-libriparty",
savedir="/tmp/speechbrain_vad"
)
# Get speech segments
boundaries = VAD_model.get_speech_segments(audio_path)
Advantages:
- Well-tested and reliable
- Good for longer audio files
- Part of comprehensive SpeechBrain toolkit
3. WebRTC VAD
Best for: Lightweight applications, real-time processing
import webrtcvad
vad = webrtcvad.Vad(2) # Aggressiveness: 0-3 (higher = more aggressive)
# Process audio frames (must be 10ms, 20ms, or 30ms)
is_speech = vad.is_speech(frame_bytes, sample_rate)
Advantages:
- Very lightweight
- Fast processing
- Good for real-time applications
Postprocessing VAD Boundaries
After VAD, you should postprocess boundaries to:
- Merge close segments
- Remove very short segments
- Smooth boundaries
def postprocess_boundaries(boundaries, min_dur=0.30, merge_gap=0.25):
"""
boundaries: list of [start_sec, end_sec]
min_dur: drop segments shorter than this (sec)
merge_gap: merge segments if silence gap <= this (sec)
"""
# Sort by start time
boundaries = sorted(boundaries, key=lambda x: x[0])
# Remove short segments
boundaries = [(s, e) for s, e in boundaries if (e - s) >= min_dur]
# Merge close segments
merged = [list(boundaries[0])]
for s, e in boundaries[1:]:
prev_s, prev_e = merged[-1]
if s - prev_e <= merge_gap:
merged[-1][1] = max(prev_e, e)
else:
merged.append([s, e])
return merged
Choosing the Right VAD
| Tool | Best For | Pros | Cons | |------|----------|------|------| | Silero VAD | Short segments, real-time | Better short-segment detection | Requires PyTorch | | SpeechBrain VAD | General purpose | Reliable, well-tested | May miss short segments | | WebRTC VAD | Lightweight apps | Fast, lightweight | Less accurate, requires specific frame sizes |
Common Issues and Solutions
- Too many false alarms: Increase threshold or min_speech_duration_ms
- Missing short segments: Use Silero VAD or decrease threshold
- Over-segmentation: Increase merge_gap in postprocessing
- Missing speech at boundaries: Decrease min_silence_duration_ms
Integration with Speaker Diarization
VAD boundaries are used to:
- Extract speech segments for speaker embedding extraction
- Filter out non-speech regions
- Improve clustering by focusing on actual speech
# After VAD, extract embeddings only for speech segments
for start, end in vad_boundaries:
segment_audio = waveform[:, int(start*sr):int(end*sr)]
embedding = speaker_model.encode_batch(segment_audio)
# ... continue with clustering
Related Skills
LocalAI
49.3kLocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
design
130.2kComprehensive design skill: brand identity, design tokens, UI styling, logo generation (55 styles, Gemini, Atlas Cloud, or MuAPI AI), corporate identity program (50 deliverables, CIP mockups), HTML presentations (Chart.js), banner design (22 styles, social/ads/web/print), icon design (15 styles, SVG…
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
