SkillAgentSearch skills...

filler-word-processing

Process filler word annotations to generate video edit lists

Install / Use

npx skills add benchflow-ai/skillsbench --skill filler-word-processing

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

86/100

Supported Platforms

Universal

Tags

Our assessment of filler-word-processing

filler-word-processing scores 86/100 on our quality scale, 1550th of 3,997 Development & Engineering skills we index (top 39%).

Its SKILL.md is 4.1 KB long, well organised into 9 sections with 4 code examples: a solid amount of guidance for an agent.

With 1,813 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
26/30
Structure
20/20
Description
12/15
Adoption
14/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated about 2 months ago, so filler-word-processing is actively maintained.
  • It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

filler-word-processing compared with similar skills

All 4 of these similar skills score higher than filler-word-processing; compare them before choosing.

SkillScoreStarsUpdatedFormat
filler-word-processing (this skill)by benchflow-ai861.8k2mo agoSKILL.md
ai-job-searchby MadsLorentzen10044.6ktodayCLAUDE.md
claude-howtoby luongnv8910041.7ktodayCLAUDE.md
algorithmic-artby anthropics100177.9k7d agoSKILL.md
pptxby anthropics100177.9k7d agoSKILL.md

Frequently asked questions

How do I install filler-word-processing?
Run npx skills add benchflow-ai/skillsbench --skill filler-word-processing. The install tabs above show the steps for each supported agent.
Which AI agents does filler-word-processing work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is filler-word-processing safe to use?
It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is filler-word-processing still maintained?
The repository was last updated about 2 months ago, so filler-word-processing is actively maintained.

name: filler-word-processing description: "Process filler word annotations to generate video edit lists. Use when working with timestamp annotations for removing speech disfluencies (um, uh, like, you know) from audio/video content."

Filler Word Processing

Annotation Format

Typical annotation JSON structure:

[
  {"word": "um", "timestamp": 12.5},
  {"word": "like", "timestamp": 25.3},
  {"word": "you know", "timestamp": 45.8}
]

Converting Annotations to Cut Segments

Each filler word annotation marks when the word starts. To remove it, use word-specific durations since different fillers have different lengths:

import json

# Word-specific durations (in seconds)
WORD_DURATIONS = {
    "uh": 0.3,
    "um": 0.4,
    "hum": 0.6,
    "hmm": 0.6,
    "mhm": 0.55,
    "like": 0.3,
    "yeah": 0.35,
    "so": 0.25,
    "well": 0.35,
    "okay": 0.4,
    "basically": 0.55,
    "you know": 0.55,
    "i mean": 0.5,
    "kind of": 0.5,
    "i guess": 0.5,
}
DEFAULT_DURATION = 0.4

def annotations_to_segments(annotations_file, buffer=0.05):
    """
    Convert filler word annotations to (start, end) cut segments.

    Args:
        annotations_file: Path to JSON annotations
        buffer: Small buffer before the word (seconds)

    Returns:
        List of (start, end) tuples representing segments to remove
    """
    with open(annotations_file) as f:
        annotations = json.load(f)

    segments = []
    for ann in annotations:
        word = ann.get('word', '').lower().strip()
        timestamp = ann['timestamp']
        # Use word-specific duration, fall back to default
        word_duration = WORD_DURATIONS.get(word, DEFAULT_DURATION)
        # Cut starts slightly before the word
        start = max(0, timestamp - buffer)
        # Cut ends after word duration
        end = timestamp + word_duration
        segments.append((start, end))

    return segments

Merging Overlapping Segments

When filler words are close together, merge their cut segments:

def merge_overlapping_segments(segments, min_gap=0.1):
    """
    Merge segments that overlap or are very close together.

    Args:
        segments: List of (start, end) tuples
        min_gap: Minimum gap to keep segments separate

    Returns:
        Merged list of segments
    """
    if not segments:
        return []

    # Sort by start time
    sorted_segs = sorted(segments)
    merged = [sorted_segs[0]]

    for start, end in sorted_segs[1:]:
        prev_start, prev_end = merged[-1]

        # If this segment overlaps or is very close to previous
        if start <= prev_end + min_gap:
            # Extend the previous segment
            merged[-1] = (prev_start, max(prev_end, end))
        else:
            merged.append((start, end))

    return merged

Complete Processing Pipeline

def process_filler_annotations(annotations_file, word_duration=0.4):
    """Full pipeline: load annotations -> create segments -> merge overlaps"""

    # Load and create initial segments
    segments = annotations_to_segments(annotations_file, word_duration)

    # Merge overlapping cuts
    merged = merge_overlapping_segments(segments)

    return merged

Tuning Parameters

| Parameter | Typical Value | Notes | |-----------|---------------|-------| | word_duration | varies | Short fillers (um, uh) ~0.25-0.3s, single words (like, yeah) ~0.3-0.4s, phrases (you know, i mean) ~0.5-0.6s | | buffer | 0.05s | Small buffer captures word onset | | min_gap | 0.1s | Prevents micro-segments between close fillers |

Word Duration Guidelines

| Category | Words | Duration | |----------|-------|----------| | Quick hesitations | uh, um | 0.3-0.4s | | Sustained hums (drawn out while thinking) | hum, hmm, mhm | 0.55-0.6s | | Quick single words | like, yeah, so, well | 0.25-0.35s | | Longer single words | okay, basically | 0.4-0.55s | | Multi-word phrases | you know, i mean, kind of, i guess | 0.5-0.55s |

Quality Considerations

  • Too aggressive: Cuts into adjacent words, sounds choppy
  • Too conservative: Filler words partially audible
  • Sweet spot: Clean cuts with natural-sounding result

Test with a few samples before processing full video.

Related Skills

View on GitHub
GitHub Stars1.8k
CategoryDevelopment
Updated2mo ago
Forks368

Languages

PDDL

Trust signals

100/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

No cautions