talking-head-recut
Package an existing talking-head / interview / podcast video with timed, designed GRAPHIC OVERLAY cards — kinetic titles, lower-thirds, data callouts, quotes, side panels, picture-in-picture — synced to the transcript, on a 16:9 / 9:16 / 4:5 canvas of your choice; the clip plays untouched underneath…
Install / Use
npx skills add heygen-com/hyperframes --skill talking-head-recutInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
DesignSupported Platforms
Our assessment of talking-head-recut
talking-head-recut scores 90/100 on our quality scale, 19th of 133 Design skills we index (top 15%).
Its SKILL.md is 64 KB long, well organised into 35 sections with 20 code examples: long enough that it reads more like full documentation than a focused instruction file, which agents can find harder to follow.
With 52,778 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated today, so talking-head-recut is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
talking-head-recut compared with similar skills
All 4 of these similar skills score higher than talking-head-recut; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| talking-head-recut (this skill)by heygen-com | 90 | 52.8k | today | SKILL.md |
| algorithmic-artby anthropics | 100 | 177.9k | 2d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 2d ago | SKILL.md |
| designby nextlevelbuilder | 100 | 130.2k | 3d ago | SKILL.md |
| ui-ux-pro-maxby nextlevelbuilder | 100 | 130.2k | 3d ago | SKILL.md |
Frequently asked questions
- How do I install talking-head-recut?
- Run
npx skills add heygen-com/hyperframes --skill talking-head-recut. The install tabs above show the steps for each supported agent. - Which AI agents does talking-head-recut work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is talking-head-recut safe to use?
- It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is talking-head-recut still maintained?
- The repository was last updated today, so talking-head-recut is actively maintained.
Skill content
View source on GitHubname: talking-head-recut description: Package an existing talking-head / interview / podcast video with timed, designed GRAPHIC OVERLAY cards — kinetic titles, lower-thirds, data callouts, quotes, side panels, picture-in-picture — synced to the transcript, on a 16:9 / 9:16 / 4:5 canvas of your choice; the clip plays untouched underneath. Trigger on "graphic overlays", "on-screen graphics", "package / dress up my video". Not plain subtitles (/embedded-captions). Unclear → /hyperframes.
First, keep this skill fresh — confirm with the user before running:
npx hyperframes skills update talking-head-recut. A fast no-op when everything is current; otherwise it refreshes this skill plus the core domain skills it depends on before you rely on them.
Talking Head Recut
Talking Head Recut takes a local video that plays in full and layers a sequence of
timed, designed graphic cards onto it — titles, lower-thirds, data callouts,
quotes, side panels, picture-in-picture — synced to what's being said. The agent
designs the cards (timing + content) and writes each card's HTML directly in the
conversation, then assembles a single composition HTML and renders it to MP4 via
hyperframes. There is no fixed archetype list and no prescribed card structure —
the overlays emerge from what the transcript actually says.
The front door is
/hyperframes. This skill packages an existing talking-head clip with designed graphic cards (titles, lower-thirds, data callouts, quotes, side panels, PiP) — not plain captions (the spoken words as text). The clip plays untouched. Any other intent — plain subtitles, a standalone graphic, a from-scratch video — or any uncertainty → read/hyperframesfirst: the intent layer owns every route decision.
Graphic-packaging sibling of
embedded-captions. Captions add the spoken words as a readable subtitle; this adds designed graphics on top of the playing video. Plain subtitles →embedded-captions. Build a video from scratch → the creation workflows (product-launch-video/faceless-explainer/ …).
Routed through /hyperframes, the intent layer confirms only the input (which clip) and announces the render-strategy questions as deferred asks — aspect, layout, style group, and card count stay at Step 7, where the probed footage and transcript ground the recommendations; the layer's run-shape questions don't apply. A BRIEF.md, when present, carries the confirmed input and any user notes — read it first.
Inspectable intermediate files in the work directory:
metadata.json— duration / width / height / fpsaudio.mp3— extracted audiotranscript.json— a flat word array[{ text, start, end }, …](Whisper; nosegments, nowordswrapper)storyboard.json— lightweight card outline (the agent's plan)public/cards/card-XX.html— one HTML fragment per cardpublic/index.html— final assembled compositionoutput.mp4— rendered video
CLI Resolution
# hyperframes — transcription (local Whisper) + rendering the assembled HTML to MP4
npx hyperframes --help
This skill runs entirely on the hyperframes CLI plus system ffmpeg / ffprobe.
Transcription is local Whisper via hyperframes transcribe — no third-party
service, API key, or rate-limited proxy.
Workflow
1. Check Environment
npx hyperframes doctor # ffmpeg, headless browser, render deps
# confirm bundled assets:
ls "<SKILL_DIR>/assets/fonts" "<SKILL_DIR>/assets/vendor/gsap.min.js"
Required:
ffmpeg/ffprobe(system)<SKILL_DIR>/assets/fonts/*.woff2,<SKILL_DIR>/assets/vendor/gsap.min.js(bundled inside this skill, staged to work dir in Step 9)
Transcription needs no key — hyperframes transcribe runs Whisper locally (Step 4).
Strongly recommended on macOS for hyperframes render:
export PRODUCER_BROWSER_GPU_MODE=hardware
2. Create a Work Directory
All artifacts live under videos/<project-name>/ — the same convention as the other
video workflows (product-launch-video / faceless-explainer / pr-to-video). Keep
the cwd at the workspace root; everything below writes under this one subdirectory.
VIDEO_PATH="/absolute/path/input.mp4"
WORK_DIR="videos/$(basename "$VIDEO_PATH" | sed 's/\.[^.]*$//')"
mkdir -p "$WORK_DIR"
3. Extract Audio and Metadata
# metadata — duration / width / height / fps
ffprobe -v error -select_streams v:0 \
-show_entries stream=width,height,r_frame_rate \
-show_entries format=duration -of json "$VIDEO_PATH" > "$WORK_DIR/metadata.json"
# audio
ffmpeg -y -i "$VIDEO_PATH" -vn -acodec libmp3lame -q:a 2 "$WORK_DIR/audio.mp3"
Outputs: metadata.json (read width/height/duration; fps = the r_frame_rate
fraction evaluated, e.g. 30000/1001 → 29.97) + audio.mp3.
4. Transcribe
npx hyperframes transcribe "$WORK_DIR/audio.mp3" -d "$WORK_DIR" --json --model small.en
Local Whisper — no API key, no proxy, no rate limit. Writes a word-level
transcript.json into the work dir (word text + start / end timestamps).
Read it for the word / sentence timings that drive card timing in Step 6; group
words into sentences yourself at punctuation / pauses if you need segment-level
chunks.
Clamp to media duration. Whisper can return the final word's end a hair past the
actual clip length — clamp every card endSec and composition.durationSeconds to the
metadata.json duration, or the render will show a black tail past the video.
5. Correct Transcript
transcript.json is a flat array of word objects — [{ "text": "...", "start": s, "end": s }, …] (no segments array, no words wrapper; the per-word key is text). Read it and fix obvious ASR errors:
- Homophones, product names, technical terms, punctuation
- Edit a word's
textin place; preserve itsstart/endtimestamps - There is no pre-grouped
segmentsarray — group words into sentences yourself (split at terminal punctuation / pauses) when you need segment-level chunks for card timing
6. Draft a Lightweight Storyboard (in chat)
No CLI involved. Read transcript.json + metadata.json and design
cards directly. storyboard.json is an agent-internal planning artifact
— no CLI command consumes it; it exists so you can think clearly
about timing and content before writing each card's HTML. Keep the
shape consistent with the example below so the same outline can drive
the composition you author in Step 9:
{
"schemaVersion": 3,
"composition": {
"fps": 30,
"width": 1080,
"height": 1920,
"durationSeconds": 121.2,
"layout": "portrait",
"themeId": "noir",
"seed": 42
},
"videoTrack": {
"sourcePath": "input-video.mp4",
"startSec": 0,
"endSec": 121.2,
"bounds": { "x": 0, "y": 0, "width": 1080, "height": 1920 }
},
"subtitles": { "enabled": false },
"cards": [
{
"id": "card-01",
"intent": "Hook with the speaker's anxious midnight question",
"startSec": 0.5,
"endSec": 13.0,
"accentIndex": 0,
"zone": "fullscreen",
"contentHints": {
"kicker": "AN HONEST QUESTION",
"title": "The soul-searching question at 11 PM",
"detail": "Client's 60-second voice message: 'If the RMB appreciates, does that mean my USD policy is a terrible loss?'"
}
}
]
}
Required Card fields:
| field | type | purpose |
| ----------------------- | ------------------------------------------ | ----------------------------------------------------------------------------------------------------- |
| id | string | stable id used in card HTML & GSAP selectors |
| intent | string | natural-language description; fed to card synthesis |
| startSec / endSec | number | times in seconds (endSec > startSec) |
| accentIndex | 0 | 1 | 2 | 3 | 4 | which of the 5 theme accent colors this card pulls |
| zone | enum (see below) | where on the canvas the card lives |
| contentHints | object | free-form bag; agent puts kicker/title/detail/data/quote here |
| archetype (optional) | string | free-form label you may attach to remember a card's pattern; absent = free-form, which is the default |
| transition (optional) | enum: cut | fade | slide | wipe | declarative card-to-card transition |
Five zone values:
| zone | resolved bounds | when to use |
| ----------------- | ---------------------------------------------- | --------------------------------------- |
| fullscreen | covers whole canvas | hero moments, big numbers, mantras |
| whiteboard-area | inset 40px margin (or 45% of portrait height) | dense data / annotated content |
| lower-third | bottom 30% band | annotation over visible video |
| side-panel | right 42% (landscape) or bottom 40% (portrait) | data side, video other side |
| video-overlay | full canvas, expects mostly-transparent card | annotation overlays on full-bleed video |
When you assemble the composition in Step 9, resolve each card's zone
into pixel bounds on the card-host wrapper following the table above.
Video bounds are set once at composition level (videoTrack.bounds);
to make video appear to "move between cards", author GSAP tweens against
#video-wrap in the composition's <script> (see Step 9).
No prescribed card roles, no prescribed narrative arc. Cards emerge from what the video actually says — could be all quotes or all data, could open with a number or with a story. Let the transcript drive the rhythm.
How many takeaways? — auto-infer from duration + density. No fixed upper limit. Pick a base pace from the video duration, then adjust by information density. Only floor is fixed: minimum 5 cards so even short videos have rhythm.
Step 1 — base pace by duration (the natural sec/card for medium density):
| video duration | base pace (sec per card) | rationale | | ------------------ | ------------------------ | ------------------------------------------- | | < 60s (short reel) | 6–8s | viewers expect fast cuts in short-form | | 60s – 3 min | 8–12s | normal social pace | | 3 – 10 min | 12–20s | give breathing room; each card carries more | | 10 – 30 min | 20–35s | long-form lecture / interview rhythm | | > 30 min | 30–60s | episodic, near-chapter feel |
Step 2 — density multiplier (multiplies the base pace):
| signal in the transcript | multiplier | effect | | --------------------------------------------------------------------------------------------------------------------------- | --------
Truncated for display — read the full file on GitHub.
Related Skills
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
design
130.2kComprehensive design skill: brand identity, design tokens, UI styling, logo generation (55 styles, Gemini, Atlas Cloud, or MuAPI AI), corporate identity program (50 deliverables, CIP mockups), HTML presentations (Chart.js), banner design (22 styles, social/ads/web/print), icon design (15 styles, SVG…
ui-ux-pro-max
130.2kUI/UX design intelligence for web, mobile, and desktop. This skill should be used when designing, building, reviewing, or fixing interfaces, including pages, components, design systems, accessibility, interaction, responsive layout, typography, color, charts, and stack-specific UI implementation.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
