SkillAgentSearch skills...

Video Podcast Maker

Topic → 4K narrated video for coding agents. v4.0: all TTS via the ttsCN engine component (11 platforms incl. MiniMax voice clone, native word-level subtitle sync), manifest-based Asset Engine, Remotion composition, cost-gated AI generation

Install / Use

npx skills add Agents365-ai/video-podcast-maker

Installs into whichever agent you are using.

About this skill

Quality Score

0/100

Supported Platforms

Claude Code
Claude Desktop

README

Video Podcast Maker

License: CC BY-NC 4.0 GitHub stars GitHub forks Latest Release Last Commit

SkillsMP Claude Code Plugin Agent Skills

中文文档

Automated pipeline to create professional video podcasts from a topic. Supports Bilibili, YouTube, Xiaohongshu, Douyin, and WeChat Channels with multi-language output (zh-CN, en-US). Combines research, script generation, multi-engine TTS (11 backends via the ttscn bridge), Remotion rendering, and FFmpeg mixing. Current release: v5.2.1 — see CHANGELOG.md for version history.

Works with: Claude Code · OpenClaw · OpenCode · Codex · Pi — any coding agent that supports SKILL.md

Publish to: Bilibili · YouTube · Xiaohongshu · Douyin · WeChat Channels

No coding required! Just describe your topic in plain language — the coding agent guides you through each step interactively. You make creative decisions, the agent handles all the technical details.

Note: This project is still under active development and may not be fully mature yet. Your feedback is greatly appreciated — feel free to open an issue.

Features

  • Topic → 4K video - research, narration script, TTS audio, Remotion composition, 4K render + BGM in one pipeline
  • 11 TTS backends - Edge (free), Azure, CosyVoice, Doubao, Tencent, Baidu, MiniMax, Xunfei, ElevenLabs, OpenAI, Google — all synthesized by the required ttscn component skill
  • Asset engine - per-video manifest with license provenance; producers are user files, assetseeker stock, imagencn AI stills, videogencn AI B-roll, and Hyperframes overlays — paid generation always asks first
  • 4K output + Remotion-native subtitles - 3840×2160; SRT rendered in React at 4K (legacy FFmpeg burn-in available)
  • Design learning - extract style profiles from reference videos/images; auto-applied when topics match
  • Vertical shorts - 9:16 highlight clips generated from long-form sections
  • Multi-platform & multi-language - Bilibili / YouTube / Xiaohongshu / Douyin / WeChat Channels × zh-CN / en-US, with per-platform publish info
  • Pronunciation control - global + per-project phoneme dictionaries for Chinese polyphones

Quick Start

1. Install — via the 365-skills marketplace (recommended) or by cloning this repo.

2. Set up — Python 3.8+, Node.js 18+, FFmpeg, and a Remotion project:

brew install ffmpeg node python3          # macOS (Ubuntu: sudo apt install ffmpeg nodejs python3)
pip install -r skills/video-podcast-maker/requirements.txt
npx create-video@latest my-video-project   # or reuse an existing Remotion project
cd my-video-project && npm i

3. Configure — set TTS_BACKEND plus its API keys (see TTS Backends and Environment Variables).

4. Tell your agent:

"Create a video podcast about [your topic]"

The agent runs the whole workflow (research → script → TTS → Remotion composition → Studio review → 4K render + BGM). Preview and iterate in Remotion Studio (npx remotion studio src/remotion/index.ts); the agent waits for your explicit "render 4K" confirmation before the final render.

⚠️ For the human reading this (not the AI): manually polish podcast.txt, repeatedly

This section is for you, the human — not the agent. Every downstream step — TTS narration, subtitles, section transitions, animation timing, final cut — is derived from this single podcast.txt. A weak script renders into 4K garbage. No amount of polish downstream saves it.

The AI-generated draft is a starting point, nothing more. Do these yourself — don't hand them off to the AI:

  1. Mentally read it as the narrator. Treat each sentence as one breath — if a line forces you to "catch your breath" or backtrack to parse, fix it. Where you stumble silently is where TTS stumbles audibly.
  2. Revise at least three times.
    • Pass 1: typos, awkward phrasing, tongue-twisters
    • Pass 2: cut filler, cut throat-clearing intros ("So today we're going to talk about…"), cut redundancy
    • Pass 3: tune rhythm — where to pause, where to break a long sentence, which word carries the stress
  3. Read each [SECTION:xxx] block end-to-end. Confirm each section opens with a hook and lands a clean transition into the next — not a bullet-point dump.
  4. Audit numbers, proper nouns, and English terms separately. ~90% of TTS mispronunciations live here. If pronunciation is wrong, add it to phonemes.json; if it just sounds awkward, rewrite it.
  5. Know your length budget. Estimate ~280 zh-CN chars/min or ~150 en words/min. A 5–10 min video means ~1400–2800 chars / 750–1500 words. Don't pad to fill time.

The only acceptance test: read through it once in your head — does any line make you wince? If yes, don't move on to Step 7 (TTS) yet. Otherwise you're just rendering 4K of something even you don't want to hear.

Workflow

Pipeline

Component Skills

Asset Flow

Related Skills

  • remotion-best-practices - required; core Remotion patterns and guidelines
  • ttscn - required; the TTS engine behind all 11 backends (install under ~/.claude/skills/ttscn, as a Pi skill, or set TTSCN_HOME)
  • assetseeker - optional; license-vetted stock photos/video/BGM/SFX/icons/fonts
  • imagencn - optional; AI stills and thumbnails (paid APIs)
  • videogencn - optional; AI video clips for B-roll (paid APIs)
  • Hyperframes - optional; transparent overlay animations (Node 22+)

Requirements

| Software | Version | Purpose | | ---------- | --------- | --------- | | macOS / Linux | - | Tested on macOS, Linux compatible | | Python | 3.8+ | TTS script, automation | | Node.js | 18+ | Remotion video rendering | | FFmpeg | 4.0+ | Audio/video processing |

Marketplace install (recommended): users typically install this skill via the 365-skills marketplace rather than cloning. SKILL.md, scripts, and templates then live under the agent's ${SKILL_DIR}; paths in this README are written from the repo-root perspective for contributors.

TTS Backends (all via ttscn)

All 11 platforms are synthesized by the required ttscn component skill. Set TTS_BACKEND to any platform id; only the active platform's env vars are needed:

| TTS_BACKEND | Provider | Required env vars | Get Key | | --------------- | ---------- | ------------------- | --------- | | edge (default) | Microsoft Edge TTS | (none — free) | — | | azure | Microsoft Azure Speech | AZURE_SPEECH_KEY (+ optional AZURE_SPEECH_REGION, default eastasia) | Azure Portal | | cosyvoice | Aliyun CosyVoice | DASHSCOPE_API_KEY | Aliyun Bailian | | doubao | Volcengine Doubao | VOLCENGINE_APPID, VOLCENGINE_ACCESS_TOKEN | Volcengine Console | | tencent | Tencent Cloud TTS | TENCENT_SECRET_ID, TENCENT_SECRET_KEY | Tencent Console | | baidu | Baidu AI TTS | BAIDU_APP_ID, BAIDU_API_KEY, BAIDU_SECRET_KEY | Baidu Console | | minimax | MiniMax TTS | MINIMAX_API_KEY | MiniMax Platform | | xunfei | iFlytek Xunfei TTS | XUNFEI_APP_ID, XUNFEI_API_KEY, XUNFEI_API_SECRET | Xfyun | | elevenlabs | ElevenLabs | ELEVENLABS_API_KEY | ElevenLabs | | openai | OpenAI TTS | OPENAI_API_KEY | OpenAI Platform | | google | Google Cloud TTS | GOOGLE_TTS_API_KEY | Google Cloud Console |

Non-TTS keys (optional): GEMINI_API_KEY / DASHSCOPE_API_KEY for AI thumbnails (imagencn).

Environment Variables

Add to ~/.zshrc or ~/.bashrc:

export TTS_BACKEND="edge"                  # azure / cosyvoice / doubao / tencent / baidu / minimax / xunfei / elevenlabs / openai / google
export TTS_VOICE="zh-CN-XiaoxiaoNeural"    # optional; unset = platform default
export TTS_RATE="+5%"                      # optional; also settable in user_prefs.json (g

Related Skills

View on GitHub
GitHub Stars1.5k
CategoryDevelopment
Updated14h ago
Forks159

Languages

Python

Security Score

85/100

Audited on Aug 7, 2026

No findings