SkillAgentSearch skills...

Youtube Automation Agent

🎬 Fully automated YouTube channel management with AI agents. Creates, optimizes & publishes videos 24/7. Works with FREE Gemini API or OpenAI. No coding required!

Install / Use

npx skills add darkzOGx/youtube-automation-agent

Installs into whichever agent you are using.

README

YouTube Automation Agent

What's New in v2.4

  • Guided walkthrough for first-time setup β€” npm run walkthrough (also offered when you run npm run setup). It explains every choice in plain English, shows exactly where to get each key (and opens the page in your browser), live-tests keys the moment you paste them, walks you click-by-click through Google Cloud for the YouTube connection, and signs you in via your browser instead of copy-pasting auth codes. Every step is skippable and progress is saved β€” re-run it any time.
  • .env files actually work now β€” dotenv was a dependency but was never loaded, so .env settings (API keys, API_KEY, FFMPEG_PATH…) were silently ignored unless exported in your shell. index.js and the setup tools now load .env on start.
  • .env.example no longer poisons setup β€” the uncommented OPENAI_API_KEY=your-openai-api-key-here placeholder would have been picked up as a real key; all placeholders are now commented out.
  • Browser OAuth opens automatically β€” the YouTube authorization URL now opens in your default browser.

What's New in v2.3

  • Full free-tier pipeline with Gemini β€” image generation (gemini-3.1-flash-image) and native voice narration (gemini-3.1-flash-tts-preview) now run on your Gemini key. A Gemini-only setup produces complete narrated videos end to end; OpenAI/ElevenLabs are used first when configured. Models and voice are configurable via GEMINI_IMAGE_MODEL, GEMINI_TTS_MODEL, GEMINI_TTS_VOICE. (Thanks to PR #6 for demonstrating the demand and fallback-chain direction.)
  • ~50Γ— faster slideshow rendering β€” instead of screenshotting a headless browser at 30fps (~10 minutes for a 30-second video), the renderer captures one still per slide and lets FFmpeg build the video with crossfades (seconds).
  • No more junk template topics β€” template mode (no AI key) previously scraped single keywords from trending titles and produced videos like "crown: The Complete Guide". It now uses a curated evergreen topic list and only accepts trending topics that read like real subjects.
  • Model catalog corrections β€” replaced the nonexistent gemini-3.5-pro picker entry with gemini-3.1-pro-preview / gemini-2.5-pro (verified against Google's current model list).

What's New in v2.2

This release resolves every open GitHub issue (#1, #2, #3, #4, #8, #9, #13):

  • Gemini (and every other provider) now passes credential validation β€” startup and setup no longer demand an OpenAI key. Any one configured AI provider (OpenAI, Gemini, OpenRouter, Kimi, MiMo, or GLM) is enough. (#3, #9)
  • FFmpeg is bundled β€” npm install now pulls a prebuilt FFmpeg binary via ffmpeg-static, so 'ffmpeg' is not recognized is gone. A system install on your PATH or FFMPEG_PATH in .env still takes precedence. (#1)
  • Generated content actually reaches the publish queue β€” the /generate pipeline previously produced a video and then never scheduled it, so "Processing publish queue" ran forever with nothing to do. It now queues every successful production. (#2)
  • Real .mp4 output without paid keys β€” if TTS isn't configured, the slideshow renders as a silent video instead of dying on a placeholder file. Placeholder .info assets are filtered out of slides. (#4)
  • No more silent failures β€” a capability check at startup shows exactly which pipeline stages will run for real (βœ“) vs. what's missing and how to fix it (βœ—). Productions that only produced placeholders are marked simulated, are never scheduled for upload, and log a loud warning. (#4, #8, #13)
  • Setup wizard no longer hard-aborts β€” missing credentials or FFmpeg produce warnings with fix instructions instead of ❌ Setup failed!. (#9)
  • Publish-queue logging is informative β€” shows how many items are waiting and when the next publish happens, instead of an identical line every 15 minutes. (#2)

What's New in v2.1

  • Real AI generation wired in β€” the Content Strategy, Script Writer, and SEO agents now call your configured AI provider (OpenAI, OpenRouter, Kimi, MiMo, GLM, or Gemini) for topics, scripts, titles, descriptions, and tags. If no provider key is set, they fall back to the built-in templates so the pipeline still runs.
  • API protection β€” set API_KEY in .env and the mutating endpoints (POST /generate, POST /publish/:id) require a matching x-api-key header. Request bodies are validated and size-limited.
  • Safer publishing β€” default privacy is now private (set DEFAULT_PRIVACY_STATUS=public to opt in), and the uploader streams the real video file β€” it refuses to upload placeholder assets from simulated runs.
  • Startup and scheduler fixes β€” added the missing sharp dependency (the app previously crashed on boot), created the missing automation_events table (every scheduled task previously threw on logging), fixed the double-insert in the content pipeline, and fixed the publish-queue removal.
  • No more fabricated statistics β€” template scripts no longer invent numbers like "90% of people…".
  • Cleaner repo β€” removed two dead OAuth flows (authenticate.js, simple-auth.js used Google's long-deprecated OOB flow), dead dependencies (cron, jimp), broken npm scripts, and committed build artifacts. Added ESLint (npm run lint) and GitHub Actions CI.

What's New in v2.0

  • Model upgrades across the board β€” GPT-5.5 / GPT-5.5 Instant replace GPT-4-turbo, GPT Image 2 replaces DALL-E 3, Gemini 3.5 Flash/Pro replace Gemini 1.x, ElevenLabs Eleven v3 replaces v1, Wan 2.7 replaces Stable Video Diffusion
  • OpenAI SDK v6 β€” upgraded from v4, along with @google/genai v2.9, replicate v1.4, googleapis v173
  • Revamped setup wizard β€” new TTS service picker (OpenAI TTS / ElevenLabs / Azure), ElevenLabs credential setup, updated model selection menus
  • Fixed deprecated API patterns β€” OpenAI v3 SDK calls in credential testing replaced with v4+ patterns
  • Dynamic year in content strategy β€” no more hardcoded "2025" in trend analysis prompts
  • README rewrite β€” developer-focused docs with Mermaid architecture diagrams, no fluff

Fully automated YouTube channel management system. AI agents handle content strategy, scriptwriting, thumbnail generation, SEO, publishing, and analytics β€” end to end, on a daily schedule.

Built by

@darkzOGx. Solo builder shipping AI automation and developer tools.

Find me on X and laderalabs.io.

If this saves you time, a star helps it reach more developers.

Architecture

graph TD
    A[Content Strategy Agent] --> B[Script Writer Agent]
    B --> C[Thumbnail Designer Agent]
    B --> D[SEO Optimizer Agent]
    C --> E[Production Management Agent]
    D --> E
    E --> F[Publishing & Scheduling Agent]
    F --> G[Analytics & Optimization Agent]
    G -->|feedback loop| A

How It Works

Each agent handles one stage of the pipeline:

| Agent | Role | |-------|------| | Content Strategy | Analyzes YouTube trends, identifies topics, plans content calendar | | Script Writer | Generates scripts with hooks, storytelling, CTAs | | Thumbnail Designer | Creates thumbnails, runs A/B variations | | SEO Optimizer | Keywords, titles, descriptions, tags | | Production | Coordinates TTS audio, image assets, video assembly | | Publishing | Uploads, schedules, manages playlists | | Analytics | Tracks performance, feeds insights back to strategy |

AI Providers

All OpenAI-compatible providers work out of the box β€” the system auto-configures the SDK base URL. Pick one, or use OpenRouter to access everything through a single key.

graph LR
    subgraph Direct
        OA[OpenAI<br/>GPT-5.5]
        GM[Gemini<br/>3.5 Flash/Pro]
        KM[Kimi<br/>K2.6]
        MM[MiMo<br/>V2.5 Pro]
        GL[GLM<br/>GLM-5]
    end
    subgraph Router
        OR[OpenRouter<br/>300+ models]
    end
    Direct --> YAA[YouTube Automation Agent]
    Router --> YAA

| Provider | Models | Base URL | Cost | |----------|--------|----------|------| | OpenAI | GPT-5.5, GPT-5.5 Instant | api.openai.com/v1 | ~$0.05–0.20/video | | OpenRouter | 300+ (GPT, Claude, Gemini, Kimi, GLM, etc.) | openrouter.ai/api/v1 | varies by model | | Google Gemini | Gemini 3.5 Flash, 3.5 Pro | via @google/genai SDK | free tier available | | Kimi (Moonshot AI) | Kimi K2.6, K2.5 | api.moonshot.ai/v1 | ~80% cheaper than GPT-5.5 | | MiMo (Xiaomi) | MiMo V2.5 Pro, V2.5 | api.xiaomimimo.com/v1 | competitive | | GLM (Zhipu AI) | GLM-5, GLM-5.1 | api.z.ai/api/paas/v4/ | ~$1/M input tokens |

Additional integrations: Anthropic Claude (claude-opus-4-8), ElevenLabs (Eleven v3 TTS), Replicate (Wan 2.7 video), local models via Ollama, any OpenAI-compatible endpoint.

Quick Start

git clone https://github.com/darkzOGx/youtube-automation-agent.git
cd youtube-automation-agent
npm install
npm run walkthrough   # guided first-time setup: explains everything, tests your keys live
npm start

Dashboard runs at http://localhost:3456.

Already know what you're doing? npm run setup offers a classic quick mode, and .env.example documents every setting.

Prerequisites

  • Node.js 18+
  • FFmpeg β€” bundled automatically via ffmpeg-static on npm install; a system install on your PATH or an FFMPEG_PATH env var takes precedence
  • Google account (YouTube Data API β€” free)
  • At least one AI provider key (OpenAI, Gemini, OpenRouter, Kimi, MiMo, or GLM) β€” without one, agents fall back to template-based generation
  • Images and narration come from your AI key: OpenAI or Gemini both cover image generation and TTS (ElevenLabs / Azure Speech optional for premium voices) β€” with no media provider at all you get gradient slides and silent video

Configuration

API Keys

YouTube Data API (required, free)

  1. Create a project in [Google Cloud Console](https://

Related Skills

View on GitHub
GitHub Stars1.8k
CategoryMarketing
Updated2h ago
Forks431

Languages

JavaScript

Security Score

100/100

Audited on Aug 8, 2026

No findings