SkillAgentSearch skills...

video-fetcher-to-markdown

Portable AI-agent skill: turn YouTube, Instagram, TikTok, X and other video links into Obsidian-ready Markdown notes: captions or local Whisper transcripts, frames, metadata and timestamps

Install / Use

npx skills add JimmySadek/video-fetcher-to-markdown

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

91/100

Supported Platforms

Claude Code
OpenAI Codex

Tags

Our assessment of video-fetcher-to-markdown

video-fetcher-to-markdown scores 91/100 on our quality scale, 329th of 1,174 Content & Media skills we index (top 29%).

Its SKILL.md is 20 KB long, well organised into 34 sections with 16 code examples: a thorough specification that gives an agent plenty to work with.

It has 485 GitHub stars, a meaningful sign that others use it.

Substance
30/30
Structure
20/20
Description
15/15
Adoption
11/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated today, so video-fetcher-to-markdown is actively maintained.
  • It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

Safety scan

No issues found

Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. An AI review of the same text found nothing harmful.

AI review by kimi-k2.7-code on 2026-10-08. Automated pattern scan on 2026-10-08. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.

video-fetcher-to-markdown compared with similar skills

All 4 of these similar skills score higher than video-fetcher-to-markdown; compare them before choosing.

SkillScoreStarsUpdatedFormat
video-fetcher-to-markdown (this skill)by JimmySadek91485todaySKILL.md
siyuanby siyuan-note10046.7ktodayMCP Server
algorithmic-artby anthropics100177.9k15d agoSKILL.md
pptxby anthropics100177.9k15d agoSKILL.md
designby nextlevelbuilder100133.6k4d agoSKILL.md

Frequently asked questions

How do I install video-fetcher-to-markdown?
Run npx skills add JimmySadek/video-fetcher-to-markdown. The install tabs above show the steps for each supported agent.
Which AI agents does video-fetcher-to-markdown work with?
It is written for Claude Code and OpenAI Codex, as a SKILL.md file. Other agents that read the same format can often use it too.
Is video-fetcher-to-markdown safe to use?
Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. An AI review of the same text found nothing harmful. It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is video-fetcher-to-markdown still maintained?
The repository was last updated today, so video-fetcher-to-markdown is actively maintained.

Video Fetcher to Markdown

<p align="center"> <img src="assets/banner.png" alt="Video Fetcher to Markdown: YouTube, TikTok and Instagram links into Markdown notes with frames" width="100%"> </p>

A video link in, a structured archival Markdown note out. Capture the transcript, creator metadata, description, chapters, actual language, and provenance in one Obsidian-ready file, without an API key. YouTube captions are read directly; Instagram, TikTok, X, Vimeo, Facebook and other sites are transcribed on your own machine with Whisper, with a contact sheet of frames for short videos.

npx skills add JimmySadek/video-fetcher-to-markdown

Read the v2.0.0 release notes for other video sites, local Whisper transcription, frames, and the login-wall fallback.

Formerly YouTube Fetcher to Markdown. Existing installs keep working and keep updating: the skill is still named youtube-fetcher, and GitHub redirects the old address. npx skills update brings you the latest version.

An independent open-source tool, not affiliated with or endorsed by YouTube, Google, or any other video platform it reads.

What you get

Paste a YouTube link and receive a file such as:

~/yt_transcripts/2026-03-04_obsidian-the-king-of-learning-tools_[hSTy_BInQs8].md
---
title: "Obsidian: The King of Learning Tools (FULL GUIDE + SETUP)"
channel: "Odysseas"
url: "https://www.youtube.com/watch?v=hSTy_BInQs8"
video_id: "hSTy_BInQs8"
fetched: "2026-03-04"
source_project: "my-project"
language: "en"
caption_type: "manual"
duration: "36m 26s"
upload_date: "2024-04-24"
tags:
  - yt-transcript
---

# Obsidian: The King of Learning Tools (FULL GUIDE + SETUP)

## Video Details
| Field    | Value |
|----------|-------|
| URL      | https://www.youtube.com/watch?v=hSTy_BInQs8 |
| Channel  | Odysseas |
| Duration | 36m 26s |
| Uploaded | 2024-04-24 |
| Fetched  | 2026-03-04 |
| Source   | my-project |
| Language | en (manual) |

## Video Description
The creator's description, links, and chapter markers...

## Transcript
The complete caption text...

A video from another site gives the same kind of note, tagged media-transcript, with platform, creator, transcription_engine and transcription_model in the frontmatter, a Frames section that embeds the contact sheet with each tile's time, and a Whisper transcript with timestamps:

~/yt_transcripts/2026-10-07_claude-motion-tips_[instagram-dehp8dpsimi].md
~/yt_transcripts/2026-10-07_claude-motion-tips_[instagram-dehp8dpsimi].frames.jpg

The YAML frontmatter makes a collection queryable through tools such as Dataview, while the Markdown remains portable to Logseq, other knowledge bases, and plain text workflows.

Why this exists

Most transcript extractors stop at raw caption text. An archival knowledge note also needs the source URL, creator, capture date, actual language, description, chapters, and a predictable filename. Video Fetcher to Markdown keeps that complete record in one local file. Short social videos often show the real content on screen (tool names, prompts, links) rather than saying it, so notes from those sites include frames as well as words.

Features

  • Manual and auto-generated captions with optional timestamps
  • Clickable timestamps and chapters that jump to the moment in the video
  • Ordered language preferences, regional variants, automatic selection, and strict language matching
  • Explicit YouTube translation, labeled with source language and machine-translation provenance
  • Title, channel, duration, upload date, description, and chapters when available
  • Safe YAML frontmatter and Markdown tables for dynamic metadata
  • File protection for every format, safe replacement, and refreshes that update existing notes in place
  • Obsidian-vault and custom-directory output
  • Plain text, JSON, SRT, and WebVTT export
  • Bounded network requests, useful errors, and optional metadata-free capture
  • No API keys and no hosted service
  • Other sites (new): Instagram, TikTok, X, Vimeo, Facebook and anything else yt-dlp supports, plus YouTube videos without captions and local files, transcribed on your machine with Whisper
  • Frames (new): a contact sheet of the video for reading on-screen text, by default for videos up to 3 minutes
  • Login walls (new): a clear exit code and a browser fallback for sites such as Instagram; your browser login is used only when you ask

Installation

Install the skill

npx skills add JimmySadek/video-fetcher-to-markdown

Or clone the canonical repository:

git clone https://github.com/JimmySadek/video-fetcher-to-markdown.git

Install runtime dependencies

Python 3.8–3.14 is supported for captions. Python 3.10 or newer is recommended for current optional yt-dlp releases. From the cloned or installed skill directory, use an isolated environment so your system Python stays unchanged:

python3 -m venv .venv
.venv/bin/python -m pip install -r requirements.txt
.venv/bin/python scripts/fetch_transcript.py --check-deps

On Windows PowerShell:

py -m venv .venv
.venv\Scripts\python.exe -m pip install -r requirements.txt
.venv\Scripts\python.exe scripts\fetch_transcript.py --check-deps

Activate that environment before using the python3 examples below (source .venv/bin/activate on macOS/Linux), or use the full interpreter path each time. An agent should also use that interpreter. If your skill installation is read-only, create the environment in a writable location and pass the full path to requirements.txt.

yt-dlp is optional for descriptions, chapters, duration, and upload dates:

.venv/bin/python -m pip install yt-dlp
# Windows: .venv\Scripts\python.exe -m pip install yt-dlp

Put its executable on PATH by activating the environment. Without it, oEmbed still supplies title and channel when accessible. The script never installs packages automatically. --no-metadata skips both metadata providers.

Other video sites need a few more command-line tools (ffmpeg, yt-dlp and a Whisper tool); see Other video sites.

Usage

python3 scripts/fetch_transcript.py "https://youtu.be/VIDEO_ID"

An agent using the skill resolves scripts/fetch_transcript.py relative to its installed SKILL.md; it does not depend on one fixed home-directory path.

Output location

The first configured option wins:

  1. --output for one exact file
  2. --output-dir for this run
  3. VIDEO_FETCHER_DIR for a persistent directory (the older YOUTUBE_FETCHER_DIR still works)
  4. ~/yt_transcripts/ by default
# Save this note to an Obsidian vault
python3 scripts/fetch_transcript.py URL --output-dir ~/Notes/MyVault

# Set a persistent default
export VIDEO_FETCHER_DIR=~/Notes/MyVault
python3 scripts/fetch_transcript.py URL

# Save to one exact file
python3 scripts/fetch_transcript.py URL --output ~/Notes/video.md

Every format preserves an existing destination and exits with code 3, before making a network request when the destination is already known. This is the same in terminals and agent sessions; there is no hidden interactive prompt. --force replaces the chosen file completely, including any annotations. A default Markdown refresh reuses the existing note's path even if its title or capture date has changed. An explicit --output is honored independently of other notes for the same video, so distinct files can hold different languages or versions.

Writes use a temporary file beside the destination. Where the filesystem supports hard links, a new file appears only once its UTF-8 content is complete. Other filesystems use exclusive creation: they still refuse to open an existing file for writing, but a new file can be visible during the write. Handled write failures remove that partial file; abrupt termination or disk failure can leave it behind. Forced refreshes replace a completed temporary file and preserve existing POSIX permission modes. New notes use normal file-creation permissions.

If another process creates the destination during a fetch, the non-force write still refuses to overwrite it. --stdout prints only the result and creates no file, even when output-path options are present; diagnostics go to stderr.

Languages and translation

# Prefer French, then German, with English as the final fallback
python3 scripts/fetch_transcript.py --lang fr,de -- URL

# Require Japanese captions; fail clearly if unavailable
python3 scripts/fetch_transcript.py --lang ja --strict-lang -- URL

# Capture available captions when you do not know their language
python3 scripts/fetch_transcript.py --lang auto -- URL

# Explicit YouTube machine translation of an available track into English
python3 scripts/fetch_transcript.py --lang auto --translate en -- URL

# Inspect caption tracks and supported translation targets
python3 scripts/fetch_transcript.py --list -- URL

Language preference outranks caption type. For each requested language, exact codes are tried before regional variants (es can select es-MX); manual captions win within that match. All requested languages precede English fallback. --strict-lang disables that fallback, while still allowing regional variants. The default remains --lang en for compatibility.

auto selects a manual track if available, otherwise a generated track, using YouTube's track order for ties. It does not establish the original audio language. Selection and fallback are reported to stderr and recorded in the note. Translation happens only with --translate, requires support from YouTube, and is never described as a human translation.

Markdown records requested_language, source_language, language (the actual output language), caption_type (the original track's type), translated, and translation_provider when applicable. metadata_source distinguishes yt-dlp, oembed, unavailable, and deliberately skipped metadata. Existing frontmatter keys remain compatible. JSON keeps its existing array of {text, start, duration} objects; raw exports have no provenance wrapper, so retain stderr or use Markdown when that context matters.

Transcript and subtitle exports

python3 scripts/fetch_transcript.py --stdout --timestamps -- URL
python3 scripts/fetch_transcript.py --format txt --stdout -- URL
python3 scripts/fetch_transcript.py --format json --output captions.json -- URL
python3 scripts/fetch_transcript.py --format srt -- URL
python3 scripts/fetch_transcript.py --format vtt -- URL

text and markdown both produce an archival Markdown note; txt produces plain caption text. Timestamped Markdown and chapter lists link to the corresponding video time. Raw exports skip metadata requests. A raw ID starting with - works after --; put all options before that separator.

Options

| Flag | What it does | |------|-------------| | --output / -o | Save to one exact file | | --output-dir | Save inside a directory or knowledge vault | | --timestamps / -t | Add linked Markdown timestamps or plain timestamps in txt | | --lang / -l | One code, ordered comma-separated codes, or auto; default en | | --strict-lang | Disable English fallback | | --translate | Explicit YouTube machine translation to a target language | | --source / -s | Override the capture-project name | | --format / -f | text/markdown (default), txt, json, srt, or vtt | | --no-description | Skip the description and chapters section | | --no-metadata | Skip yt-dlp and oEmbed while retaining captions and source URL | | --timeout | Connect/read timeout per HTTP request, in seconds; default 15 | | --stdout | Print the result instead of saving it | | --list | Show available caption languages | | --force | Replac

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars485
CategoryContent
Updated17h ago
Forks39

Languages

Python

Trust signals

100/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

No cautions