gemini-video-understanding
Analyze videos with Google Gemini API (summaries, Q&A, transcription with timestamps + visual context, scene/timeline detection, video clipping, FPS control, multi-video comparison, and YouTube URL analysis).
Install / Use
npx skills add benchflow-ai/skillsbench --skill gemini-video-understandingInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AutomationSupported Platforms
Our assessment of gemini-video-understanding
gemini-video-understanding scores 92/100 on our quality scale, 792nd of 3,055 Automation skills we index (top 26%).
Its SKILL.md is 9.4 KB long, well organised into 25 sections with 11 code examples: a thorough specification that gives an agent plenty to work with.
With 1,813 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated about 2 months ago, so gemini-video-understanding is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-10-02. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
gemini-video-understanding compared with similar skills
All 4 of these similar skills score higher than gemini-video-understanding; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| gemini-video-understanding (this skill)by benchflow-ai | 92 | 1.8k | 2mo ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 87.6k | 16d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.3k | today | CLAUDE.md |
| rufloby ruvnet | 100 | 73.7k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 85.1k | 1d ago | MCP Server |
Frequently asked questions
- How do I install gemini-video-understanding?
- Run
npx skills add benchflow-ai/skillsbench --skill gemini-video-understanding. The install tabs above show the steps for each supported agent. - Which AI agents does gemini-video-understanding work with?
- It is written for Gemini CLI, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is gemini-video-understanding safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is gemini-video-understanding still maintained?
- The repository was last updated about 2 months ago, so gemini-video-understanding is actively maintained.
Skill content
View source on GitHubname: gemini-video-understanding description: Analyze videos with Google Gemini API (summaries, Q&A, transcription with timestamps + visual context, scene/timeline detection, video clipping, FPS control, multi-video comparison, and YouTube URL analysis).
Gemini Video Understanding Skill
Purpose
This skill enables video understanding workflows using the Google Gemini API, including video summarization, question answering, transcription with optional visual descriptions, timestamp-based queries (MM:SS), scene/timeline detection, video clipping, custom FPS sampling, multi-video comparison, and YouTube URL analysis.
When to Use
- Summarizing a video into key points or chapters
- Answering questions about what happens at specific timestamps (MM:SS)
- Producing a transcript (optionally with visual context) and speaker labels
- Detecting scene changes or building a timeline of events
- Analyzing long videos by clipping to relevant segments or reducing FPS
- Comparing multiple videos (up to 10 videos on Gemini 2.5+)
- Analyzing public YouTube videos directly via URL
Required Libraries
The following Python libraries are required:
from google import genai
from google.genai import types
import os
import time
Input Requirements
- File formats: MP4, MPEG, MOV, AVI, FLV, MPG, WebM, WMV, 3GPP
- Size constraints:
- Use inline bytes for small files (rule of thumb: <20MB).
- Use the File API upload flow for larger videos (most real videos).
- YouTube:
- Video must be public (not private/unlisted) and not age-restricted.
- Provide a valid YouTube URL.
- Duration / context window (model-dependent):
- 2M-token models: ~2 hours (default resolution) or ~6 hours (low-res).
- 1M-token models: ~1 hour (default) or ~3 hours (low-res).
- Timestamps: Use MM:SS (e.g.,
01:15) when requesting time-based answers.
Output Schema
All extracted/derived content should be returned as valid JSON conforming to this schema:
{
"success": true,
"source": {
"type": "file|youtube",
"id": "video.mp4|VIDEO_ID_OR_URL",
"model": "gemini-2.5-flash"
},
"summary": "Concise summary of the video...",
"transcript": {
"available": true,
"text": "Full transcript text (may include speaker labels)...",
"includes_visual_descriptions": true
},
"events": [
{
"timestamp": "MM:SS",
"description": "What happens at this time",
"category": "scene_change|key_point|action|other"
}
],
"warnings": [
"Optional warnings about limitations, missing timestamps, or low confidence areas"
]
}
Field Descriptions
success: Whether the analysis completed successfullysource.type:filefor uploaded/local content,youtubefor YouTube analysissource.id: Filename for local uploads, or URL/ID for YouTubesource.model: Gemini model used for the requestsummary: High-level video summarytranscript.*: Transcript payload (may be omitted oravailable=falseif not requested)events: Timeline items with MM:SS timestamps (chapters, scene changes, key actions)warnings: Any issues that could affect correctness (e.g., “timestamp not found”, “long video clipped”)
Code Examples
Basic Video Analysis (Local Video + File API)
from google import genai
import os
import time
client = genai.Client(api_key=os.getenv("GEMINI_API_KEY"))
# Upload video (File API for >20MB)
myfile = client.files.upload(file="video.mp4")
# Wait for processing
while myfile.state.name == "PROCESSING":
time.sleep(1)
myfile = client.files.get(name=myfile.name)
if myfile.state.name == "FAILED":
raise ValueError("Video processing failed")
response = client.models.generate_content(
model="gemini-2.5-flash",
contents=["Summarize this video in 3 key points", myfile],
)
print(response.text)
YouTube Video Analysis (Public Videos Only)
from google import genai
from google.genai import types
import os
client = genai.Client(api_key=os.getenv("GEMINI_API_KEY"))
response = client.models.generate_content(
model="gemini-2.5-flash",
contents=[
"Summarize the main topics discussed",
types.Part.from_uri(
uri="https://www.youtube.com/watch?v=VIDEO_ID",
mime_type="video/mp4",
),
],
)
print(response.text)
Inline Video (<20MB)
from google import genai
from google.genai import types
import os
client = genai.Client(api_key=os.getenv("GEMINI_API_KEY"))
with open("short-clip.mp4", "rb") as f:
video_bytes = f.read()
response = client.models.generate_content(
model="gemini-2.5-flash",
contents=[
"What happens in this video?",
types.Part.from_bytes(data=video_bytes, mime_type="video/mp4"),
],
)
print(response.text)
Video Clipping (Analyze a Segment Only)
from google.genai import types
response = client.models.generate_content(
model="gemini-2.5-flash",
contents=[
"Summarize this segment",
types.Part.from_video_metadata(
file_uri=myfile.uri,
start_offset="40s",
end_offset="80s",
),
],
)
Custom Frame Rate (Token/Cost Control)
from google.genai import types
# Lower FPS for static content (saves tokens)
response = client.models.generate_content(
model="gemini-2.5-flash",
contents=[
"Analyze this presentation",
types.Part.from_video_metadata(file_uri=myfile.uri, fps=0.5),
],
)
# Higher FPS for fast-moving content
response = client.models.generate_content(
model="gemini-2.5-flash",
contents=[
"Analyze rapid movements in this sports video",
types.Part.from_video_metadata(file_uri=myfile.uri, fps=5),
],
)
Timeline / Scene Detection (MM:SS)
response = client.models.generate_content(
model="gemini-2.5-flash",
contents=[
"""Create a timeline with timestamps:
- Key events
- Scene changes
- Important moments
Format: MM:SS - Description
""",
myfile,
],
)
Transcription (Optional Visual Descriptions)
response = client.models.generate_content(
model="gemini-2.5-flash",
contents=[
"""Transcribe with visual context:
- Audio transcription
- Visual descriptions of important moments
- Timestamps for salient events
""",
myfile,
],
)
Structured JSON Output (Schema-Guided)
from pydantic import BaseModel
from typing import List
from google.genai import types as genai_types
class VideoEvent(BaseModel):
timestamp: str # MM:SS
description: str
category: str
class VideoAnalysis(BaseModel):
summary: str
events: List[VideoEvent]
duration: str
response = client.models.generate_content(
model="gemini-2.5-flash",
contents=["Analyze this video", myfile],
config=genai_types.GenerateContentConfig(
response_mime_type="application/json",
response_schema=VideoAnalysis,
),
)
Best Practices
- Use the File API for most videos (>20MB) and wait for processing to complete before analysis.
- Reduce token usage by clipping to the relevant segment and/or lowering FPS for static content.
- Improve accuracy by being explicit about the desired output (timestamps format, number of events, whether you want scene changes vs actions vs chapters).
- Use
gemini-2.5-prowhen you need highest-quality reasoning over complex, long, or visually dense videos.
Error Handling
import time
def upload_and_wait(client, file_path: str, max_wait_s: int = 300):
myfile = client.files.upload(file=file_path)
waited = 0
while myfile.state.name == "PROCESSING" and waited < max_wait_s:
time.sleep(5)
waited += 5
myfile = client.files.get(name=myfile.name)
if myfile.state.name == "FAILED":
raise ValueError(f"Video processing failed: {myfile.state.name}")
if myfile.state.name == "PROCESSING":
raise TimeoutError(f"Processing timeout after {max_wait_s}s")
return myfile
Common issues:
- Upload processing stuck: wait and poll; fail after a max timeout.
- YouTube errors: verify the video is public and not age-restricted.
- Rate limits: retry with exponential backoff.
- Incorrect timestamps: re-prompt with strict “MM:SS” formatting and request fewer events.
Limitations
- Long-video support is limited by model context and token budget (default vs low-res modes).
- YouTube analysis requires public videos; live streaming analysis is not supported.
- Very long videos may require chunking (clip by time range and process in segments).
- Multi-video comparison is limited (up to 10 videos per request on Gemini 2.5+).
Version History
- 1.0.0 (2026-01-15): Initial release focused on Gemini video understanding (summaries, Q&A, timestamps, clipping, FPS control, YouTube, and structured outputs).
Resources
Related Skills
Agent-Reach
87.6kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.3kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
ruflo
73.7k🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
Scrapling
85.1k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
