gemini-omni-flash-api
Use this skill for generative video editing, text-to-video, image-referenced video generation, first-frame-to-video, first-and-last-frame transitions, and video extensions using Gemini Omni 1.1 Flash (gemini-omni-1.1-flash) via the official google-genai SDK.
Install / Use
npx skills add google-gemini/gemini-skills --skill gemini-omni-flash-apiInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AutomationSupported Platforms
Our assessment of gemini-omni-flash-api
gemini-omni-flash-api scores 95/100 on our quality scale, 282nd of 2,124 Automation skills we index (top 14%).
Its SKILL.md is 25 KB long, well organised into 20 sections with 5 code examples: a thorough specification that gives an agent plenty to work with.
With 4,205 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 5 days ago, so gemini-omni-flash-api is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
gemini-omni-flash-api compared with similar skills
All 4 of these similar skills score higher than gemini-omni-flash-api; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| gemini-omni-flash-api (this skill)by google-gemini | 95 | 4.2k | 5d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 86.0k | 13d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.0k | today | CLAUDE.md |
| rufloby ruvnet | 100 | 73.4k | today | CLAUDE.md |
| crawl4aiby unclecode | 100 | 84.4k | 3d ago | MCP Server |
Frequently asked questions
- How do I install gemini-omni-flash-api?
- Run
npx skills add google-gemini/gemini-skills --skill gemini-omni-flash-api. The install tabs above show the steps for each supported agent. - Which AI agents does gemini-omni-flash-api work with?
- It is written for Gemini CLI, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is gemini-omni-flash-api safe to use?
- It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is gemini-omni-flash-api still maintained?
- The repository was last updated 5 days ago, so gemini-omni-flash-api is actively maintained.
Skill content
View source on GitHubname: gemini-omni-flash-api description: Use this skill for generative video editing, text-to-video, image-referenced video generation, first-frame-to-video, first-and-last-frame transitions, and video extensions using Gemini Omni 1.1 Flash (gemini-omni-1.1-flash) via the official google-genai SDK. Includes workflows for pre-processing/optimizing high-resolution or long source videos with ffmpeg, stripping audio for full sound regeneration, and handling turn-by-turn video editing and parallel execution.
Gemini Omni Flash Skill
This skill uses the Gemini Omni 1.1 Flash model (gemini-omni-1.1-flash) to perform text to video generation, image to video generation (first frame and last frame transitions), video extensions (up to 40s), and video editing.
[!WARNING] Important Regional Restrictions: Uploading videos to use for video edits or extensions is NOT available in the EEA, Switzerland, the United Kingdom, and some US states. If a video-to-video edit completes quickly with empty outputs (
total_output_tokens: 0or no video content), it is likely due to this restriction.
Core capabilities
- Text to video: Generating videos from a text prompt.
- First frame to video: Generating videos from a starting image (
--first-frame). - First and last frame transition: Generating videos interpolating between a starting image and a final image (
--first-frameand--last-frame; note:--last-framemust be used with--first-frame). - Video extensions: Extending existing videos by up to 10 seconds per turn, up to a total length of 40 seconds (
--extendor--previous-interaction-id). - Video editing and refinement: Editing existing videos (maximum duration 10 seconds), applying stylistic changes, or performing inpainting/outpainting.
- Image and video referenced generation: Using style, character, or object references from images or videos to guide video generation.
Workflow
-
Analyze request: Determine the target task (e.g., first-frame-to-video, first-and-last-frame transition, video extension, reference-guided editing) and identify any input media assets.
-
Run SDK scripts:
- Directly run the appropriate utility (
scripts/video/generate_video.pyorscripts/upload_file.py). - Configure settings like
--aspect-ratio(e.g.16:9,9:16),--resolution(360p,720p,1080p,4k; default:720p), and--duration(any integer between3and10seconds, e.g.3,5,10). Note:4krequests take longer to generate.
- Directly run the appropriate utility (
-
Retrieve and process output: Outputs are saved to the local filesystem (e.g.
media/). Report back the completed media path to the user.
Reference Documentation
- Interactions API: All operations and state management for the Gemini Omni 1.1 Flash model (
gemini-omni-1.1-flash) are handled via the Interactions API. - Files API: Input media files (such as reference images and videos) must be uploaded via the Files API first before being referenced in generations. The uploaded file URI and MIME type are then included in the
interactions.createinput parts array. - Gemini API Skill Reference: Platform-wide guidelines, current model specifications, and SDK usage rules for the Gemini API.
Dependencies and Prerequisites
- Python SDK (
google-genai): Requiresgoogle-genai >= 2.19.0(Python) to support theinteractionsclient and full video output resolution configuration (360p,720p,1080p,4k). Install or upgrade using:pip install -U google-genai - Python Runtime: Requires Python >= 3.10 (for compatibility with modern
google-genaiSDK types and methods). - ffmpeg & ffprobe:
prep_video.py,inspect_video.py, andgenerate_video.py(when stripping audio via--strip-audio) requireffmpegandffprobebinaries installed and available in your systemPATH. - API Key: Set the
GEMINI_API_KEYenvironment variable:export GEMINI_API_KEY="your-api-key"
Available scripts
Use the following Python scripts to upload media with the Files API, prepare input videos with ffmpeg, and generate video outputs using the Interactions API.
-
upload_file.py: Uploads local media (images and videos) to the Files API and polls until
ACTIVE. If uploading a video larger than 25MB, it prints an informative warning/tip highlighting that Gemini Omni Flash is optimized for editing 10s videos at 720p/24fps, and recommends pre-processing withprep_video.pyfirst to speed up the upload../scripts/upload_file.py path/to/image.png -
generate_video.py: Performs end-to-end video generation and downloads the output video. It detects and uploads local media references (images or videos) before calling the Interactions API. Large video assets (>25MB) will trigger informative pre-processing recommendations without blocking the upload.
-
Text to video:
./scripts/video/generate_video.py "A close-up of a cat drinking tea" --output media/cat_tea.mp4 -
Output resolution options (
--resolution):Gemini Omni 1.1 Flash natively supports four output resolutions across both landscape (
16:9) and portrait (9:16) aspect ratios:360p:640x360(16:9) or360x640(9:16)720p:1280x720(16:9) or720x1280(9:16) — (default)1080p:1920x1080(16:9) or1080x1920(9:16)4k:3840x2160(16:9) or2160x3840(9:16)
# High-definition (1080p) ./scripts/video/generate_video.py "A cinematic drone shot over misty mountains at sunrise" --resolution 1080p --output media/mountains_1080p.mp4 # Ultra-high-definition 4K (Note: 4K requests take longer to generate; pass --timeout if needed) ./scripts/video/generate_video.py "A macro shot of a dewdrop on a flower petal in golden sunlight" --resolution 4k --timeout 900 --output media/flower_4k.mp4 -
Configurable request timeouts (
--timeout):Default HTTP timeout is
600seconds (10 minutes). For computationally intensive requests — such as extending a 30s video in 4K by 10s (up to the maximum 40s total video length) — generation can take several minutes. Use--timeout 900(or1200) to provide an extended execution budget. -
First frame to video:
./scripts/video/generate_video.py "The waves crash against the shore." --first-frame start.png --output media/waves.mp4 -
First and last frame transition:
Provide a starting frame and an ending frame to generate a smooth transition between them (note:
--last-framemust be used together with--first-frame):./scripts/video/generate_video.py "A smooth timelapse from sunrise to sunset" --first-frame start.png --last-frame end.png --output media/interpolation.mp4 -
Looping video (identical start and end frame):
./scripts/video/generate_video.py "A crystal orb spinning continuously in place" --first-frame orb.png --last-frame orb.png --output media/loop.mp4 -
Image-referenced video generation:
./scripts/video/generate_video.py "A cybernetic warrior in the style of <IMAGE_REF_0>" --image reference.png --output media/warrior.mp4 -
Video-referenced video generation:
Provide one or more reference videos (
--video-reference/-vr) to guide character, object, or motion style (ideal duration is ~3s, up to 3 reference videos recommended):./scripts/video/generate_video.py "A musician playing cello in the style of <VIDEO_REF_0>" --video-reference ref_dance.mp4 --output media/cello.mp4 -
Video extension (extend an existing video):
Extend an existing video by up to 10 seconds (total duration up to 40 seconds):
./scripts/video/generate_video.py "The scene continues as the sun sets over the horizon" --extend media/sunset.mp4 --output media/sunset_extended.mp4 -
Video extension with reference images and reference videos:
Prompt-based extension allows passing reference images and reference videos simultaneously:
./scripts/video/generate_video.py "Extend this video. The character in <IMAGE_REF_0> enters dancing like the dancer in <VIDEO_REF_0>." --extend media/sunset.mp4 --image character.png --video-reference dance_ref.mp4 --output media/sunset_extended_with_refs.mp4 -
Video editing (keep original audio):
./scripts/video/generate_video.py "Transform the style to Japanese anime" --video input.mp4 --output media/anime_style.mp4 -
Video editing (regenerate all audio from scratch):
./scripts/video/generate_video.py "Transform the style to Japanese anime" --video input.mp4 --strip-audio --output media/anime_style_new_audio.mp4 -
Turn-by-turn video editing (edit previous interaction):
Edit a prior video generation without re-uploading assets by passing the interaction ID:
./scripts/video/generate_video.py "Change the setting to a snowy winter wonderland." --previous-interaction-id "v1_..." --output media/winter_wonderland.mp4 -
Turn-by-turn video extension (extend previous interaction):
Extend a prior video generation by passing the previous interaction ID:
./scripts/video/generate_video.py "Extend this video. The character turns around and begins to run." --previous-interaction-id "v1_..." --output media/extended_turn.mp4 -
Parallel batch execution (prompts file): Run multiple prompts from a line-by-line text file concurrently:
./scripts/video/generate_video.py --prompts-file prompts.txt --concurrency 3 -
Parallel batch execution (JSON config): Execute fully configured, distinct generation and editing jobs in parallel:
./scripts/video/generate_video.py --batch jobs.json --concurrency 3Example
jobs.json:[ { "prompt": "A smooth timelapse from sunrise to sunset.", "first_frame": "start.png", "last_frame": "end.png", "resolution": "1080p", "output": "media/interpolation.mp4" }, { "prompt": "Extend this video. The scene continues with the character in <IMAGE_REF_0> dancing like <VIDEO_REF_0>.", "extend": "media/sunset.mp4", "image": "character.png", "video_reference": "dance_ref.mp4", "output": "media/extended_with_refs.mp4" }, { "prompt": "A macro shot of a crystal orb refracting cosmic nebula colors.", "resolution": "4k", "output": "media/nebula_orb_4k.mp4" }, { "prompt": "A musician playing cello in the style of <VIDEO_REF_0>.", "video_reference": "cello_ref.mp4", "output": "media/cello.mp4" }, { "prompt": "Transform the style to Japanese anime.", "video": "input.mp4", "output": "media/anime_style.mp4", "strip_audio": false, "aspect_ratio": "16:9" } ]
-
-
inspect_video.py: Inspects a local video file (using
ffprobe) to check its duration, resolution, frame rate (FPS), audio stream presence, and format details../scripts/video/inspect_video.py media/output.mp4-
To get a pre-parsed, structured JSON summary:
./scripts/video/inspect_video.py media/output.mp4 --json -
To get the complete, unmodified
ffproberaw
-
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
86.0kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.0kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
ruflo
73.4k🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
crawl4ai
84.4kOpen-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
