SkillAgentSearch skills...

gemini-omni-flash-api

Use this skill for generative video editing, text-to-video, image-referenced video generation, first-frame-to-video, first-and-last-frame transitions, and video extensions using Gemini Omni 1.1 Flash (gemini-omni-1.1-flash) via the official google-genai SDK.

Install / Use

npx skills add google-gemini/gemini-skills --skill gemini-omni-flash-api

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

95/100

Category

Automation

Supported Platforms

Gemini CLI

Our assessment of gemini-omni-flash-api

gemini-omni-flash-api scores 95/100 on our quality scale, 282nd of 2,124 Automation skills we index (top 14%).

Its SKILL.md is 25 KB long, well organised into 20 sections with 5 code examples: a thorough specification that gives an agent plenty to work with.

With 4,205 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
30/30
Structure
20/20
Description
15/15
Adoption
15/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 5 days ago, so gemini-omni-flash-api is actively maintained.
  • It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

gemini-omni-flash-api compared with similar skills

All 4 of these similar skills score higher than gemini-omni-flash-api; compare them before choosing.

SkillScoreStarsUpdatedFormat
gemini-omni-flash-api (this skill)by google-gemini954.2k5d agoSKILL.md
Agent-Reachby Panniantong10086.0k13d agoCLAUDE.md
headroomby headroomlabs-ai10074.0ktodayCLAUDE.md
rufloby ruvnet10073.4ktodayCLAUDE.md
crawl4aiby unclecode10084.4k3d agoMCP Server

Frequently asked questions

How do I install gemini-omni-flash-api?
Run npx skills add google-gemini/gemini-skills --skill gemini-omni-flash-api. The install tabs above show the steps for each supported agent.
Which AI agents does gemini-omni-flash-api work with?
It is written for Gemini CLI, as a SKILL.md file. Other agents that read the same format can often use it too.
Is gemini-omni-flash-api safe to use?
It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is gemini-omni-flash-api still maintained?
The repository was last updated 5 days ago, so gemini-omni-flash-api is actively maintained.

name: gemini-omni-flash-api description: Use this skill for generative video editing, text-to-video, image-referenced video generation, first-frame-to-video, first-and-last-frame transitions, and video extensions using Gemini Omni 1.1 Flash (gemini-omni-1.1-flash) via the official google-genai SDK. Includes workflows for pre-processing/optimizing high-resolution or long source videos with ffmpeg, stripping audio for full sound regeneration, and handling turn-by-turn video editing and parallel execution.

Gemini Omni Flash Skill

This skill uses the Gemini Omni 1.1 Flash model (gemini-omni-1.1-flash) to perform text to video generation, image to video generation (first frame and last frame transitions), video extensions (up to 40s), and video editing.

[!WARNING] Important Regional Restrictions: Uploading videos to use for video edits or extensions is NOT available in the EEA, Switzerland, the United Kingdom, and some US states. If a video-to-video edit completes quickly with empty outputs (total_output_tokens: 0 or no video content), it is likely due to this restriction.

Core capabilities

  1. Text to video: Generating videos from a text prompt.
  2. First frame to video: Generating videos from a starting image (--first-frame).
  3. First and last frame transition: Generating videos interpolating between a starting image and a final image (--first-frame and --last-frame; note: --last-frame must be used with --first-frame).
  4. Video extensions: Extending existing videos by up to 10 seconds per turn, up to a total length of 40 seconds (--extend or --previous-interaction-id).
  5. Video editing and refinement: Editing existing videos (maximum duration 10 seconds), applying stylistic changes, or performing inpainting/outpainting.
  6. Image and video referenced generation: Using style, character, or object references from images or videos to guide video generation.

Workflow

  1. Analyze request: Determine the target task (e.g., first-frame-to-video, first-and-last-frame transition, video extension, reference-guided editing) and identify any input media assets.

  2. Run SDK scripts:

    • Directly run the appropriate utility (scripts/video/generate_video.py or scripts/upload_file.py).
    • Configure settings like --aspect-ratio (e.g. 16:9, 9:16), --resolution (360p, 720p, 1080p, 4k; default: 720p), and --duration (any integer between 3 and 10 seconds, e.g. 3, 5, 10). Note: 4k requests take longer to generate.
  3. Retrieve and process output: Outputs are saved to the local filesystem (e.g. media/). Report back the completed media path to the user.

Reference Documentation

  • Interactions API: All operations and state management for the Gemini Omni 1.1 Flash model (gemini-omni-1.1-flash) are handled via the Interactions API.
  • Files API: Input media files (such as reference images and videos) must be uploaded via the Files API first before being referenced in generations. The uploaded file URI and MIME type are then included in the interactions.create input parts array.
  • Gemini API Skill Reference: Platform-wide guidelines, current model specifications, and SDK usage rules for the Gemini API.

Dependencies and Prerequisites

  • Python SDK (google-genai): Requires google-genai >= 2.19.0 (Python) to support the interactions client and full video output resolution configuration (360p, 720p, 1080p, 4k). Install or upgrade using:
    pip install -U google-genai
    
  • Python Runtime: Requires Python >= 3.10 (for compatibility with modern google-genai SDK types and methods).
  • ffmpeg & ffprobe: prep_video.py, inspect_video.py, and generate_video.py (when stripping audio via --strip-audio) require ffmpeg and ffprobe binaries installed and available in your system PATH.
  • API Key: Set the GEMINI_API_KEY environment variable:
    export GEMINI_API_KEY="your-api-key"
    

Available scripts

Use the following Python scripts to upload media with the Files API, prepare input videos with ffmpeg, and generate video outputs using the Interactions API.

  1. upload_file.py: Uploads local media (images and videos) to the Files API and polls until ACTIVE. If uploading a video larger than 25MB, it prints an informative warning/tip highlighting that Gemini Omni Flash is optimized for editing 10s videos at 720p/24fps, and recommends pre-processing with prep_video.py first to speed up the upload.

    ./scripts/upload_file.py path/to/image.png
    
  2. generate_video.py: Performs end-to-end video generation and downloads the output video. It detects and uploads local media references (images or videos) before calling the Interactions API. Large video assets (>25MB) will trigger informative pre-processing recommendations without blocking the upload.

    • Text to video:

      ./scripts/video/generate_video.py "A close-up of a cat drinking tea" --output media/cat_tea.mp4
      
    • Output resolution options (--resolution):

      Gemini Omni 1.1 Flash natively supports four output resolutions across both landscape (16:9) and portrait (9:16) aspect ratios:

      • 360p: 640x360 (16:9) or 360x640 (9:16)
      • 720p: 1280x720 (16:9) or 720x1280 (9:16) — (default)
      • 1080p: 1920x1080 (16:9) or 1080x1920 (9:16)
      • 4k: 3840x2160 (16:9) or 2160x3840 (9:16)
      # High-definition (1080p)
      ./scripts/video/generate_video.py "A cinematic drone shot over misty mountains at sunrise" --resolution 1080p --output media/mountains_1080p.mp4
      
      # Ultra-high-definition 4K (Note: 4K requests take longer to generate; pass --timeout if needed)
      ./scripts/video/generate_video.py "A macro shot of a dewdrop on a flower petal in golden sunlight" --resolution 4k --timeout 900 --output media/flower_4k.mp4
      
    • Configurable request timeouts (--timeout):

      Default HTTP timeout is 600 seconds (10 minutes). For computationally intensive requests — such as extending a 30s video in 4K by 10s (up to the maximum 40s total video length) — generation can take several minutes. Use --timeout 900 (or 1200) to provide an extended execution budget.

    • First frame to video:

      ./scripts/video/generate_video.py "The waves crash against the shore." --first-frame start.png --output media/waves.mp4
      
    • First and last frame transition:

      Provide a starting frame and an ending frame to generate a smooth transition between them (note: --last-frame must be used together with --first-frame):

      ./scripts/video/generate_video.py "A smooth timelapse from sunrise to sunset" --first-frame start.png --last-frame end.png --output media/interpolation.mp4
      
    • Looping video (identical start and end frame):

      ./scripts/video/generate_video.py "A crystal orb spinning continuously in place" --first-frame orb.png --last-frame orb.png --output media/loop.mp4
      
    • Image-referenced video generation:

      ./scripts/video/generate_video.py "A cybernetic warrior in the style of <IMAGE_REF_0>" --image reference.png --output media/warrior.mp4
      
    • Video-referenced video generation:

      Provide one or more reference videos (--video-reference / -vr) to guide character, object, or motion style (ideal duration is ~3s, up to 3 reference videos recommended):

      ./scripts/video/generate_video.py "A musician playing cello in the style of <VIDEO_REF_0>" --video-reference ref_dance.mp4 --output media/cello.mp4
      
    • Video extension (extend an existing video):

      Extend an existing video by up to 10 seconds (total duration up to 40 seconds):

      ./scripts/video/generate_video.py "The scene continues as the sun sets over the horizon" --extend media/sunset.mp4 --output media/sunset_extended.mp4
      
    • Video extension with reference images and reference videos:

      Prompt-based extension allows passing reference images and reference videos simultaneously:

      ./scripts/video/generate_video.py "Extend this video. The character in <IMAGE_REF_0> enters dancing like the dancer in <VIDEO_REF_0>." --extend media/sunset.mp4 --image character.png --video-reference dance_ref.mp4 --output media/sunset_extended_with_refs.mp4
      
    • Video editing (keep original audio):

      ./scripts/video/generate_video.py "Transform the style to Japanese anime" --video input.mp4 --output media/anime_style.mp4
      
    • Video editing (regenerate all audio from scratch):

      ./scripts/video/generate_video.py "Transform the style to Japanese anime" --video input.mp4 --strip-audio --output media/anime_style_new_audio.mp4
      
    • Turn-by-turn video editing (edit previous interaction):

      Edit a prior video generation without re-uploading assets by passing the interaction ID:

      ./scripts/video/generate_video.py "Change the setting to a snowy winter wonderland." --previous-interaction-id "v1_..." --output media/winter_wonderland.mp4
      
    • Turn-by-turn video extension (extend previous interaction):

      Extend a prior video generation by passing the previous interaction ID:

      ./scripts/video/generate_video.py "Extend this video. The character turns around and begins to run." --previous-interaction-id "v1_..." --output media/extended_turn.mp4
      
    • Parallel batch execution (prompts file): Run multiple prompts from a line-by-line text file concurrently:

      ./scripts/video/generate_video.py --prompts-file prompts.txt --concurrency 3
      
    • Parallel batch execution (JSON config): Execute fully configured, distinct generation and editing jobs in parallel:

      ./scripts/video/generate_video.py --batch jobs.json --concurrency 3
      

      Example jobs.json:

      [
        {
          "prompt": "A smooth timelapse from sunrise to sunset.",
          "first_frame": "start.png",
          "last_frame": "end.png",
          "resolution": "1080p",
          "output": "media/interpolation.mp4"
        },
        {
          "prompt": "Extend this video. The scene continues with the character in <IMAGE_REF_0> dancing like <VIDEO_REF_0>.",
          "extend": "media/sunset.mp4",
          "image": "character.png",
          "video_reference": "dance_ref.mp4",
          "output": "media/extended_with_refs.mp4"
        },
        {
          "prompt": "A macro shot of a crystal orb refracting cosmic nebula colors.",
          "resolution": "4k",
          "output": "media/nebula_orb_4k.mp4"
        },
        {
          "prompt": "A musician playing cello in the style of <VIDEO_REF_0>.",
          "video_reference": "cello_ref.mp4",
          "output": "media/cello.mp4"
        },
        {
          "prompt": "Transform the style to Japanese anime.",
          "video": "input.mp4",
          "output": "media/anime_style.mp4",
          "strip_audio": false,
          "aspect_ratio": "16:9"
        }
      ]
      
  3. inspect_video.py: Inspects a local video file (using ffprobe) to check its duration, resolution, frame rate (FPS), audio stream presence, and format details.

    ./scripts/video/inspect_video.py media/output.mp4
    
    • To get a pre-parsed, structured JSON summary:

      ./scripts/video/inspect_video.py media/output.mp4 --json
      
    • To get the complete, unmodified ffprobe raw

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars4.2k
CategoryAutomation
Updated5d ago
Forks443

Languages

Python

Trust signals

100/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

No cautions