gpt-multimodal
Analyze images and multi-frame sequences using OpenAI GPT series
Install / Use
npx skills add benchflow-ai/skillsbench --skill gpt-multimodalInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AI & Machine LearningSupported Platforms
Our assessment of gpt-multimodal
gpt-multimodal scores 91/100 on our quality scale, 253rd of 875 AI & Machine Learning skills we index (top 29%).
Its SKILL.md is 19 KB long, well organised into 48 sections with 17 code examples: a thorough specification that gives an agent plenty to work with.
With 1,813 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated about 2 months ago, so gpt-multimodal is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
gpt-multimodal compared with similar skills
All 4 of these similar skills score higher than gpt-multimodal; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| gpt-multimodal (this skill)by benchflow-ai | 91 | 1.8k | 2mo ago | SKILL.md |
| claude-memby thedotmack | 100 | 95.0k | today | CLAUDE.md |
| Understand-Anythingby Egonex-AI | 100 | 84.8k | 2d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.2k | today | CLAUDE.md |
| CowAgentby zhayujie | 100 | 47.2k | today | CLAUDE.md |
Frequently asked questions
- How do I install gpt-multimodal?
- Run
npx skills add benchflow-ai/skillsbench --skill gpt-multimodal. The install tabs above show the steps for each supported agent. - Which AI agents does gpt-multimodal work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is gpt-multimodal safe to use?
- It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is gpt-multimodal still maintained?
- The repository was last updated about 2 months ago, so gpt-multimodal is actively maintained.
Skill content
View source on GitHubname: gpt-multimodal description: Analyze images and multi-frame sequences using OpenAI GPT series
OpenAI Vision Analysis Skill
Purpose
This skill enables image analysis, scene understanding, text extraction, and multi-frame comparison using OpenAI's vision-capable GPT models (e.g., gpt-4o, gpt-5). It supports single and multiple images analysis and sequential frames for temporal analysis.
When to Use
- Analyzing image content (objects, scenes, colors, spatial relationships)
- Extracting and reading text from images (OCR via vision models)
- Comparing multiple images to detect differences or changes
- Processing video frames to understand temporal progression
- Generating detailed image descriptions or captions
- Answering questions about visual content
Required Libraries
The following Python libraries are required:
from openai import OpenAI
import base64
import json
import os
from pathlib import Path
Input Requirements
- File formats: JPG, JPEG, PNG, WEBP, non-animated GIF
- Image quality: Clear and legible; minimum 512×512px recommended
- File size: Under 20MB per image recommended
- Maximum per request: Up to 500 images, 50MB total payload
- URL or Base64: Images can be provided as URLs or base64-encoded data
Output Schema
All analysis results should be returned as valid JSON conforming to this schema:
{
"success": true,
"model": "gpt-5",
"analysis": "Detailed description or analysis of the image content...",
"metadata": {
"image_count": 1,
"detail_level": "high",
"tokens_used": 850,
"processing_time_ms": 1234
},
"extracted_data": {
"objects": ["car", "person", "building"],
"text_found": "Sample text from image",
"colors": ["blue", "white", "gray"],
"scene_type": "urban street"
},
"warnings": []
}
Field Descriptions
success: Boolean indicating whether the API call succeededmodel: The GPT model used for analysis (e.g., "gpt-4o", "gpt-5")analysis: Complete textual analysis or description from the modelmetadata.image_count: Number of images analyzed in this requestmetadata.detail_level: Detail parameter used ("low", "high", or "auto")metadata.tokens_used: Approximate token count for the requestmetadata.processing_time_ms: Time taken to process the requestextracted_data: Structured information extracted from the image(s)warnings: Array of issues or limitations encountered
Code Examples
Basic Image Analysis
from openai import OpenAI
import base64
def analyze_image(image_path, prompt="What's in this image?"):
"""Analyze a single image using GPT-5 Vision."""
client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))
# Read and encode image
with open(image_path, "rb") as image_file:
base64_image = base64.b64encode(image_file.read()).decode('utf-8')
response = client.chat.completions.create(
model="gpt-5",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": prompt},
{
"type": "image_url",
"image_url": {
"url": f"data:image/jpeg;base64,{base64_image}"
}
}
]
}
],
max_tokens=300
)
return response.choices[0].message.content
Using Image URLs
from openai import OpenAI
def analyze_image_url(image_url, prompt="Describe this image"):
"""Analyze an image from a URL."""
client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))
response = client.chat.completions.create(
model="gpt-5",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": prompt},
{
"type": "image_url",
"image_url": {"url": image_url}
}
]
}
]
)
return response.choices[0].message.content
Multiple Images Analysis
from openai import OpenAI
import base64
def analyze_multiple_images(image_paths, prompt="Compare these images"):
"""Analyze multiple images in a single request."""
client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))
# Build content array with text and all images
content = [{"type": "text", "text": prompt}]
for image_path in image_paths:
with open(image_path, "rb") as image_file:
base64_image = base64.b64encode(image_file.read()).decode('utf-8')
content.append({
"type": "image_url",
"image_url": {
"url": f"data:image/jpeg;base64,{base64_image}"
}
})
response = client.chat.completions.create(
model="gpt-5",
messages=[{"role": "user", "content": content}],
max_tokens=500
)
return response.choices[0].message.content
Full Analysis with JSON Output
from openai import OpenAI
import base64
import json
import time
def analyze_image_to_json(image_path, prompt="Analyze this image"):
"""Analyze image and return structured JSON output."""
client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))
start_time = time.time()
warnings = []
try:
# Read and encode image
with open(image_path, "rb") as image_file:
base64_image = base64.b64encode(image_file.read()).decode('utf-8')
# Make API call
response = client.chat.completions.create(
model="gpt-5",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": prompt},
{
"type": "image_url",
"image_url": {
"url": f"data:image/jpeg;base64,{base64_image}",
"detail": "high"
}
}
]
}
],
max_tokens=500
)
analysis = response.choices[0].message.content
tokens_used = resp…[redacted]
processing_time = int((time.time() - start_time) * 1000)
result = {
"success": True,
"model": "gpt-5",
"analysis": analysis,
"metadata": {
"image_count": 1,
"detail_level": "high",
"tokens_used": tokens_used,
"processing_time_ms": processing_time
},
"extracted_data": {},
"warnings": warnings
}
except Exception as e:
result = {
"success": False,
"model": "gpt-5",
"analysis": "",
"metadata": {
"image_count": 0,
"detail_level": "high",
"tokens_used": 0,
"processing_time_ms": 0
},
"extracted_data": {},
"warnings": [f"API call failed: {str(e)}"]
}
return result
# Usage
result = analyze_image_to_json("photo.jpg", "Describe what you see in detail")
print(json.dumps(result, indent=2))
Batch Processing with Sequential Frames
from openai import OpenAI
import base64
from pathlib import Path
def process_video_frames(frames_directory, analysis_prompt):
"""Process sequential video frames for temporal analysis."""
client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))
image_extensions = {'.jpg', '.jpeg', '.png', '.webp'}
frame_paths = sorted([
f for f in Path(frames_directory).iterdir()
if f.suffix.lower() in image_extensions
])
# Analyze frames in groups (e.g., 5 frames at a time)
batch_size = 5
results = []
for i in range(0, len(frame_paths), batch_size):
batch = frame_paths[i:i+batch_size]
# Build content with all frames in batch
content = [{"type": "text", "text": analysis_prompt}]
for frame_path in batch:
with open(frame_path, "rb") as f:
base64_image = base64.b64encode(f.read()).decode('utf-8')
content.append({
"type": "image_url",
"image_url": {
"url": f"data:image/jpeg;base64,{base64_image}",
"detail": "low" # Use low detail for video frames to save tokens
}
})
response = client.chat.completions.create(
model="gpt-5",
messages=[{"role": "user", "content": content}],
max_tokens=800
)
results.append({
"batch_index": i // batch_size,
"frame_range": f"{batch[0].name} to {batch[-1].name}",
"analysis": response.choices[0].message.content
})
return results
Text Extraction from Images (OCR Alternative)
from openai import OpenAI
import base64
def extract_text_with_gpt(image_path):
"""Extract text from image using GPT Vision as OCR alternative."""
client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))
with open(image_path, "rb") as image_file:
base64_image = base64.b64encode(image_file.read()).decode('utf-8')
response = client.chat.completions.create(
model="gpt-5",
messages=[
{
"role": "user",
"content": [
{
"type": "text",
"text": "Extract all text from this image. Return only the text content, preserving the layout and structure."
},
{
"type": "image_url",
"image_url": {
"url": f"data:image/jpeg;base64,{base64_image}",
"detail": "high"
}
}
]
}
],
max_tokens=1000
)
return response.choices[0].message.content
Model Selection and Configuration
Available Models
# GPT-4o - Best for general vision tasks, fast and cost-effective
model = "gpt-4o"
# GPT-5-nano - Faster and cheaper for simple vision tasks
model = "gpt-5-nano"
# GPT-5 - More capable for complex reasoning
model = "gpt-5"
Detail Level Configuration
Control how much visual detail the model processes:
# Low detail - 512×512px resolution, fewer tokens, faster
"image_url": {
"url": image_url,
"detail": "low"
}
# High detail - Full resolution with tiling, more tokens, better accuracy
"image_url": {
"url": image_url,
"detail": "high"
}
# Auto - Model chooses appropriate detail level
"image_url": {
"url": image_url,
"detail": "auto"
}
When to use each detail level:
- Low: Video frames, simple scene classification, color/shape detection
- High: Text extraction, detailed object detection, fine-grained analysis
- Auto: General purpose when unsure; model optimizes cost vs. quality
Token Cost Management
Understanding Image Tokens
Image tokens count toward your request limits and costs:
- Low detail: Fixed ~85 tokens per image (gpt-5)
- High detail: Base tokens + tile tokens based on image dimensions
- Images are scaled to fit within 2048×2048px
- Divided into 512×512px tiles
- Each tile costs additional tokens
Cost Calculation Examples
# For gpt-5 with high detail:
# - Base: 85 tokens
# - Per tile: 170 tokens
# - Example: 1024×1024 image = 85 + (2×2 tiles × 170) =
Truncated for display — read the full file on GitHub.
Related Skills
claude-mem
95.0kPersistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Understand-Anything
84.8kGraphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
headroom
74.2kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
CowAgent
47.2kOpen-source personal AI assistant & Agent Harness. Plans tasks, runs tools and skills, self-evolves with memory and knowledge. Multi-agent, multi-model, multi-channel. Lightweight, extensible, one-line install.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
