image
Image prompting skill for Nano Banana (NBP/NB2) and GPT Image 2.5 (Flare/Sunburst). Writes ready-to-use prompts with model/quality/size recommendations
Install / Use
npx skills add smixs/visual-skills --skill imageInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AI & Machine LearningSupported Platforms
Our assessment of image
image scores 83/100 on our quality scale, 692nd of 961 AI & Machine Learning skills we index.
Its SKILL.md is 8.4 KB long, well organised into 11 sections with 2 code examples: a thorough specification that gives an agent plenty to work with.
It has 431 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated 20 days ago, so image is actively maintained.
- It is released under the CC-BY-4.0 license; check its terms before commercial use.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
image compared with similar skills
All 4 of these similar skills score higher than image; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| image (this skill)by smixs | 83 | 431 | 20d ago | SKILL.md |
| claude-memby thedotmack | 100 | 97.1k | today | CLAUDE.md |
| Agent-Reachby Panniantong | 100 | 92.6k | 21d ago | CLAUDE.md |
| Understand-Anythingby Egonex-AI | 100 | 85.4k | today | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.5k | today | CLAUDE.md |
Frequently asked questions
- How do I install image?
- Run
npx skills add smixs/visual-skills --skill image. The install tabs above show the steps for each supported agent. - Which AI agents does image work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is image safe to use?
- It is CC-BY-4.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is image still maintained?
- The repository was last updated 20 days ago, so image is actively maintained.
Skill content
View source on GitHubname: image license: CC-BY-4.0 (attribution required — Serge Shima, github.com/smixs/visual-skills) description: > Image prompting skill for Nano Banana (NBP/NB2) and GPT Image 2.5 (Flare/Sunburst). Writes ready-to-use prompts with model/quality/size recommendations. Use when: "нарисуй", "сгенерируй картинку", "image prompt", "промпт для картинки", blog covers, slides, posters, product shots, UI mockups, storyboards, character sheets, edit/colorize, style transfer, vision analysis, image-to-prompt, nb, NBP, NB2, gpt-image-2.5, multi-panel grids, ecommerce product photography, fashion editorial, food/beverage ads, cinematic portraits. Do NOT use for: video (use video skill), 3D models, audio, non-image tasks.
Image Prompting — Nano Banana & GPT Image 2.5
This skill writes image prompts. It does not generate images. The output is: model name + quality / size / aspect ratio + the prompt itself.
The body of this SKILL.md is intentionally thin so you cannot fake a result by reading it alone. The actual rules — what the models reward, what they punish, how to phrase a 5-slot template, when to add quality: high, when to use image grounding — live only in the reference files.
Route first — is this actually an image-prompt task?
- Motion, clips, montage (Seedance, Kling, Veo, any image-to-video): use the sibling
videoskill. This skill's storyboard and keyframe outputs feed it. - No idea or script yet (user wants a concept or an ad scenario, not a picture): if the
creative-directorskill is installed, start there — it develops ideas and scripts for commercials and beyond (github.com/smixs/creative-director-skill). - A concrete image is needed — this skill. Continue below.
Mandatory reading order — DO NOT WRITE A PROMPT WITHOUT THIS
Past attempts to write prompts directly from this skill body produced lazy, generic results. Each model has its own physics; common rules collapse into mush when applied without model-specific syntax. Read in this order before producing any prompt:
Step 1 — always read first → models.md
Decide: Nano Banana (NB2 or NBP) or GPT Image 2.5 (Flare for speed, Sunburst for precision edits). The choice changes the prompt syntax fundamentally — natural-language paragraphs vs. labeled 5-slot template, quality settings, which features exist (image grounding only on NB, EXACT TEXT discipline only on GPT Image, etc.).
If the user named a model — confirm and proceed. If not — pick using the table in models.md, then state your choice in the output header.
Step 2 — read one model file (the one you picked)
-
Nano Banana → nano-banana.md Image grounding for real locations. Extreme aspect ratios (1:8, 8:1, 4:1). Thinking mode. JSON for 5+ elements. Up to 14 reference images. Why you must NOT write
50mm / f-stop / ISOnumbers. -
GPT Image 2.5 → gpt-image.md 5-slot template (Scene / Subject / Important Details / Use Case / Constraints). Anti-slop banned-words list.
quality: low / medium / high / xhigh / maxas a deliberate fidelity lever. Size constraints (multiples of 16, max 3:1, up to 4K 3840×2160). Two-column edit logic (Change / Preserve / Constraints). Up to 16 reference images with explicit roles.
The model file is non-negotiable. Skipping it is the single biggest cause of weak prompts.
Step 3 — always read after the model file → golden-rules.md
Universal rules that apply to both models: start with a verb, positive framing, hex colors, quote text, edit don't re-roll, one change per iteration, reference images.
Step 4 — task-shaped reading (load only what matches the request)
Pick zero or more, depending on what the user asked for:
- Text in image, infographic, diagram, multilingual rendering → text-rendering.md
- Edit existing image (object removal, lighting swap, colorization, restoration, localization) → editing.md
- Character continuity across multiple images / panels → characters.md
- The image must pass as a real photograph (portrait, reportage, UGC, casting, product-in-hand) — or the user says the result "looks AI", "too glossy", "not like the reference" → de-slop.md. Model default priors, banned booster words, capture pipeline instead of adjectives, located imperfections.
- Presentation slides → slides.md
- Sequential narrative (storyboard, comic, panel sequence) → storyboards.md
- Sketch → final, wireframes, structural input → structural.md
- 2D → 3D, floor plans, isometric → dimensional.md
- Vision analysis / image-to-prompt / style transfer from a reference image → vision-decomposer.md. Load this whenever the user attaches an image and asks to recreate, match, decompose, or transfer its style.
- Multi-panel compositions (grids, collages, storyboard sheets in ONE image) → multi-panel.md. 9-cell TVC grids, 2x2 portrait grids, 3-panel campaign collages, 4x3 borderless grids, 6-frame cinematic sequences, before/after splits, 12-panel storyboard posters.
- Industry pattern libraries — proven prompt templates by vertical. Load the matching file:
- E-commerce product shots → patterns/ecommerce.md
- Fashion editorial campaigns → patterns/fashion-editorial.md
- Food & beverage advertising → patterns/food-beverage.md
- Cinematic portraits → patterns/portrait-cinema.md
- Posters & illustration → patterns/poster-illustration.md. Also holds the Style DNA + Reject Checklist method — how to fix a strong visual style in four lines before generating and how to judge the returned image against it. Usable with any pattern in this list, not only posters.
- Character design (turnarounds, expression sheets, outfit grids) → patterns/character-design.md
- UI mockups & social media formats → patterns/ui-social.md
Step 5 — read for production language → creative-direction.md
Studio-quality vocabulary for lighting design, camera and hardware, color grading and film stock, materiality and texture. Read when you need precise terms beyond what golden-rules.md covers.
Step 6 — read if structuring a complex prompt → prompt-framework.md
Universal element checklist (subject, context, action, environment, camera, lighting, mood, materials, palette, format), detail modes (concise / standard / verbose / cinematic verbose), parameterized templates, output structure with parameters and exclusions.
Output format
When you return the prompt, structure it like this:
Model: <nano-banana-2 | nano-banana-pro | gpt-image-2.5-flare | gpt-image-2.5-sunburst>
Quality: <low | medium | high | xhigh | max> (only for gpt-image-2.5)
Size / Ratio: <e.g. 1536×1024 or 16:9>
Prompt:
<the prompt text, ready to copy>
Notes:
- <anything you inferred or assumed because the user did not specify>
For edits, also include an explicit preserve-list (mandatory for gpt-image-2.5, recommended for nano-banana):
Change: <one concrete thing>
Preserve: <face, pose, lighting, framing, geometry, ...>
Constraints: <no extra objects, no drift, ...>
Final response style
Prefer: ready-to-copy prompts, hex colors, concrete materials, named compositions, model-specific syntax (5-slot for GPT Image, natural prose for Nano Banana).
Avoid: tag soup ("cool, modern, 4k"), vague praise ("stunning, epic, masterpiece" — actively hurts GPT Image 2.5), negative framing ("no people, no cars" — invert to positive), external comparisons ("like Apple ad" — describe the visual properties instead), numerical lens parameters in Nano Banana prompts (it ignores them).
Author: Serge Shima (t.me/aimastersme · sergeshima.com · aimasters.me) · License: CC BY 4.0 — attribution required · Source: smixs/visual-skills
Related Skills
claude-mem
97.1kPersistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Agent-Reach
92.6kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
Understand-Anything
85.4kGraphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
headroom
74.5kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
