Best for
- Use when the user needs a blog header, social graphic, product shot, hero image, banner, thumbnail, or any generated image.
MoizIbnYousaf/marketing-cli/skills/image-gen/SKILL.md
Generate images using the brand's visual identity and Gemini API. Reads brand/creative-kit.md for visual style, crafts narrative prompts, and produces images via Nano Banana Pro (gemini-3-pro-image-preview). Supports on-brand and freestyle modes. Use when the user needs a blog header, social graphic, product shot, hero image, banner, thumbnail, or any generated image. Also use proactively when building content that would benefit from visuals. Triggers on "generate image", "create image", "make m
Decision brief
Describe what you need. Get an image that looks like your brand made it.
Compatibility matrix
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Not declared | No explicit evidence | Portability before use |
| Claude Code | Not declared | No explicit evidence | Portability before use |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/MoizIbnYousaf/marketing-cli --skill "skills/image-gen"Inspect the Agent Skill "image-gen" from https://github.com/MoizIbnYousaf/marketing-cli/blob/3074fe0eb48483c4fb63126e1643b512b561c46b/skills/image-gen/SKILL.md at commit 3074fe0eb48483c4fb63126e1643b512b561c46b. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
Use AskUserQuestion. Ask one at a time. Skip questions the user already answered in their request.
You are an expert Nano Banana prompt engineer. Your job is to turn the user's brief into a single, high-quality prompt for Nano Banana 2 (or Pro), a "thinking" image model used for professional asset production.
Default to gemini-3.1-flash-image-preview (Nano Banana 2). Launched Feb 26, 2026. Pro quality at Flash speed/pricing. 4K support, up to 14 reference images for style consistency. Use gemini-3-pro-image-preview (Nano Banana Pro) only when text-heavy infographics or premium qualit…
After generating, show the image and offer:
1. Check GEMINIAPIKEY environment variable. - If missing: "Image generation requires a Gemini API key. Set GEMINIAPIKEY in your environment. Get one at ai.google.dev." - Do not proceed without it.
Permission review
The documentation asks the agent to create, modify, or delete local files.
Save image to project directory (e.g., `images/`, `assets/`, or wherever the project keeps visuals)Evidence record
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 94/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 30 | Source | Repository attention, not individual Skill quality |
| Compatibility | 0 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
Describe what you need. Get an image that looks like your brand made it.
This skill reads your visual brand identity from brand/creative-kit.md, crafts a narrative prompt that bakes in your style constraints, and generates the image via Gemini API. No brand style defined yet? It still works — just at a lower enhancement level. Run /visual-style first for the best results.
Check GEMINI_API_KEY environment variable.
Read brand files in priority order:
brand/creative-kit.md — look for ## Visual Brand Style sectionbrand/voice-profile.md — personality informs image tonebrand/positioning.md — angles inform visual metaphorsbrand/landscape.md — Claims Blacklist (don't generate imagery that visually implies blacklisted claims)Determine enhancement level:
| Level | Context available | Image quality |
|---|---|---|
| L0 | No brand files | Good generic images — uses discovery questions + prompt craft |
| L1 | voice-profile.md | Personality-aligned — playful brand gets warm/bright images |
| L2 | + creative-kit.md colors/typography | Color-constrained — palette woven into prompts |
| L3 | + Visual Brand Style section | Fully on-brand — style anchors, lighting, mood, composition all applied |
Use AskUserQuestion. Ask one at a time. Skip questions the user already answered in their request.
Question 1: Purpose "What's this image for?"
This determines aspect ratio:
| Use case | Default ratio | Resolution |
|---|---|---|
| Blog header | 16:9 | 2K |
| Social square | 1:1 | 2K |
| Social story | 9:16 | 2K |
| Hero / banner | 21:9 | 2K |
| Product shot | 4:3 | 4K |
| Thumbnail | 16:9 | 1K |
Question 2: Feeling "What should someone feel when they see this?" (Free text — this drives the prompt's emotional anchor)
Question 3: Style override (only if L3 brand style exists) "Use your on-brand style, or something different?"
If user picks "different," their description overrides the brand style for this image only.
Skip discovery when: The user's request is specific enough. "Generate a 16:9 blog header showing a glowing terminal on a dark background, warm rim lighting" — don't ask what they want, they just told you.
You are an expert Nano Banana prompt engineer. Your job is to turn the user's brief into a single, high-quality prompt for Nano Banana 2 (or Pro), a "thinking" image model used for professional asset production.
Core principle: brief a senior art director, don't list keywords. Write natural language in full sentences. Never use "tag soup" like "dog, park, 4k, realistic."
1. General style. Be specific and descriptive about subject, setting, composition, camera/viewpoint, lighting, mood, materials, and textures. Full sentences, not comma lists.
2. Context and purpose. Always encode the purpose and audience (YouTube thumbnail, app icon, hero banner, tweet graphic, 4K wallpaper). Let purpose guide style, polish level, and framing.
3. Text and infographics. If text must appear, put it clearly in quotes in the prompt. Ask for legible, clean typography and specify style (bold sans-serif, monospace, handwritten). For data, ask the model to compress into infographics, diagrams, or whiteboards.
4. Character and brand consistency. When reference images exist, explicitly refer to them: "Keep the person's facial features exactly the same as Image 1." Allow changes in pose, expression, angle while preserving identity.
5. Grounding and realism. For real data, locations, or products, tell the model to rely on up-to-date factual knowledge. Encourage coherent details consistent with physics.
6. Editing and restoration. For edits to existing images, give semantic instructions: "remove," "replace," "add," "restore," "change the season." Maintain original structure, only change what's intended.
7. Dimensional and structural control. For floor plans, schematics, wireframes, grids, tell the model to follow that layout closely. For 2D↔3D, describe how the new representation should look while preserving key relationships.
8. Resolution, detail, and format. Specify resolution ("high detail suitable for 4K wallpaper," "clean 16:9 thumbnail"). Call out micro details and textures when needed (brushed steel, cracked paint, mossy stone).
9. Narrative and sequences. For multiple images, describe the story arc, emotional beats, what stays consistent across images. Specify count, format, and identity/style consistency.
10. Output rules. Do not ask follow-up questions about the prompt. Resolve small ambiguities with sensible professional defaults. Output a single flowing narrative prompt.
Build the narrative in this order, woven into flowing prose:
When Visual Brand Style exists, weave constraints INTO the narrative — don't add as a separate block:
See references/prompt-patterns.md for proven patterns. See references/visual-metaphors.md for concept-to-metaphor mapping.
import os
from google import genai
from google.genai import types
client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
response = client.models.generate_content(
model="gemini-3.1-flash-image-preview",
contents=["<narrative prompt>"],
config=types.GenerateContentConfig(
response_modalities=['TEXT', 'IMAGE'],
image_config=types.ImageConfig(
aspect_ratio="<ratio>",
image_size="<resolution>",
),
),
)
for part in response.parts:
if part.text:
print(part.text)
elif part.inline_data:
image = part.as_image()
image.save("output.png")
Default to gemini-3.1-flash-image-preview (Nano Banana 2). Launched Feb 26, 2026. Pro quality at Flash speed/pricing. 4K support, up to 14 reference images for style consistency. Use gemini-3-pro-image-preview (Nano Banana Pro) only when text-heavy infographics or premium quality justify the higher cost.
images/, assets/, or wherever the project keeps visuals)brand/assets.md:
| <date> | image | <file-path> | image-gen | <1-line description of what was generated> |
After generating, show the image and offer:
"Image generated. Want me to:"
For adjustments, use Gemini's multi-turn chat API for iterative refinement — pass the previous image + adjustment prompt.
chat = client.chats.create(model="gemini-3-pro-image-preview")
response = chat.send_message("Make the lighting warmer and add more negative space on the left for text")
When the user provides an existing image to modify:
from PIL import Image
img = Image.open("input.png")
response = client.models.generate_content(
model="gemini-3-pro-image-preview",
contents=["<edit instruction>", img],
config=types.GenerateContentConfig(
response_modalities=['TEXT', 'IMAGE'],
),
)
Supports: style transfer, background replacement, object removal, color grading, text overlay.
If brand/landscape.md exists, check the Claims Blacklist before generating:
| Anti-pattern | Why it fails | Instead |
|---|---|---|
| Listing comma-separated keywords | "professional, modern, tech, blue, minimal" produces generic AI art | Write a narrative: "A single glowing terminal on a dark slate desk, soft amber rim lighting catching the screen edges" |
| Ignoring aspect ratio | Generating 1:1 for a blog header means the user has to crop, losing composition | Always ask where the image goes FIRST — ratio determines everything |
| Overriding brand style silently | User defined a visual brand identity for consistency — ignoring it defeats the purpose | Default to brand style. Only override when user explicitly asks for something different |
| Generating without purpose | "A nice image" produces nothing distinctive | Every image needs a purpose (blog header, social, hero) and a feeling (trust, excitement, calm) |
| Saving as JPEG when PNG needed | Gemini returns JPEG by default — transparency and quality loss | Always convert to PNG with image.save("output.png", format="PNG") |
| Skipping the assets log | Future agents won't know what images exist, leading to duplicate generation | Always append to brand/assets.md after generating |
Frequently asked questions
Describe what you need. Get an image that looks like your brand made it.
The source record exposes this install command: npx skills add https://github.com/MoizIbnYousaf/marketing-cli --skill "skills/image-gen". Inspect the command and pinned source before running it.
Static rules flagged write-files in the source; the page lists the matching lines and excerpts.
Alternatives
MoizIbnYousaf/marketing-cli
Use when the user wants to generate an image or video via Higgsfield AI. Covers 30+ models: Soul V2, Seedance 2.0, Kling 3.0, Veo 3.1, GPT Image 2, Nano Banana 2. Also covers Marketing Studio — branded ad video/image with avatars and products. Use whenever: "generate an image", "make a video", "animate this photo", "image-to-video", "img2vid", "edit this image with AI", "produce a clip", "create an ad", "make a UGC video", "marketing video", "brand video", "TV spot", "import product from URL", "
vibeeval/vibecosystem
Full-stack frontend development combining premium UI design, cinematic animations, AI-generated media assets, persuasive copywriting, and visual art. Builds complete, visually striking web pages with real media, advanced motion, and compelling copy. Use when: building landing pages, marketing sites, product pages, dashboards, generating media assets (image/video/audio/music), writing conversion copy, creating generative art, or implementing cinematic scroll animations.
wshobson/agents
Brand-first landing page designer — runs a brand-identity interview (colors, typography, shape language), then generates and iterates on a polished landing page via Stitch with deployment-ready HTML. Use when the user asks to create, design, or build a landing page, homepage, or marketing page and has no established visual direction. Skip when they have a design mockup, need a dashboard or app UI, are working at component level, building a multi-page app, or restyling with known design tokens —
nexscope-ai/Amazon-Skills
Amazon PPC campaign builder and optimizer for sellers. Two modes: (A) Build — design a complete campaign structure from scratch with keyword groupings, bid calculations, and negative keyword lists, (B) Optimize — audit existing campaigns using search term reports, identify keyword funnel opportunities, calculate bid adjustments, and generate a week-by-week action plan. Integrates with amazon-keyword-research for keyword input. No API key required. Use when: (1) setting up Amazon PPC campaigns fo