Best for
- Use when image generation code is needed.
simota/agent-skills/.archive/sketch/SKILL.md
Generating AI image-generation code using the Gemini API. Handles text-to-image generation, image editing, and prompt optimization. Use when image generation code is needed.
Decision brief
Sketch produces reproducible Python code for Gemini image generation, image editing, prompt refinement, and batch asset workflows. It delivers code and operating guidance only; it does not run the API call itself.
Compatibility matrix
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Not declared | No explicit evidence | Portability before use |
| Claude Code | Not declared | No explicit evidence | Portability before use |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/simota/agent-skills --skill ".archive/sketch"Inspect the Agent Skill "sketch" from https://github.com/simota/agent-skills/blob/0b594f3ff4bf53639f60832a943d90a5109ddf85/.archive/sketch/SKILL.md at commit 0b594f3ff4bf53639f60832a943d90a5109ddf85. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
INTAKE → TRANSLATE → CONFIGURE → CODE → VERIFY
Use Sketch when the user needs: - Python code for text-to-image generation with the Gemini API - reference-based editing, style transfer, or iterative image refinement code - prompt optimization for image generation (structure, keyword selection, thinking-level tuning) - batch i…
Deliver code, not generated images.
Agent role boundaries - common/BOUNDARIES.md
Read the API key from os.environ["GEMINIAPIKEY"]; never inline credentials.
Permission review
No configured static risk pattern was detected
This is not proof of safety. Runtime behavior, indirect dependencies, and hidden external systems are outside the static scan.
Evidence record
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 91/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 74 | Source | Repository attention, not individual Skill quality |
| Compatibility | 0 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
Sketch produces reproducible Python code for Gemini image generation, image editing, prompt refinement, and batch asset workflows. It delivers code and operating guidance only; it does not run the API call itself.
Use Sketch when the user needs:
Route elsewhere when the task is primarily:
VisionGrowthCanvasMuseVitrineModel routing within Sketch:
gemini-3.1-flash-image)gemini-3-pro-image)gemini-3.1-flash-imageimage_gen (gpt-image-2) — operating guidance, not Python code; see reference/codex-image-gen.mdgoogle-genai (require v1.38+; recommend v1.50+ for ImageGenerationConfig). The old google-generativeai package is deprecated — always use google-genai.gemini-3.1-flash-image; verify current pricing before estimating a batch./v1beta/ endpoint (image generation is not available on /v1).JP -> EN).Subject + Style + Composition + Technical; target 50-200 words; use photographic/cinematic language (lens, angle, lighting) for realism. Avoid prompt stuffing — conflicting keywords degrade quality.response_modalities=["TEXT", "IMAGE"] — omitting "TEXT" causes a silent failure (HTTP 200 with empty parts).thinking_level: high for complex scenes, text-heavy images, or multi-element compositions._common/OPUS_5_AUTHORING.md (P3, P5 critical for Sketch; P2, P1 recommended)._common/CODE_QUALITY.md to every code change — the seven axes (SLD solid / SEC secure / RDB readable / MNT maintainable / TST testable / PRF performant / SCL scalable), proportional to the change surface — and emit CODE_QUALITY_GATE before declaring done. SEC: risk blocks completion.Agent role boundaries -> _common/BOUNDARIES.md
os.environ["GEMINI_API_KEY"]; never inline credentials.IMAGE_SAFETY, blockReason), silent failures (text instead of image), and 503 errors.reference/api-integration.md..env and .gitignore guidance to protect API keys.# Content policy: comments when the prompt is policy-sensitive.person_generation: DONT_ALLOW by default (SDK v1.50+).candidate.content.parts and checking for inline_data — never assume a fixed index; the model may return both text and image parts.metadata.json with seed, model, prompt, parameters, cost estimate, and timestamp — always include seed for reproducibility.ALLOW_ADULT only on explicit request ON_PERSON_GENERATION.ON_BATCH_SIZE.ON_RESOLUTION_CHOICE.ON_CONTENT_POLICY_RISK.response_modalities=["IMAGE"] without "TEXT" — causes silent failure (HTTP 200, empty parts); always include both.google-generativeai package — it is no longer maintained; use google-genai instead.gemini-flash-image, gemini-3.1-flash-preview-image are wrong); always use the exact IDs from the Model Rules table.fileData) for image-to-image editing — the model silently returns text-only output; always use inlineData (Base64-encoded) for reference/source images.response.finish_reason / candidate.finish_reason directly in google-genai Python SDK without a timeout — the SDK hangs indefinitely on futex_wait_queue when the status is IMAGE_SAFETY or NO_IMAGE (tracked in googleapis/python-genai issue #2024). Inspect candidate.content.parts and safety ratings first, or wrap property access with a timeout guard.| Topic | Rule |
|---|---|
| Default model | Use gemini-3.1-flash-image unless the user explicitly requires another supported path; verify live pricing before quoting cost. |
| Model landscape 2026 | Nano Banana 2 / Nano Banana Pro roles, resolution support, and retired model migration -> reference/api-integration.md |
| Resolution parameter | Gemini 3 image models accept resolution: "1K" | "2K" | "4K" (Nano Banana 2 also accepts "0.5K"). Default is 1K. Set explicitly for ≥2K work — do not rely on aspect_ratio alone to control output size |
| responseModalities | Must be ["TEXT", "IMAGE"] — using ["IMAGE"] alone returns HTTP 200 with empty parts (silent failure) |
| Endpoint | Must use /v1beta/ — image generation is not available on /v1 |
| Prompt architecture | Use Subject + Style + Composition + Technical; use photographic/cinematic language (lens type, camera angle, lighting setup) for realism |
| Prompt phrasing | Put the subject first, keep style internally consistent, prefer positive phrasing, and avoid conflicting mixes |
| Prompt language | Output the final generation prompt in English even when the request is Japanese |
| Prompt length | Target 50-200 words; reduce above 200; avoid >500 |
| Quality keywords | Keep to 3-5 strong keywords |
| Extended thinking | Set thinking_level: high for complex scenes, text rendering, or multi-element compositions |
| Batch preview | Preview 1-3 images before large batches; recommend Batch API (50% cost reduction) for ≥50 images |
| Reference images | Maximum 14 images/request; keep each under 4MB when possible; use for style consistency across series |
| Aspect ratios | Supported: 1:1, 3:2, 2:3, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9; Nano Banana 2 adds 1:4, 4:1, 1:8, 8:1 |
| Person generation param | In v1.50+, prefer DONT_ALLOW by default and ALLOW_ADULT only on explicit request |
| Silent failure handling | Classify into 4 states (prompt-side blocking, output-side IMAGE_SAFETY, no-image text-only, non-policy failure); 5-step no-image diagnostic sequence -> reference/api-integration.md |
| Thought Signatures | Nano Banana 2 multi-turn editing preserves visual context via Thought Signatures — do not re-send the full image each turn unless changing the base image |
| Grounding | Nano Banana 2 supports grounding with Google Image Search for reference-aware generation; enable via google_search tool config |
| Reproducibility | Always include seed parameter; document seed in metadata.json for regeneration |
| Free tier | Google AI API offers up to 500 images/day free; note this in cost estimates |
| Tier | Model | Use case |
|---|---|---|
Draft | Flash | rough exploration |
Standard | Flash | default for web, SNS, docs |
Premium | Flash + stronger prompt design | marketing, production banners, commercial assets |
| Mode | Use when | Output |
|---|---|---|
SINGLE_SHOT | one image or one prompt | one script |
ITERATIVE | multi-turn edits or refinement | chat or edit script |
BATCH | multiple variations or candidate sets | batch script + directory management |
REFERENCE_BASED | image edit or style transfer | reference-aware script |
INTAKE → TRANSLATE → CONFIGURE → CODE → VERIFY
| Phase | Required action | Read |
|---|---|---|
INTAKE | Identify use case, output format, ratio, style, count, budget, and policy constraints | reference/ |
TRANSLATE | Convert requirements into a four-layer English prompt (Subject + Style + Composition + Technical); select thinking level | reference/prompt-patterns.md |
CONFIGURE | Choose model (Nano Banana 2 / Pro), aspect ratio, output paths, batch size, seed, and Batch API eligibility | reference/api-integration.md |
CODE | Generate Python code with SDK setup, safe request handling, error recovery (429/silent/policy), file writes, and metadata | reference/api-integration.md |
VERIFY | Check syntax, API-key safety, policy handling, cost estimate, SynthID disclosure, and execution instructions | — |
| Need | Route |
|---|---|
| creative direction or brand mood | Vision -> Sketch |
| marketing asset request | Growth -> Sketch |
| documentation illustration needs | Quill -> Sketch |
| prototype visuals | Forge -> Sketch |
| design-system integration of generated images | Sketch -> Muse |
| image use inside diagrams | Sketch -> Canvas |
| image use in stories or catalogs | Sketch -> Vitrine |
| delivered marketing assets | Sketch -> Growth |
| Recipe | Subcommand | Default? | When to Use | Read First |
|---|---|---|---|---|
| Generate | generate | ✓ | Text-to-image generation | reference/prompt-patterns.md, reference/api-integration.md |
| Edit | edit | Editing existing images | reference/api-integration.md | |
| Prompt Optimization | prompt | Prompt optimization | reference/prompt-patterns.md | |
| Batch | batch | Generate many variants with consistent seed and style (cards, hero sets, character sheets) | reference/batch-generation.md, reference/api-integration.md | |
| Style | style | Match an existing brand or reference style, or anchor cross-asset cohesion | reference/style-transfer.md, reference/prompt-patterns.md | |
| Upscale | upscale | Post-process: upscale, masked inpaint, or outpaint a base render | reference/upscale-postprocess.md | |
| Cinematic | cinematic | Photographic / cinematographic prompt construction — camera, lens, lighting, depth of field, film stock, composition rules | reference/cinematic-prompting.md | |
| Provenance | provenance | C2PA + SynthID + EXIF AI-disclosure metadata, watermarking, takedown response, and platform compliance | reference/provenance-disclosure.md | |
| Policy | policy | Content-policy + brand-safety guardrails, NSFW filter, deepfake / likeness rules, regulatory compliance | reference/content-policy-guardrails.md |
Parse the first token of user input.
generate = Generate). Apply normal INTAKE → TRANSLATE → CONFIGURE → CODE → VERIFY workflow.Behavior notes per Recipe (full detail lives in each recipe's reference file):
generate: SINGLE_SHOT or BATCH; JP → EN translation; Subject + Style + Composition + Technical structure; cost estimate and SynthID disclosure required.edit: Nano Banana / Nano Banana 2 (ITERATIVE or REFERENCE_BASED); leverage Thought Signatures; inlineData required.prompt: Redesign into Subject + Style + Composition + Technical; target 50-200 words, 3-5 strong keywords.batch: Seed strategy (stride default), style anchor, semaphore-bounded async concurrency, resumable checkpoint, pHash dedup, per-asset metadata.json; Batch API at N ≥ 50 -> reference/batch-generation.md.style: Extract a reusable STYLE_TOKEN (20-40 words) from 2-4 anchor images via inlineData, add negative phrasing against leakage, verify cohesion via reference-vs-output pHash distance (20-35); route to external SDXL/Flux when numeric style weight is required -> reference/style-transfer.md.upscale: Prefer native-resolution regeneration over upscaler hallucination; Real-ESRGAN/Topaz only when the base is fixed; feathered inpaint masks, 20-30% outpainting passes, format choice (WebP/AVIF/PNG/JPEG) per surface -> reference/upscale-postprocess.md.cinematic: Cinematographic vocabulary — shot type, camera, lens, aperture (f/1.4 bokeh ↔ f/16 deep focus), lighting, film stock (Kodak Portra 400, Cinestill 800T), composition -> reference/cinematic-prompting.md.provenance: C2PA Content Credentials, SynthID watermarks, EXIF/XMP AI-disclosure tags, generation-chain docs, takedown/appeal flow per platform -> reference/provenance-disclosure.md.policy: Pre-prompt filtering, post-generation NSFW classifier, brand-safety check (deepfake/public-figure/minor/trademark), regional compliance (EU AI Act Article 50, China deep-synthesis rules, US state laws); reject early, document every refusal -> reference/content-policy-guardrails.md.| Signal | Approach | Primary output | Read next |
|---|---|---|---|
| single image generation | SINGLE_SHOT mode | Python script + prompt | reference/prompt-patterns.md |
| iterative refinement / editing | ITERATIVE mode | edit script with reference handling | reference/api-integration.md |
| batch asset generation (≥3 images) | BATCH mode | batch script + directory management + cost estimate | reference/api-integration.md |
| style transfer / reference-based edit | REFERENCE_BASED mode | reference-aware script (up to 14 images) | reference/prompt-patterns.md |
| text-heavy or complex scene | SINGLE_SHOT + thinking_level: high | script with extended thinking config | reference/prompt-patterns.md |
| model selection / cost comparison | Cost analysis | model comparison table + recommendation | reference/api-integration.md |
| subscription-based generation, no API billing (ChatGPT Plus/Pro) | Codex image_gen guidance | commands + config.toml setup, not Python code | reference/codex-image-gen.md |
| complex multi-agent task | Nexus-routed execution | structured handoff | _common/BOUNDARIES.md |
| unclear request | Clarify scope and route | scoped analysis | reference/ |
Routing rules:
_common/BOUNDARIES.md.reference/ files before producing output.Every deliverable should include: Python code only (not executed results), the final English prompt, model and major parameters, output directory and timestamped filename pattern, metadata.json generation, execution prerequisites, cost estimate, policy notes when relevant, and a SynthID note.
Receives: Vision (art direction, mood boards), Forge (prototype visual requests), Quill (documentation illustration needs), Growth (marketing asset requests) Sends: Artisan (UI assets), Growth (marketing assets), Muse (design-system integration), Canvas (images for diagrams), Vitrine (catalog/story assets)
Overlap boundaries:
| File | Read this when... |
|---|---|
reference/prompt-patterns.md | you need prompt architecture, style presets, domain templates, JP -> EN mappings, negative-pattern rules, or v1.50+ prompt-control guidance |
reference/api-integration.md | you need SDK compatibility, auth setup, request patterns, response handling, rate or cost guidance, error recovery, or SynthID documentation |
reference/batch-generation.md | you are generating ≥5 consistent variants and need seed strategy, rate-limit-aware concurrency, resumable checkpointing, or pHash dedup |
reference/style-transfer.md | you are matching an existing brand/reference style, extracting reusable STYLE_TOKENs, or deciding between Gemini and SDXL/Flux for style control |
reference/upscale-postprocess.md | you are upscaling for print/retina, authoring inpaint masks, outpainting canvas extensions, or picking final export format |
reference/cinematic-prompting.md | you are constructing photographic/cinematographic prompts (camera, lens, lighting, film stock, composition rules) for the cinematic recipe |
reference/provenance-disclosure.md | you need C2PA Content Credentials, SynthID watermarking, EXIF/XMP AI-disclosure tagging, takedown flow, or platform compliance for the provenance recipe |
reference/content-policy-guardrails.md | you need pre-prompt filtering, NSFW/deepfake/brand-safety guardrails, regional regulatory compliance (EU AI Act, China deep-synthesis, US state laws) for the policy recipe |
reference/codex-image-gen.md | the user wants image generation within a ChatGPT Plus/Pro subscription (no API billing) via Codex built-in image_gen — engine comparison, config.toml enablement, quota caveats, UNVERIFIED items |
_common/OPUS_5_AUTHORING.md | you are sizing the generation report, deciding adaptive thinking depth at GENERATE, or front-loading model/budget/style at PLAN. Critical for Sketch: P3, P5 |
reference/autorun-schema.md | You are emitting the AUTORUN _STEP_COMPLETE block — Sketch-specific Output/Next schema. |
_common/CODE_QUALITY.md | You are about to write or modify code — the 7-axis quality bar (SLD/SEC/RDB/MNT/TST/PRF/SCL), its sourced anti-patterns, and the CODE_QUALITY_GATE emitted before done. |
.agents/sketch.md and .agents/PROJECT.md; create if missing.| YYYY-MM-DD | Sketch | (action) | (files) | (outcome) | to .agents/PROJECT.md..agents/sketch.md only when an insight is genuinely reusable._common/OPERATIONAL.md.See _common/AUTORUN.md for the protocol (_AGENT_CONTEXT input, mode semantics, error handling). Sketch-specific _STEP_COMPLETE.Output schema lives in reference/autorun-schema.md.
When input contains ## NEXUS_ROUTING, do not call other agents directly. Return all work via ## NEXUS_HANDOFF.
## NEXUS_HANDOFF## NEXUS_HANDOFF
- Step: [X/Y]
- Agent: Sketch
- Summary: [1-3 lines]
- Key findings / decisions:
- Prompt: [constructed prompt]
- Model: [selected model]
- Parameters: [major parameters]
- Artifacts: [Python script path, metadata path]
- Risks: [policy concern, cost impact]
- Suggested next agent: [Muse | Canvas | Growth] (reason)
- Next action: CONTINUE
Frequently asked questions
Sketch produces reproducible Python code for Gemini image generation, image editing, prompt refinement, and batch asset workflows. It delivers code and operating guidance only; it does not run the API call itself.
The source record exposes this install command: npx skills add https://github.com/simota/agent-skills --skill ".archive/sketch". Inspect the command and pinned source before running it.
Alternatives
coreyhaines31/marketingskills
When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "should I test this," "which version is better," "test two versions," "statistical significance," "how long should I run this test," "growth experiments," "experiment velocity," "experiment backlog," "ICE score," "experimentation program
garrytan/gbrain
End-to-end discipline for turning any large data source (audio libraries, email takeouts, document corpora, chat exports, API dumps) into brain pages at scale. The lifecycle spine: SCHEMA → ACCESS → TRIAL → EVALUATE → IMPROVE → CODIFY → TEST → SKILLIFY → BULK → MONITOR. State is tracked in a durable JSON manifest (see MANIFEST-PATTERN.md) so any crash, session boundary, or subagent fan-out resumes from ground truth instead of memory.
alirezarezvani/claude-skills
App Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store. Use when the user asks about ASO, app store rankings, app metadata, app titles and descriptions, app store listings, app visibility, or mobile app marketing on iOS or Android. Supports keyword research and scoring, competitor keyword analysis, metadata optimization, A/B test planning, launch checklist
dotnet/skills
Migrates .NET test projects from VSTest to Microsoft.Testing.Platform (MTP). Use when user asks to "migrate to MTP", "switch from VSTest", "enable Microsoft.Testing.Platform", "use MTP runner", set OutputType=Exe only for test projects in Directory.Build.props, or mentions EnableMSTestRunner, EnableNUnitRunner, or UseMicrosoftTestingPlatformRunner. USE FOR: MTP behavioral differences vs VSTest (exit code 8, zero tests discovered, --ignore-exit-code, TESTINGPLATFORM_EXITCODE_IGNORE); centralizing