gemini-nanobanana
SkillMediaThis Gemini Claude Skill lets your agent generate and edit images through Google's Nano Banana model.
Available today. Use it from your connected AI after setup.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the gemini-nanobanana skill
About this skill
Use this skill when users ask to generate, edit, or compose images with Gemini Nano Banana 2, including text-to-image, image editing, multi-image composition, grounding, and output sizing/saving controls.
What this skill tells your AI
The instructions your AI receives, as published by duotify/githubclawtoolkit in skills/gemini-nanobanana/SKILL.md and read by ahel’s review.
Do this first
- Use a Node.js wrapper (
@google/genai) as the primary flow (multi-turn edits, grounding, advancedgenerationConfig). - Use
scripts/gemini-nanobanana-cli.jsfor quick single-turn generation/editing runs. - Do not teach raw
curl; keep guidance in JS CLI/wrapper form.
Enforce these defaults
- API key source:
GEMINI_NANOBANANA_API_KEYwith fallbackGEMINI_API_KEY. - Default model:
gemini-3.1-flash-image-preview(allow env override viaGEMINI_NANOBANANA_MODEL). - Reference images: support up to 14 total.
- Thinking strength: configurable (
minimal|low|medium|high), default High. - Aspect ratio: default Auto.
- Resolution: default 1K.
- Output mode: default Images only.
- Google Search grounding tool: default Disabled; enable with
--google-searchwhen needed. - Output directory: default
nanobanana-output/, unless the prompt explicitly asks for another location. - 512px rule: send
imageConfig.imageSizeas string"512"in API calls (never numeric512).
Read references intentionally
- Start in
references\image-generation-api.mdfor operational payload rules, model behavior, sizing, thinking, grounding, and limits. - Use
references\sources.mdto verify source provenance and jump to upstream docs. - For canonical API behavior, read:
Output formatting rules
After successful generation, always report results in this format:
✅ 圖片已產出- Each image as a markdown embed with absolute GitHub URL:
 - File metadata: format (
JPEG/PNG), dimensions, file size - Artifact metadata block (enables Telegram relay to send the actual photo):
<!-- githubclaw-artifacts: {"images":[{"branch":"{BRANCH}","path":"{relative_path}"}],"html":[]} -->
Example (single image)
Assuming GITHUB_REPO=test/baoclaw-5, BRANCH=issue-3:
✅ 圖片已產出

- 格式:JPEG · 1408×768 · 757 KB
<!-- githubclaw-artifacts: {"images":[{"branch":"issue-3","path":"issue-3/artifacts/4153431460/matcha-latte-01.jpg"}],"html":[]} -->
Example (multiple images)
✅ 圖片已產出


- 圖 1:JPEG · 1408×768 · 703 KB
- 圖 2:JPEG · 1408×768 · 512 KB
<!-- githubclaw-artifacts: {"images":[{"branch":"issue-3","path":"issue-3/artifacts/4153431460/cute-puppy-01.jpg"},{"branch":"issue-3","path":"issue-3/artifacts/4153431460/cute-puppy-02.jpg"}],"html":[]} -->
Why this format
- GitHub Issue comments require absolute URLs to render images inline (relative paths won't display).
- Telegram relay detects
githubclaw-artifactsmetadata → downloads image via GitHub API → sends as photo. ?raw=trueensures GitHub serves raw image bytes instead of the HTML file viewer.
Execution pattern
- Default to Node.js wrapper flows for regular usage, especially when payload control is needed.
- Quick path (agent runs from
issue-N/; resolve the repo root first):REPO_ROOT=$(git rev-parse --show-toplevel 2>/dev/null || echo "../..")thennode "$REPO_ROOT/.agents/skills/gemini-nanobanana/scripts/gemini-nanobanana-cli.js" --prompt "..."- Add references via repeated
-i/--image(up to 14). - Enable grounding via
--google-searchwhen prompt needs fresh web context.
- The API key (
GEMINI_NANOBANANA_API_KEY/GEMINI_API_KEY) is injected by the workflow environment; do not hardcode it. The CLI reads it automatically from the environment.
Signals
- GitHub stars
- 253
- Forks
- 53
- Last commit
- Sep 2026
Advanced
- Item type
- skill
- Key
gemini-nanobanana- Source
- github.com/duotify/githubclawtoolkit