Image-2 Skill

👤 gpt 📦 v1.0.1 ⭐ 4.6 ⬇️ 711 下載
🎨 設計多媒體 免費 🔑 需 API Key

📖 技能介紹


name: image-2 version: 1.1.0 description: "GPT-4o Image Generation & Editing Skill - Create, edit, transform, and analyze images using GPT-4o native image-2 API. Supports text-to-image, inpainting, outpainting, style transfer, background removal, and intelligent image analysis. Ideal for marketing, product photos, illustrations, UI mockups, and visual content creation." metadata: openclaw: emoji: "🎨" homepage: "https://clawhub.ai/gpt/image-2" always: false skillKey: "image-2" requires: env: - OPENAI_API_KEY primaryEnv: OPENAI_API_KEY install: - kind: node package: openai bins: []


Image-2 Skill

Create, edit, transform, and analyze images with GPT-4o's native image generation API

When to Use This Skill

Use this skill whenever the user needs to: - Generate images from text descriptions ("畫一張...", "生成圖片...", "create an image of...") - Edit existing images with natural language ("把背景去掉", "add a sunset", "換成藍色") - Create variations of an image ("生成幾個變體", "make 4 variations") - Analyze/describe images ("這張圖是什麼", "describe this image", "提取文字") - Remove backgrounds ("去除背景", "remove background") - Style transfer ("變成水彩風格", "make it look like Van Gogh") - Create marketing visuals ("設計海報", "make a social media post") - Product photography ("產品圖", "product shot on white background") - UI/UX mockups ("介面設計", "app mockup", "website screenshot")

Core Workflows

Workflow 1: Text-to-Image Generation

When the user describes an image they want to create:

  1. Enhance the prompt — Automatically add quality boosters:
  2. Append professional photography/art terms based on context
  3. Add lighting, composition, and mood details if not specified
  4. Specify output format and dimensions if needed

  5. Call the API — Use generateImage() with the enhanced prompt: javascript const result = await generateImage(enhancedPrompt, { size, quality, style });

  6. Save and present — Download the image to the project directory and show the user:

  7. Save to ./generated-images/ by default
  8. Return the file path and a brief description

Workflow 2: Image Editing

When the user wants to modify an existing image:

  1. Locate the source image — Find the image file path from the conversation context
  2. Parse the edit intent — Understand what changes the user wants
  3. Call the edit API — Use editImage() with the source and instruction: javascript const result = await editImage(imagePath, editInstruction, { mask: maskPath });
  4. Present the result — Show the edited image and describe what changed

Workflow 3: Image Analysis

When the user asks about an image:

  1. Get the image — From file path or URL
  2. Analyze with GPT-4o Vision — Use describeImage(): javascript const result = await describeImage(imageSource, question);
  3. Report findings — Present the analysis in a structured format

Workflow 4: Batch Generation

When the user needs multiple images:

  1. Parse the batch request — Understand variations needed
  2. Generate in parallel — Call generateImage() for each variant
  3. Organize results — Save with descriptive filenames

Prompt Enhancement Rules

When generating images, automatically enhance the user's prompt:

Quality Boosters (always append unless user specifies quality)

professional quality, high resolution, sharp details

Context-Based Additions

User Intent Auto-Add
Product photo "studio lighting, clean background, commercial photography"
Portrait "professional portrait photography, natural lighting"
Social media "eye-catching, vibrant colors, modern design"
Illustration "detailed illustration, professional artist quality"
Logo/branding "clean vector style, scalable, minimal details"
Architecture "architectural visualization, realistic rendering"
Food "appetizing, food styling, professional food photography"
UI mockup "clean design, modern interface, pixel-perfect"

Size Recommendations

Use Case Recommended Size
Social media post 1024x1024 (square)
Story/vertical 1024x1792
Banner/landscape 1792x1024
Product listing 1024x1024
Presentation 1792x1024
Wallpaper 1792x1024

Style Presets

Quick style references for common requests:

Preset Name Style Description
product Clean white background, studio lighting, commercial photography
lifestyle Natural setting, warm lighting, aspirational mood
minimalist Simple composition, negative space, clean lines
vintage Retro color grading, film grain, nostalgic mood
futuristic Neon accents, dark background, sci-fi aesthetic
watercolor Soft edges, pastel palette, artistic brush strokes
3d-render Octane render, realistic materials, dramatic lighting
anime Japanese animation style, vibrant, expressive
sketch Pencil drawing, hand-drawn, artistic
flat-design Vector style, bold colors, geometric shapes

API Reference

generateImage(prompt, options)

Generate a new image from text description.

Parameters: - prompt (string) — Image description (auto-enhanced by this skill) - options (object): - size1024x1024 | 1024x1792 | 1792x1024 (default: 1024x1024) - qualitystandard | hd (default: standard) - stylevivid | natural (default: vivid) - modelgpt-image-2 | dall-e-3 (default: gpt-image-2) - saveTo — File path to save the image (default: ./generated-images/)

Returns: { success, url, localPath, revisedPrompt }

editImage(imagePath, prompt, options)

Edit an existing image with natural language instructions.

Parameters: - imagePath (string) — Path to the source image - prompt (string) — Edit instruction - options (object): - mask — Path to mask image (white = edit area, black = keep) - size — Output size - modelgpt-image-2 | dall-e-3 (default: gpt-image-2)

Returns: { success, url, localPath }

generateVariations(imagePath, options)

Generate creative variations of an existing image.

Parameters: - imagePath (string) — Path to the source image - options (object): - count — Number of variations 1-4 (default: 2) - size — Output size

Returns: { success, variations: [{ url, localPath }] }

describeImage(imageSource, question)

Analyze an image using GPT-4o Vision.

Parameters: - imageSource (string) — File path or URL of the image - question (string|null) — Specific question about the image (default: general description)

Returns: { success, description }

downloadImage(url, savePath)

Download a generated image to local storage.

Parameters: - url (string) — Image URL from generation API - savePath (string|null) — Local file path (default: auto-generated in ./generated-images/)

Returns: { success, localPath }

Error Handling

Error Cause Resolution
Invalid API key OPENAI_API_KEY not set or invalid Check environment variable
Content policy violation Prompt violates safety guidelines Rephrase the prompt
Rate limit exceeded Too many requests Wait and retry with backoff
Image too large Source image exceeds size limit Resize to under 4MB
Timeout Generation took too long Simplify prompt or retry

Best Practices

  1. Always enhance prompts — Don't pass raw user input directly to the API
  2. Save locally — Download generated images; URLs expire after 1 hour
  3. Use appropriate sizes — Match the output size to the use case

    推薦訪問7w4.net獲取更多AI技能。

  4. Prefer gpt-image-2 — Better quality and text rendering than dall-e-3
  5. Batch thoughtfully — Generate 2-4 images max per request to avoid rate limits
  6. Describe edits clearly — Be specific about what to change and where

Changelog

v1.1.0

  • Added GPT-4o native image generation support (gpt-image-2 model)
  • Added automatic prompt enhancement workflow
  • Added image download and local save functionality
  • Added style presets for quick reference
  • Added batch generation workflow
  • Improved error handling and documentation

v1.0.0

  • Initial release with DALL-E 3 support
  • Basic generate, edit, variations, and describe functions

Tags: image-generation AI-art GPT-4o image-2 gpt-image-2 visual-creation marketing product-photos illustration design openai dall-e image-editing background-removal style-transfer ui-mockup

🤖 AI 評測

這個 Skill 質量不錯,文件寫得很詳細,圖片生成流程完整,使用起來比較省心。初始化配置方便,支援批次生成和參考圖上傳,輪詢任務帶進度顯示。價格策略也比較靈活。不足之處是依賴項只有一個 requests,缺少使用驗證機制,新手遇到問題時可能需要自行排查。

📊 多維度評分

適應性4.4
規範性4.5
有效性4.7
可靠性4.3
可信度5

📁 包含檔案 (7 個)

📄 README.md 5.5 KB
📄 SKILL.md 9 KB
📄 _meta.json 126 B
📄 examples/prompts-gallery.md 6.5 KB
📄 examples/quick-starts.md 6.2 KB
📄 package.json 622 B
📄 scripts/image-generator.js 18.4 KB