Skip to main content

GPT-Image-2.5 画像生成ガイド(日本語訳・Part 1)

· 8 min read
AICU API Team
api.aicu.ai

OpenAI 公式「Image generation」ガイド(原文)の日本語訳です。 Part 1 は 概要/Image API と Responses API の選び方/画像生成/マルチターン/ストリーミングまで。 編集・出力のカスタマイズ・対応モデルなどは Part 2 で続けます。

api.aicu.ai は Image API(/v1/images/generations / /v1/images/edits)を OpenAI 互換で提供しています。 base_url を https://api.aicu.ai/v1、キーを aicu_live_... にすれば、下の Image API の例はそのまま動きます。 Responses API(会話内で画像を生成するツール)の例は OpenAI 直の呼び出しです。

概要​

この API では、テキストプロンプトから gpt-image-2.5-sunburst と gpt-image-2.5-flare で画像を生成・編集できます。 編集の精度が最重要なワークフローには Sunburst、速くて高品質な日常の生成には Flare を選びます。画像生成には 2 つの API があります。

Image API​

Image API は、役割の異なる 2 つのエンドポイントを提供します。

  • Generations: テキストプロンプトから画像をゼロから生成
  • Edits: 既存の画像を新しいプロンプトで部分的または全体的に変更

Responses API​

Responses API は、会話やマルチステップのフローの一部として画像を生成できます。 画像生成を組み込みツールとして扱い、文脈内で画像の入出力を受け付けます。

Image API に比べて次が加わります。

  • マルチターン編集: プロンプトで高忠実度の編集を繰り返し重ねられる
  • 柔軟な入力: バイト列だけでなく、画像の File ID も入力にできる

どちらの API を使うか​

  • 1 つのプロンプトから 1 枚を生成・編集するだけなら Image API。
  • 会話しながら編集できる画像体験を作るなら Responses API。

Image API では model に gpt-image-2.5-sunburst または gpt-image-2.5-flare を直接指定します。 Responses API では、トップレベルに対応するメインラインモデルを選び、画像生成ツールの model フィールドに gpt-image-2.5-sunburst / gpt-image-2.5-flare を指定します。

どちらの API でも、品質・サイズ・形式・圧縮を調整して出力をカスタマイズできます(Part 2)。

info

GPT Image モデルを責任を持って使うため、利用前に開発者コンソールで API Organization Verification を完了する必要がある場合があります。

画像を生成する​

画像生成エンドポイントでテキストプロンプトから生成するか、 Responses API の画像生成ツールで会話の一部として生成します。 出力のカスタマイズ(サイズ・品質・形式・圧縮)は Part 2 を参照してください。

n パラメータで 1 リクエストに複数枚を生成できます(既定は 1 枚)。

Image API — 画像を生成する(api.aicu.ai では base_url を差し替えるだけ)

from openai import OpenAI
import base64

client = OpenAI() # api.aicu.ai なら OpenAI(base_url="https://api.aicu.ai/v1", api_key="aicu_live_...")

prompt = """
A children's book drawing of a veterinarian using a stethoscope to
listen to the heartbeat of a baby otter.
"""

result = client.images.generate(model="gpt-image-2.5-sunburst", prompt=prompt)

image_base64 = result.data[0].b64_json
image_bytes = base64.b64decode(image_base64)

with open("otter.png", "wb") as f:
f.write(image_bytes)
curl -X POST "https://api.aicu.ai/v1/images/generations" \
-H "Authorization: Bearer $AICU_API_KEY" \
-H "Content-type: application/json" \
-d '{
"model": "gpt-image-2.5-sunburst",
"prompt": "A children'\''s book drawing of a veterinarian using a stethoscope to listen to the heartbeat of a baby otter."
}' | jq -r '.data[0].b64_json' | base64 --decode > otter.png
import OpenAI from "openai";
import fs from "fs";
const openai = new OpenAI();

const result = await openai.images.generate({
model: "gpt-image-2.5-sunburst",
prompt: "A children's book drawing of a veterinarian using a stethoscope to listen to the heartbeat of a baby otter.",
});

fs.writeFileSync("otter.png", Buffer.from(result.data[0].b64_json, "base64"));

(原文には Go / Java / C# / Ruby / CLI の例もあります。)

Responses API — 画像を生成する(OpenAI 直)

from openai import OpenAI
import base64

client = OpenAI()

response = client.responses.create(
model="gpt-6-astra",
input="Generate an image of gray tabby cat hugging an otter with an orange scarf",
tools=[{"type": "image_generation", "model": "gpt-image-2.5-sunburst"}],
)

image_data = [
output.result
for output in response.output
if output.type == "image_generation_call"
]

if image_data:
with open("otter.png", "wb") as f:
f.write(base64.b64decode(image_data[0]))

マルチターンの画像生成​

Responses API では、画像生成の呼び出し結果を文脈に含める(画像 ID だけでも可)か、 previous_response_id パラメータを使って、 複数ターンにわたる画像生成の会話を組み立てられます。プロンプトを練り直し、新しい指示を加え、会話の進行に合わせて画像を育てていけます。

対応するツールモデルは、新しい画像を生成するか、文脈内の画像を編集するかを選べます。任意の action パラメータで制御します: action: "auto"(モデルに任せる・既定)、action: "generate"(常に新規生成)、action: "edit"(文脈に画像があるとき編集を強制)。

action で生成を強制する

response = client.responses.create(
model="gpt-6-astra",
input="Generate an image of gray tabby cat hugging an otter with an orange scarf",
tools=[
{"type": "image_generation", "model": "gpt-image-2.5-sunburst", "action": "generate"}
],
)
note

文脈に画像が無いのに edit を強制するとエラーになります。生成と編集の判断をモデルに任せるなら action は auto のままにします。

previous_response_id を使う

response = client.responses.create(
model="gpt-6-astra",
input="Generate an image of gray tabby cat hugging an otter with an orange scarf",
tools=[{"type": "image_generation", "model": "gpt-image-2.5-sunburst"}],
)
# ...最初の画像を保存...

# フォローアップ
response_fwup = client.responses.create(
model="gpt-6-astra",
previous_response_id=response.id,
input="Now make it look realistic",
tools=[{"type": "image_generation", "model": "gpt-image-2.5-sunburst"}],
)

画像 ID を使う

response = client.responses.create(
model="gpt-6-astra",
input="Generate an image of gray tabby cat hugging an otter with an orange scarf",
tools=[{"type": "image_generation", "model": "gpt-image-2.5-sunburst"}],
)
image_generation_calls = [o for o in response.output if o.type == "image_generation_call"]

# フォローアップ: 前の画像生成呼び出しの id を文脈に渡す
response_fwup = client.responses.create(
model="gpt-6-astra",
input=[
{"role": "user", "content": [{"type": "input_text", "text": "Now make it look realistic"}]},
{"type": "image_generation_call", "id": image_generation_calls[0].id},
],
tools=[{"type": "image_generation", "model": "gpt-image-2.5-sunburst"}],
)

結果のイメージ: 「灰色のトラ猫がオレンジのマフラーのカワウソを抱きしめる」→ 2 ターン目「写実的にして」で、同じ構図のまま写実化された画像が返ります。

ストリーミング​

Responses API と Image API はどちらも画像生成のストリーミングに対応しています。生成途中の部分画像を受け取れるので、体験がインタラクティブになります。

partial_images パラメータで 0〜3 枚の部分画像を受け取れます。

  • partial_images を 0 にすると、最終画像だけを受け取ります。
  • 1 以上でも、最終画像の生成が速い場合は指定枚数より少ない部分画像しか届かないことがあります。

Responses API — 画像をストリームする

stream = client.responses.create(
model="gpt-6-astra",
input="Draw a gorgeous image of a river made of white owl feathers, snaking its way through a serene winter landscape",
stream=True,
tools=[{"type": "image_generation", "model": "gpt-image-2.5-sunburst", "partial_images": 2}],
)

for event in stream:
if event.type == "response.image_generation_call.partial_image":
save_base64_image(f"river-partial-{event.partial_image_index}.png", event.partial_image_b64)
elif event.type == "response.completed":
image_data = [o.result for o in event.response.output if o.type == "image_generation_call"]
if image_data:
save_base64_image("river-final.png", image_data[0])

Image API — 画像をストリームする

stream = client.images.generate(
prompt="Draw a gorgeous image of a river made of white owl feathers, snaking its way through a serene winter landscape",
model="gpt-image-2.5-sunburst",
stream=True,
partial_images=2,
)

for event in stream:
if event.type == "image_generation.partial_image":
with open(f"river{event.partial_image_index}.png", "wb") as f:
f.write(base64.b64decode(event.b64_json))

結果のイメージ: 部分画像 1 → 部分画像 2 → 最終画像、の順にだんだん精細になっていきます。

:::tip api.aicu.ai での長い生成 api.aicu.ai では、100 秒を超える生成は stream の X-AICU-Image-Id で id を受け取り、GET /v1/images/status/:id でポーリングして回収できます(Cloudflare のエッジ制限を越えるため)。status/:id は API キーが要ります。Cloudflare Pages Function から呼ぶ場合、stream: true で応答本文を途中で捨てる(body.cancel)と 502 になった実例があるので、Pages Function では stream を使わず同期応答の id を控えるか、flare のように 100 秒に収まるモデルを既定にしてください。 :::


Part 2(出力のカスタマイズ・画像の編集・マスク・透過・対応モデル・制約)に続きます。原文: Image generation。