Skip to main content

Vision (image input) Guide

You can pass images into messages on POST /v1/chat/completions and have them analyzed (OpenAI-compatible). There is no extra endpoint — an API key with the llm scope works as-is.

Which scope you need​

api.aicu.ai API keys carry scopes (permission groups), and different endpoint families require different ones.

ScopeWhat it unlocksNeeded for Vision?
llm/v1/chat/completions (text and image input alike)✅
images/v1/images/* (image generation)No
tts/v1/tts/* and /v1/audio/transcriptions (speech synthesis, transcription)No
  • Vision is just "pass an image to a chat completion," so the llm scope alone is enough. images is the scope for creating images
  • Keys issued from the dashboard come with llm / tts / images by default
  • If a company-issued key (least-privilege) returns 403, the scope is missing. Ask whoever issued it to add it

Calling it (curl, image URL)​

curl -X POST https://api.aicu.ai/v1/chat/completions \
-H "Authorization: Bearer $AICU_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "Describe this image."},
{"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}}
]
}],
"max_tokens": 300
}'

The trick is to make content an array instead of a string, and mix in parts of type: "image_url". messages is forwarded to the provider verbatim, so OpenAI's Vision spec works exactly as documented.

Calling it (Python, local image as base64)​

import base64, os
from openai import OpenAI

client = OpenAI(
api_key=os.environ["AICU_API_KEY"],
base_url="https://api.aicu.ai/v1",
)

b64 = base64.b64encode(open("frame.jpg", "rb").read()).decode()
res = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "What's the highlight of this image? One line."},
{"type": "image_url", "image_url": {
"url": f"data:image/jpeg;base64,{b64}",
"detail": "low",
}},
],
}],
max_tokens=300,
)
print(res.choices[0].message.content)
  • detail: low (low resolution, low cost) / high (tiled, high accuracy) / auto. Use low when analyzing a lot of frames

Supported models and pricing​

Measured 2026-09-22 by sending the same small image to each model and asking what it shows.

ModelImagesNote
gpt-4o-mini✅Cheapest way to read an image
gpt-4o✅
gpt-5.6-luna✅Lightest of the 5.6 family
gpt-5.6-terra✅Coding and image understanding
gpt-5.6-sol✅Same, with the deepest reasoning
gpt-6-astra✅
kimi-k2.7-codePreviewDeclared multimodal, but it returned no answer in this test

The gpt-5.6 family reading images is what makes screenshot review from a coding agent (Cline, OpenCode) practical — see Coding agents.

  • GET /v1/chat/models does not declare image support. There is no modalities or vision field, so tools that auto-detect capabilities assume text only. You have to tell them (OpenCode: declare modalities yourself — see Coding agents)
  • Passing image_url to a text-only model such as deepseek-v3 returns an error
  • Images are converted to tokens and billed (per OpenAI's rules; one image at detail: low ≈ 85 tokens. gpt-4o-mini counts about 33x the tokens but at a lower unit price, so the cost lands about the same — measured: one low image plus a short prompt ≈ 2,900 tokens). For per-model AP rates, see GET /v1/chat/models

Troubleshooting​

SymptomCause
403 ForbiddenThe key lacks the llm scope -> ask whoever issued it to add it
401 UnauthorizedKey is invalid or expired
Image fetch errorThe URL is private or requires auth -> send it as base64 (data URL) instead

OpenAI's official docs (primary source for the compatible spec)​

© 2026 AICU Inc.