Vision (image input) Guide
You can pass images into
messagesonPOST /v1/chat/completionsand have them analyzed (OpenAI-compatible). There is no extra endpoint — an API key with thellmscope works as-is.
Which scope you need
api.aicu.ai API keys carry scopes (permission groups), and different endpoint families require different ones.
| Scope | What it unlocks | Needed for Vision? |
|---|---|---|
llm | /v1/chat/completions (text and image input alike) | ✅ |
images | /v1/images/* (image generation) | No |
tts | /v1/tts/* and /v1/audio/transcriptions (speech synthesis, transcription) | No |
- Vision is just "pass an image to a chat completion," so the
llmscope alone is enough.imagesis the scope for creating images - Keys issued from the dashboard come with
llm/tts/imagesby default - If a company-issued key (least-privilege) returns
403, the scope is missing. Ask whoever issued it to add it
Calling it (curl, image URL)
curl -X POST https://api.aicu.ai/v1/chat/completions \
-H "Authorization: Bearer $AICU_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "Describe this image."},
{"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}}
]
}],
"max_tokens": 300
}'
The trick is to make content an array instead of a string, and mix in parts of
type: "image_url". messages is forwarded to the provider verbatim, so
OpenAI's Vision spec works exactly as documented.
Calling it (Python, local image as base64)
import base64, os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AICU_API_KEY"],
base_url="https://api.aicu.ai/v1",
)
b64 = base64.b64encode(open("frame.jpg", "rb").read()).decode()
res = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "What's the highlight of this image? One line."},
{"type": "image_url", "image_url": {
"url": f"data:image/jpeg;base64,{b64}",
"detail": "low",
}},
],
}],
max_tokens=300,
)
print(res.choices[0].message.content)
detail:low(low resolution, low cost) /high(tiled, high accuracy) /auto. Uselowwhen analyzing a lot of frames
Supported models and pricing
Measured 2026-09-22 by sending the same small image to each model and asking what it shows.
| Model | Images | Note |
|---|---|---|
gpt-4o-mini | ✅ | Cheapest way to read an image |
gpt-4o | ✅ | |
gpt-5.6-luna | ✅ | Lightest of the 5.6 family |
gpt-5.6-terra | ✅ | Coding and image understanding |
gpt-5.6-sol | ✅ | Same, with the deepest reasoning |
gpt-6-astra | ✅ | |
kimi-k2.7-code | Preview | Declared multimodal, but it returned no answer in this test |
The gpt-5.6 family reading images is what makes screenshot review from a coding agent
(Cline, OpenCode) practical — see Coding agents.
GET /v1/chat/modelsdoes not declare image support. There is nomodalitiesorvisionfield, so tools that auto-detect capabilities assume text only. You have to tell them (OpenCode: declaremodalitiesyourself — see Coding agents)- Passing
image_urlto a text-only model such asdeepseek-v3returns an error - Images are converted to tokens and billed (per OpenAI's rules; one image at
detail: low≈ 85 tokens.gpt-4o-minicounts about 33x the tokens but at a lower unit price, so the cost lands about the same — measured: onelowimage plus a short prompt ≈ 2,900 tokens). For per-model AP rates, seeGET /v1/chat/models
Troubleshooting
| Symptom | Cause |
|---|---|
403 Forbidden | The key lacks the llm scope -> ask whoever issued it to add it |
401 Unauthorized | Key is invalid or expired |
| Image fetch error | The URL is private or requires auth -> send it as base64 (data URL) instead |
OpenAI's official docs (primary source for the compatible spec)
- Images and vision guide: https://platform.openai.com/docs/guides/images-vision
- Chat Completions API reference: https://platform.openai.com/docs/api-reference/chat/create
© 2026 AICU Inc.