AICU LLM API Documentation
OpenAI-compatible chat completions API. An API key (Bearer auth) is required. Issue an
aicu_live_key from the dashboard.
Base URL
https://api.aicu.ai/v1
Quick Start
curl -X POST https://api.aicu.ai/v1/chat/completions \
-H "Authorization: Bearer $AICU_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v3",
"messages": [{"role": "user", "content": "Hello!"}],
"max_tokens": 100
}'
🖼️ Image input (Vision) works too. Make
contentan array and pass animage_url→ Vision guide
POST /v1/chat/completions
OpenAI-compatible chat completions endpoint.
Request:
{
"model": "deepseek-v3",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello!"}
],
"max_tokens": 1000,
"temperature": 0.7,
"stream": false
}
Parameters:
| Name | Type | Required | Description |
|---|---|---|---|
| model | string | Yes | Model ID |
| messages | array | Yes | Array of messages |
| max_tokens | number | No | Maximum tokens to generate (default: 4096) |
| temperature | number | No | Sampling temperature, 0-2 (default: 1) |
| stream | boolean | No | Stream the response (default: false) |
Response:
{
"id": "gen-xxx",
"object": "chat.completion",
"created": 1234567890,
"model": "deepseek-v3",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello! How can I help you today?"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 10,
"completion_tokens": 15,
"total_tokens": 25
}
}
Response Headers:
X-AICU-Model: Model usedX-AICU-Provider: Provider (openrouter/groq/aicu)X-AICU-Tokens: Tokens consumedX-AICU-AP-Cost: AP chargedX-AICU-Latency-Ms: Latency
GET /v1/models — every model, one list
GET /v1/models is the single source of truth for everything the platform serves: chat, image, speech (tts) and transcription (stt). Every entry has the same shape — type, the endpoint to call, stage, pricing with an explicit unit, docs_url, pricing_url and updated_at — and the list envelope carries docs, pricing_url and the latest updated_at. GET /v1/models/:id returns one model (OpenAI-style retrieve; image aliases such as flare work).
curl https://api.aicu.ai/v1/models # everything
curl https://api.aicu.ai/v1/models/deepseek-v3
{
"object": "list",
"docs": "https://api.aicu.ai/docs",
"pricing_url": "https://api.aicu.ai/docs/pricing",
"updated_at": "2026-09-11",
"data": [
{ "id": "deepseek-v3", "type": "chat", "endpoint": "POST /v1/chat/completions", "stage": "beta",
"pricing": { "unit": "per_1k_tokens", "ap_in_per_1k": 9, "usd_in_per_1k": 0.0009, "ap_out_per_1k": 27, "usd_out_per_1k": 0.0027 },
"docs_url": "https://api.aicu.ai/docs/llm-api", "updated_at": "2026-08-01" },
{ "id": "gpt-image-2.5-flare", "type": "image", "endpoint": "POST /v1/images/generations",
"pricing": { "unit": "per_image", "ap_from": 700, "ap_by_size_quality": { "1024x1024": { "low": 700, "medium": 1600, "high": 6400 } } } }
]
}
stage is beta until production starts on 2026-09-23, then production; preview marks models we may still change.
GET /v1/chat/models
The same array filtered to type=chat. Kept for compatibility; the fields below are unchanged.
curl https://api.aicu.ai/v1/chat/models
Response:
{
"object": "list",
"data": [
{
"id": "deepseek-v3",
"owned_by": "openrouter",
"pricing": {"ap_per_1k_tokens": 14, "ap_in_per_1k": 9, "ap_out_per_1k": 27},
"stage": "beta",
"description": "汎用、コスパ良好"
}
]
}
Available Models
Pricing tiers:
- 🟢 Free — permanently free (rate limits apply)
- 🟡 Free for a limited time — α / preview stage. No real billing through 2026-10-31 (Kumamoto Earthquake 2026 relief program). The AP figures in the table are the rates planned for after GA
- 💠 Paid (planned rates) — the AP figures are reference rates that apply from GA onward
The platform as a whole is currently in alpha. None of the tiers above bill end users today (usage is still logged). Rates are given as "input / output AP per 1K tokens". 1 USD = 10,000 AP (1 AP = $0.0001).
:::caution Amounts are approximate and rates get revised
The AP figures below and every amount in this document are approximate (AP→USD is fixed at 10,000 AP = $1, but yen conversion moves with the exchange rate).
Per-model AP rates are subject to price revisions. This table is maintained by hand, so always base
estimates and invoices on the current values from GET /v1/chat/models.
:::
🟢 Free (Groq — stays free after GA)
| Model | Provider | Rate | Best for |
|---|---|---|---|
groq-llama-3.3-70b | Groq | Free | High quality, rate-limited |
groq-llama-3.1-8b | Groq | Free | Fast, rate-limited |
🟡 Free for a limited time (α / preview — through 2026-10-31)
| Model | Provider | in / out (AP/1K) | Best for |
|---|---|---|---|
gpt-4o | OpenAI | 76 / 300 | High-quality multimodal |
gpt-4o-mini | OpenAI | 5 / 18 | Lightweight and fast |
gpt-5.6-sol | OpenAI | 150 / 900 | Top of the line. Hard reasoning, long context |
gpt-5.6-terra | OpenAI | 80 / 450 | Half the price of sol. The everyday workhorse |
gpt-5.6-luna | OpenAI | 30 / 180 | Light and fast. Chat and summarization |
gpt-5.4-mini | OpenAI | 30 / 140 | Inexpensive. Routine generation and extraction |
gpt-5.4-nano | OpenAI | 6 / 40 | Cheapest. Short jobs like classification and scoring |
sakura-kimi | Sakura AI Engine | 12 / 60 | Kimi-K2.6, inferred in Japan. Your data is not used for training |
kimi-k2.7-code | Sakura AI Engine | 11 / 101 | Coding-focused. OpenCode / Cline compatible |
💠 Paid (beta — planned rates)
The Gemini models moved to the rates below on 2026-08-01: the default
gemini-flashnow maps to the lighter 2.5 Flash-Lite, and a newgemini-3.6-flashcovers the latest model.
| Model | Provider | in / out (AP/1K) | Best for |
|---|---|---|---|
deepseek-v3 | OpenRouter | 9 / 27 | General-purpose, good value |
gemini-flash | OpenRouter | 3 / 12 | Lightweight and cheap (Gemini 2.5 Flash-Lite). From 2026-08-01 |
gemini-3.6-flash | OpenRouter | 45 / 225 | Latest and highest quality (Gemini 3.6 Flash). Added 2026-08-01 |
llama-3.1-8b | OpenRouter | 1 / 2 | Lightweight |
llama-3.1-70b | OpenRouter | 4 / 10 | High quality |
qwen3-32b | OpenRouter | 4 / 10 | Large context |
The previous generation —
gpt-5/gpt-5-mini/gpt-5-nano— stays available for now, but new work should target thegpt-5.6/gpt-5.4families. The always-current list isGET /v1/chat/models; this table is updated by hand and may lag behind.
Notes on reasoning models
Models that GET /v1/chat/models reports with reasoning: true (the gpt-5 family and similar) bill thinking tokens as well, and a small max_tokens can be consumed entirely by reasoning, leaving an empty body (content: "").
- Pass
reasoning_effort(minimal/low/medium/high) to control how much the model thinks — it is forwarded as-is, OpenAI-compatible. - When it is omitted for a
gpt-5-family model, api.aicu.ai applies a default ofminimalso that even short prompts come back with a body. An explicit value always wins. - For short generations where cost and predictability matter more, a non-reasoning model such as
gpt-4o-miniis the reliable choice. - What the levels mean upstream: OpenAI — Reasoning. With
toolsin the request the effort is dropped, because reasoning and tool calling cannot both run — see Coding agents.
curl -X POST https://api.aicu.ai/v1/chat/completions \
-H "Authorization: Bearer $AICU_API_KEY" -H "Content-Type: application/json" \
-d '{"model":"gpt-5-mini","messages":[{"role":"user","content":"140字で告知文を書いて"}],"max_tokens":600,"reasoning_effort":"minimal"}'
OpenAI SDK Compatibility
from openai import OpenAI
client = OpenAI(
base_url="https://api.aicu.ai/v1",
api_key="aicu_live_xxx" # API key required (issue one in the dashboard)
)
response = client.chat.completions.create(
model="deepseek-v3",
messages=[{"role": "user", "content": "Hello!"}]
)
print(response.choices[0].message.content)
import OpenAI from 'openai';
const client = new OpenAI({
baseURL: 'https://api.aicu.ai/v1',
apiKey: 'aicu_live_xxx', // API key required (issue one in the dashboard)
});
const response = await client.chat.completions.create({
model: 'deepseek-v3',
messages: [{ role: 'user', content: 'Hello!' }],
});
Where the upstream specification lives
api.aicu.ai is OpenAI-compatible, which is only a useful promise if you can see what it is compatible with. These are the upstream documents this API follows. When our page and the upstream page disagree about a parameter, the upstream page describes the parameter and this site describes what we accept — if you hit a difference, tell us at aicu.ai/contact.
| Topic | Upstream reference | How we differ |
|---|---|---|
| Request and response shape | OpenAI — Create chat completion | Same shape. Extra X-AICU-* response headers carry AP cost, model and latency |
| Tool / function calling | OpenAI — Function calling | Forwarded as-is. With a reasoning model we drop the reasoning effort so the tool call succeeds (see Coding agents) |
Reasoning models, reasoning_effort | OpenAI — Reasoning | Forwarded as-is. We default to minimal when you omit it, so short prompts still return a body |
| Image input | OpenAI — Images and vision | Same shape. Which models accept images is measured on our side — see Vision |
| Groq-hosted models | Groq — OpenAI compatibility | Groq's own limits apply on top of ours |
What is ours, not upstream: AP billing and the X-AICU-AP-Cost header, the model catalogue
(GET /v1/chat/models), key scopes, and the hold-then-settle charging described in
Pricing. Those are not in any upstream document.
Credits (billing)
- 10,000 AP = $1 (USD-denominated; 1 AP = $0.0001. Our own costs are in USD, so the unit follows)
- Yen figures are approximate and move with the exchange rate — treat every amount in these docs as an estimate
- Per-model AP rates are subject to price revisions. See
GET /v1/chat/modelsfor current values - Groq models: free (rate limits apply)
- Every endpoint requires an API key (Bearer auth)
Rate Limits
| Provider | Limit |
|---|---|
| Groq | 30 RPM |
| OpenRouter | Depends on your plan |
Troubleshooting
A 403 is about your key, not your HTTP client
Until 2026-08-23, clients that kept their default User-Agent — Python's standard-library
urllib was the usual suspect — could be blocked as bots and get a 403 (Cloudflare error 1010).
That is fixed. api.aicu.ai is exempt from the bot check, and a default urllib request now
returns 200 (verified 2026-09-22). You do not need to set a User-Agent.
If you still get a 403, it is about the key itself:
| Body | Meaning |
|---|---|
{"error": "API key does not have required scope: llm", "required_scope": "llm", "current_scopes": [...]} | The key lacks that surface. Add it with 権限変更 at dashboard/keys — no reissue needed |
{"error": {"code": "unauthorized", ...}} | The key is revoked or does not exist |
Support
- Skill: https://api.aicu.ai/skills/llm
- Dashboard: https://api.aicu.ai/dashboard
- Contact: https://aicu.ai/contact
© 2026 AICU Inc.