Skip to main content
info
The feature described on this page is in beta.

Pricing

AICU API is operated by AICU Inc. All prices are in USD. Usage is billed in prepaid credits called AP (AICU Points): 1 USD = 10,000 AP.

  • Prepaid, budget-safe: you buy credits up front; when they run out, calls stop. No surprise bills.
  • Errors are never billed. Only successful (HTTP 200) responses consume credits.
  • Cache hits are free. Repeating the same input costs nothing when the cache can serve it.
  • During the alpha/beta period, new accounts receive free trial credits (daily login bonus + first-use bonus) — you can evaluate every API below without payment. Credit purchase opens with the public beta on the dashboard.

Business customers who need a Japanese qualified invoice (適格請求書) can purchase via our reseller AICU Japan K.K. — contact info@aicu.jp.

Text to Speech — /v1/audio/speech​

High-quality character voices (AiCuty) with Japanese pronunciation enhancement, plus raw ElevenLabs voice passthrough.

MetricRate
Text to speechPer 1,000 characters of input text, not by audio length. 200 AP per 1,000 characters from 2026-09-09, stepping up every Wednesday on a schedule announced in advance, reaching 2,475 AP on 2027-04-07
Modelseleven_multilingual_v2 (default, up to 10,000 chars/request) / eleven_v3 (expressive tags, up to 5,000 chars)

Estimate before you call: POST /v1/tts/estimate.

LLM Chat — /v1/chat/completions​

OpenAI-compatible. Rates are per 1,000 tokens (snapshot 2026-08-18 — the live source of truth for every model and price, including images and audio, is GET /v1/models; each entry carries its pricing, docs_url and updated_at):

ModelUSD / 1K tokensAP / 1K tokens
groq-llama-3.3-70bFree (alpha)0
groq-llama-3.1-8bFree (alpha)0
llama-3.1-8b$0.00022
gpt-5-nano / qwen3-32b$0.00055
gemini-flash / llama-3.1-70b$0.00066
gpt-4o-mini$0.00088
deepseek-v3 / gpt-5.4-nano$0.001414
gpt-5-mini$0.002323
gpt-5.4-mini$0.00550
sakura-kimi / gpt-5.6-luna$0.00660
gemini-3.6-flash$0.00990
kimi-k2.7-code$0.0101101
gpt-5$0.0113113
gpt-4o$0.0132132
gpt-5.6-terra$0.016160
gpt-5.6-sol$0.03300

Image Generation — /v1/images/generations​

Character-consistent generation with reference images. The charge depends on the model, the quality level and the output size — not on the prompt.

Measured on 2026-09-21 with gpt-image-2.5-flare. Every response carries the exact figure in X-AICU-AP-Cost, so you never have to infer it.

quality (1024×1024)AP≈ USDTypical latency
low700$0.0710 s
medium1,600$0.1612 s
high6,400$0.6422 s
xhigh11,400$1.1424 s
max25,600$2.5641 s

Size changes the price in steps rather than continuously. At medium:

SizePixelsAP
1024×10241.05 M1,600
1536×10241.57 M2,300
1920×10882.09 M2,300

A square image is the cheapest in absolute terms. Wider canvases cost more per image but less per pixel — and 1536×1024 and 1920×1088 land in the same step, so if you want a wide image you may as well take the larger one. At high and above the relationship inverts: a non-square image is cheaper than the square one (high 5,100 vs 6,400; xhigh 9,200 vs 11,400; max 20,400 vs 25,600).

Generations that take more than 60 seconds​

Some generations take longer than 60 seconds — high quality settings (xhigh / max), requests with a reference image, and large sizes.

These are not meant to be awaited inside an app. Most HTTP clients give up after 30–60 seconds, and Cloudflare's edge cuts the connection at 100 seconds.

The result of a generation that runs past 60 seconds is delivered by email (the image attached, to the key owner; also to notify_email if you set it). You can decline with "notify_on_timeout": false, but then you have one fewer way to receive the result — keep the id from the response (or the X-AICU-Image-Id header) and fetch it with GET /v1/images/status/{id}.

If you are building an app that shows the image right away, use quality low / medium / high at 1024x1024. GET /v1/images/models reports typical_latency_s_by_quality, so you can check the expected time before you call.

Video generation is asynchronous from the start: POST /v1/video/generate returns 202 with an id, and you poll GET /v1/video/status/{id}.

xhigh and max take longer than a synchronous request usually survives — see Generations that take more than 60 seconds below.

Other models: gpt-image-2 (slower and more expensive than 2.5 at comparable settings), gpt-image-2.5-sunburst (finer editing control, longer generation), nano-banana (fast, supports aspect_ratio), sd3.5-large. Current list: GET /v1/images/models.

Transcription — /v1/audio/transcriptions​

Whisper-based speech-to-text (whisper-large-v3-turbo), OpenAI-compatible: prompt for proper nouns, verbose_json timestamps, srt / vtt output. Up to 30 MB / 30 minutes per request. Agent-readable spec: https://api.aicu.ai/skills/stt

MetricRate
Audio transcription (until 2026-09-09 05:00 UTC)1 AP per 10 seconds of audio (rounded up) — 1 hour ≈ 360 AP ≈ $0.04
Audio transcription (from 2026-09-09 05:00 UTC)3 AP per 10 seconds of audio — 1 hour ≈ 1,080 AP ≈ $0.11 (notice, cost×3 policy)

Billed on the measured duration when response_format=verbose_json. Other formats have no duration, so we estimate from the audio we received and from the transcript length, and bill on whichever is larger. The charge is returned in X-AICU-AP-Cost on every response.

How a charge is calculated​

For chat completions we hold a conservative maximum before calling the model, then settle to the real cost and refund the difference to the same balance it came from. You will see the balance dip and come back. This is normal.

The hold is the estimated input plus the maximum possible output, converted to AP at that model's current rate:

PartHow it is estimated
InputThe UTF-8 byte count of the request body (messages and tools) is treated as the token count. GPT-family models use byte-level BPE, so the real token count can never exceed the byte count — the hold always errs high. Images count as 20,000 tokens each (up to 20 per request)
OutputAs if max_tokens were used in full. If you omit max_tokens, the model's own maximum is used

The settlement replaces the hold with the actual usage once the response is complete. Anything over-held is returned, and it returns to wherever it was taken from (credits taken from TAP go back to TAP). If a streamed response does not report usage, the cost is estimated from the UTF-8 byte count of the text that was actually emitted, capped at the hold.

You are never charged more than the hold. The authoritative figure for any single call is the X-AICU-AP-Cost header on the response.

Two practical consequences:

  • Omitting max_tokens produces the largest possible hold, because the model's own maximum is assumed. Setting a realistic max_tokens lowers the hold, so a large request is less likely to be rejected with 402 insufficient_ap when your balance is close to the line. If your balance is low, set max_tokens explicitly — this is the single most effective thing you can do
  • Image generation is charged on a fixed per-request price (see above), not on a hold

How to buy​

  1. Sign in at api.aicu.ai/dashboard (email OTP — an account is created on first sign-in)
  2. Get your API key at dashboard/keys
  3. Purchase credits on the dashboard billing page (opens with public beta; alpha accounts run on free trial credits)

Payments are processed by Stripe. Terms: Terms of Service / Privacy Policy.


Pricing at a glance​

AICU API is operated by AICU Inc. Prices are in USD, and usage is billed in prepaid AP credits (1 USD = 10,000 AP).

  • Prepaid and budget-safe: when your credits run out, calls stop. No unexpected bills
  • Errors are never billed (only 200 responses are charged)
  • Cache hits are free (re-running the same input)
  • During the alpha/beta period, free trial credits (login bonus and first-use bonus) let you evaluate every API

Headline rates: text to speech = 200 AP per 1,000 characters (2026-09-09; revised weekly on the published curve); LLM = the per-model rates in the table above (always current via GET /v1/chat/models); image generation = from 700 AP per image (gpt-image-2.5-flare, quality: low, 1024×1024).

Businesses that need a Japanese qualified invoice (適格請求書) can purchase through our reseller AICU Japan K.K. (info@aicu.jp).