AICU API Platform
News & Changelog
Product updates, new models, and pricing changes. Pricing changes are tagged Pricing. The live source of truth for models and prices is GET /v1/chat/models.
Something not working right now? Check live status and incidents. This page is the record of planned changes — announcements, new models and pricing.
- Platform
The typed-answers endpoint is `POST /v1/jev` (was `/v1/decisions`)
We renamed it during the beta. `POST /v1/jev` replaces
Skill document ↗POST /v1/decisions, and the skill document is at/skills/jev. Nothing else changed — same model (jev-1.13.0), same request body, same price, samellmscope. Why: a *decision* is something a person makes — heavy, discrete, final. What comes back here is a probability: a number per option, an expected level, a 0–1 strength. The old name invited people to read0.92as "settled". It is not settled until you pick a threshold, and picking it is your call. The old path is gone, not deprecated: it had no callers outside AICU, so we cut it rather than leave a second name in place. If you wrote code against it during the beta, change the URL and the responseobjectfield ("decisions"→"jev"). - Pricing
JPY prices for AP packs move to the invoice lane rate on Sep 30; USD prices are unchanged
AP packs can be bought from two sellers: AICU Inc. in USD, and AICU Japan K.K. in JPY. The JPY lane exists for customers who need a Japanese qualified invoice (適格請求書), pay by bank transfer against an invoice, or commit an annual budget in advance. From 2026-09-30 05:00 UTC the JPY prices are set from a billing rate of 1 USD = ¥180: Pack S ¥1,700 → ¥1,800 (100,000 AP), Pack M ¥8,400 → ¥9,000 (500,000 AP), Pack L ¥33,500 → ¥36,000 (2,000,000 AP). USD prices do not change: $10.00 / $50.00 / $200.00. Validity stays 24 months. From now on the JPY price is always at least 10% above the USD price converted at the market rate, and we check it every week against the published USD/JPY table. That gap is the price of the invoice lane, not a currency surcharge — if you do not need a qualified invoice, the USD lane is the cheaper one and stays open. Announced nine days ahead, per our price-revision rule (increases get 7 days’ notice and take effect on a Wednesday at 05:00 UTC). The same notice retires two old JPY-only routes that priced the same AP *below* the USD lane: the legacy AP packages (
AP packs ↗POST /v1/stripe/checkout, ¥100 / ¥980 / ¥9,500) and the monthly JPY subscription plans (POST /v1/stripe/subscribe). Both now return HTTP 410 and point at the pack catalogue. No account was subscribed to either. - Model
New image models: GPT Image 2.5 Flare and Sunburst are serving now
Image API ↗gpt-image-2.5-flareis the one to reach for by default. It is better than the 2 series and takes about half the time (typically 85 s against 170 s), and it keeps reference-image support, so a character stays the same character across images.gpt-image-2.5-sunbursttrades time for control: finer editing, typically 200 s. Sizes are 1024×1024, 1536×1024 and 1024×1536. Flare at 1024×1024 is 700 AP onlow, 1,600 AP onmediumand 6,400 AP onhigh; the wide and tall sizes are 1,300 / 2,300 / 5,100 AP. Every price is inGET /v1/models— read it from there rather than keeping a copy. Both models regularly run past 100 seconds, which is longer than the edge will hold a connection: if the HTTP call ends in 524, the image is not lost. Keep theidfrom the response (or theX-AICU-Image-Idheader when streaming) and collect it withGET /v1/images/status/{id}. Use the id, notGET /v1/images/recent, whenever more than one generation can be in flight on the same key. - Feature
Jev: get a typed answer back with a confidence value, not prose
API reference ↗POST /v1/jevsends your state and your questions to Jev (TypeSafe AI) and returns typed answers with a confidence value — a boolean, a number or a choice, in the shape you asked for. No prompt engineering to keep the model from replying in paragraphs, and no parsing of free text at your end. It is the piece you want when a decision sits inside your control flow: should this be escalated, is this record complete, which of these three queues. It uses your existingllmscope, so keys that already work need no change, and it is fast enough to sit on a request path. Billing is per call and small — a two-question decision costs single-digit AP. - Program
Beta extended by four weeks: beta keys are issued until Oct 14 and work until Oct 31; production starts Oct 21
We are moving the beta calendar published on Sep 9 back by four weeks. Nothing changes for you before Oct 14. The revised dates, each at 09:00 JST: Sep 23 — Generative AI Expo Vol.6 (Hamamatsucho). Public beta demo at the AICU booth; keys are not switched off. Oct 7 — AP pack purchase opens in USD, and we email every account holder the production guide (commercial declaration, auto-recharge, card registration). Oct 14 — no new beta keys are issued; keys issued from here are production
Issue a key ↗aicu_live_keys. Oct 21 — production operation begins. New keys require a registered card. Oct 31 — beta keys stop working (the same day the beta and the Kumamoto relief programme end). Why: the paid purchase path is not yet open, and the week of Sep 19–24 is taken by the book launch, the expo and the Award first-round deadline. Starting production without a way to pay would be a label, not a launch. Existingaicu_alpha_andaicu_beta_keys keep working until Oct 31; expiry never deletes anything — usage history and balance stay, the key simply returns 401. We email the owner a week before any key lapses. Prices are unchanged by this notice. - Pricing
X (Twitter) gateway: from Sep 23, profile lookups and post reads are 100 AP per upstream fetch (post reads now include attached media)
The X gateway (
X gateway skill →GET /v1/x/*, scopex) lets you read X posts, profiles, timelines and search results through your AICU key. From 2026-09-23 05:00 UTC:GET /v1/x/user/:usernameis billed 100 AP (about $0.01) per upstream lookup, andGET /v1/x/status/:idis also 100 AP (about $0.01) per upstream fetch. Post reads now return the text, author,created_atand the attached images/videos (media[]with URLs and sizes), so a quote card can be built from one call without screenshotting the page — that is why the earlier 10 AP figure (announced 2026-09-13) is revised to 100 AP. Both are cached (profiles 12 hours, posts 10 minutes) and cache hits stay free; failed upstream fetches are not billed. Account-ownership verification stays free.timelineandsearchare unchanged at 2,250 AP per successful call. Revised notice published 2026-09-14, 9 days ahead, per our price-revision rule (increases get 7 days’ notice; only HTTP 200 is billed). Agent-readable spec: https://api.aicu.ai/skills/x - Platform
GET /v1/models now lists every model we serve — chat, image, speech, transcription — with prices, docs and update dates
Try GET /v1/models ↗GET /v1/modelsis now the single source of truth for everything on the platform. Until today it listed LLMs and images only, andGET /v1/chat/modelswas the only place with prices. Every entry now carries the same fields:type(chat / image / tts / stt), the endpoint to call,stage,pricingwith an explicit unit (per 1K tokens, per image by size and quality, per 1,000 characters, per 10 seconds),docs_url,pricing_urlandupdated_at. The list envelope carriesdocs,pricing_urland the latestupdated_at.GET /v1/models/:idreturns one model, OpenAI-style, and accepts image aliases such asflare.GET /v1/chat/modelskeeps working with the same fields as before; it is now atype=chatfilter of the same array. One label change:stageno longer saysalpha— the alpha programme ended on Sep 9, so models arebetauntil production starts on Sep 23. Also from today: when an account runs out of AP, the owner gets one email per day explaining the pause, sent in Japanese from our Japan reseller (info@aicu.jp) for accounts in Japan, and in English from assistant@aicu.ai everywhere else. Nothing is charged while paused. - Program
Key programme: alpha ends Sep 9, beta keys run to Sep 30
Revised on 2026-09-14 — the four dates below were moved back by four weeks (issuance stop Oct 14, production Oct 21, beta keys expire Oct 31). See the Sep 14 notice. The free alpha evaluation closes and metered usage begins on 2026-09-09 (09:00 JST). Four dates follow, each at 09:00 JST inside the Wednesday maintenance window: Sep 9 — keys issued from here carry the
Issue a key ↗aicu_beta_prefix; existingaicu_alpha_keys keep working. Sep 16 — no new beta keys are issued. Sep 23 — production operation begins. Sep 30 — beta keys stop working. The week between Sep 23 and Sep 30 is deliberate overlap: production starts at an event, and we are not switching everyone off on the morning we first run it in public. Move to a production key during that week and keep both live until the new one is confirmed. Once you are on card billing, the next key you issue is a productionaicu_live_key and is not bound to these dates. Expiry never deletes anything — usage history and balance are untouched; the key simply returns 401. We now email the owner a week before any key lapses. - Model
Fixed: four models failed unless you set reasoning_effort yourself
Calls to
Model list ↗gpt-5.6-terra,gpt-5.6-luna,gpt-5.4-miniandgpt-5.4-nanoreturned 400 — “reasoning_effort does not support ‘minimal’ with this model” — whenever the request leftreasoning_effortunset. Setting it explicitly always worked, so the models themselves were fine; the default we supply was not. That default exists so reasoning cannot silently consume your wholemax_tokens, and it was sendingminimal, a value the newer dotted generations (gpt-5.4, gpt-5.6) renamed tonone. The default is now chosen by generation instead of by model name, so the next derivative is covered before it ships.gpt-5.6-solwas unaffected and keeps itslowdefault; the cheaper, faster models now default tonone. Passingreasoning_effortyourself still overrides, as before. - Platform
Provider errors now keep their `param` and `code`
When an upstream provider rejects a request, we passed its message through but replaced
typeandcodewith our own values and droppedparamentirely — so a 400 told you something was wrong without telling you which field. Provider errors are now returned in OpenAI’s own shape,paramandcodeincluded, withprovider_statusadded alongside so you can still tell our layer from theirs. Nothing about status codes orRetry-Afterchanges. - Model
Image input confirmed on the gpt-5.6 family — and what tool-calling clients need to know
You can pass images to
OpenCode / Cline setup ↗gpt-5.6-sol,gpt-5.6-terraandgpt-5.6-lunathroughPOST /v1/chat/completionsusing the standard OpenAIimage_urlcontent part, verified end to end against production. Two things to know if you drive these from a coding tool such as Cline or OpenCode. First, tool calling and reasoning cannot be combined on this generation, so the gateway sendsreasoning_effort: “none”whenever your request carriestools— you do not need to set it yourself. Second, reasoning tokens come out ofmax_tokens: with the default effort a short image question spent 74 of 89 completion tokens on reasoning, and a request capped at 40 tokens failed outright. Give image requests room, or sendreasoning_effort: “none”. Note thatGET /v1/chat/modelsstill describes capabilities only in prose, so tools that look for a machine-readable field will assume no image support until you enable it in their config. - Model
Video generation arrives Sep 2 — Gemini Omni 1.1 Flash, from text, a still, or two keyframes
From 2026-09-02 you can generate video: text-to-video, image-to-video, first-and-last-frame interpolation, and reference-to-video that holds a character across the shot. Ten seconds per call at 360p, 720p (default) or 1080p, returned synchronously — we measured 35.8s on average for 720p, so there is no job queue to poll. A seed is supported, and the same seed returns the same shot (PSNR 41.3 dB across our runs). Two things to know before you build on it: if you write that someone is speaking but do not supply the line, the model invents dialogue — specify the line, or ask for silence; and 4K is not in this release, because a 4K clip takes about 210 seconds to generate and will not return inside a single request. Free like everything else through Sep 8; billed from Sep 9.
Measured benchmark → - Platform
Transcription moves to a new backend — more accurate, and it stops rate-limiting you
The default backend for
Japanese STT benchmark →POST /v1/audio/transcriptionsis now OpenAIgpt-4o-mini-transcribe, replacing the previous provider. Two reasons: the old backend returned429under any sustained use — we could not complete twelve retries over fourteen minutes on our own benchmark run — and it placed last in our Japanese STT benchmark at 64.97% CER, because roughly 40% of the body text was dropped at the seams where long audio is split internally. The new default measured 23.40% CER on the same audio, and it costs us slightly less. One behavioural change: the per-request size limit drops from 30 MB to 25 MB, which is the new backend’s own limit; the 30-minute duration limit is unchanged. The previous backend stays available withprovider=sakura, andmodel=whisper-large-v3-turboroutes there too. Prices are unchanged by this switch. - Pricing
Text-to-speech moves to per-character pricing on Sep 9 — 200 AP per 1,000 characters
Speech synthesis is billed per character from 2026-09-09 05:00 UTC, matching how our upstream (ElevenLabs) meters it: 200 AP per 1,000 characters, about $0.02 — roughly one-eighth of ElevenLabs’ own list price for the same voices, and with the character licensing already cleared. Until now we billed on estimated audio seconds, and the estimate ran about 2x the audio we actually produced; that is fixed in the same release. Announced 13 days ahead per our price-revision rule (increases get 7-day notice; cuts are immediate).
TTS pricing → - Pricing
Free alpha evaluation extended to Sep 8 — paid usage starts Sep 9
Everything stays free through 2026-09-08, one week longer than previously planned: we are holding rates flat across the Sep 1–4 student-lecture week. From 2026-09-09 05:00 UTC, usage is metered against your AP balance. Nothing is charged automatically — the platform is prepaid, so if your balance runs out the API simply returns 402, and no invoice or card charge follows. Free tiers, the signup bonus and the first-use bonus all continue to work after that date.
What "billing starts" means → - Pricing
Pricing notice: transcription 10 s = 3 AP from Sep 9 (05:00 UTC)
Speech-to-text (
Japanese STT benchmark →POST /v1/audio/transcriptions) moves from 1 AP to 3 AP per 10 seconds of audio on 2026-09-09 05:00 UTC — about $0.11 per hour of audio, following our cost×3 policy (the current rate has been below provider cost). Update (Aug 26): effective date moved from Sep 2 to Sep 9 — we are keeping current rates through the Sep 1–4 student-lecture week. Announced well past our 7-day-notice rule for increases. Timestamps (verbose_json),srt/vttoutput and the new/skills/sttspec shipped this week; a premium high-accuracy lane (ElevenLabs Scribe) is being evaluated — see the benchmark on the docs blog. - Platform
Fixed: non-browser clients (e.g. Python-urllib) blocked with 403 on api.aicu.ai
Cloudflare’s Browser Integrity Check was rejecting requests whose
User-Agentlooked non-browser-like — most visiblyPython-urllib/*, the default UA of Python’s standard-library HTTP client — with403(error code: 1010), even for public endpoints like/healthand/openapi.json. This is fixed forapi.aicu.ai; if you were setting a fakeUser-Agentto work around it, that’s no longer necessary. (We kept this protection in place on our other, non-API domains.) - Platform
Fixed: GET /v1/usage was missing TTS and image requests
The self-serve usage summary (
View your usage →GET /v1/usage) was silently scoped to LLM calls only, so text-to-speech and image-generation requests didn’t show up insummary.requestsorby_model— even though they were billed correctly and appeared inGET /v1/usage/logsand the dashboard. Fixed: the summary now covers every billed service. - Model
ElevenLabs v3 — emotional TTS, plus a dedicated compat endpoint
Audio API docs →POST /v1/audio/speechand the newPOST /v1/el/ttsnow acceptmodel: "eleven_v3"for ElevenLabs’ newest model — inline audio tags like[excited]or[whispers]shape delivery (5,000-char limit per request, vs 10,000 for the defaulteleven_multilingual_v2). Unrecognizedmodelvalues now return a clear 400 instead of silently generating with the wrong model. Also fixed:POST /v1/tts/generate(theslug-based legacy path, e.g.slug: "luc4") had been pointing at a retired backend and always failing — it now runs on the same ElevenLabs gateway as everything else. Verified against production with real audio for both models. - Pricing
Gemini pricing update
The default
See pricing tiers →gemini-flashnow maps to Gemini 2.5 Flash-Lite at 3 / 12 AP per 1K tokens (input / output) — lighter and cheaper for high-volume work. A newgemini-3.6-flashis added at 45 / 225 AP for top-quality reasoning. All rates follow our cost×3 policy. - Model
Kimi-K2.7-Code (Powered by Sakura AI Engine)
A coding-specialized model —
OpenCode setup →kimi-k2.7-code. OpenAI-compatible and verified working in OpenCode (v1.18.9) through api.aicu.ai. Also works from Cline and Continue. Free during public preview; domestic inference, no training on your data. - Platform
API keys required · reasoning models improved
All
Get an API key →/v1endpoints now require a Bearer API key (issue one in the dashboard; public listing endpoints stay open). Reasoning models (gpt-5*) now honorreasoning_effort(minimal/low/medium/high) and default tominimalwhen unset, so short prompts return content instead of spending the budget on reasoning tokens. - Program
Reiwa 8 Kumamoto Earthquake — free API for relief apps
Free access to our AI API infrastructure (OpenAI-compatible LLMs, multilingual text-to-speech, Whisper speech-to-text, image generation) for developers building disaster-support apps, through 2026-10-31.
AICU Call ↗ - Feature
Audio API — multilingual TTS + Whisper STT
Text-to-speech via ElevenLabs voices (elena / mei / mina / nao / saki) at
Audio API docs →POST /v1/audio/speech, and speech-to-text via Sakurawhisper-large-v3-turboatPOST /v1/audio/transcriptions. Both are OpenAI-SDK compatible. - Feature
Image API with C2PA provenance
Generate images at
Image API docs →POST /v1/images/generations(GPT-Image-2 and Nano Banana 2), and detect embedded C2PA content-credentials atPOST /v1/images/inspect.