AICU API Platform

News & Changelog

Product updates, new models, and pricing changes. Pricing changes are tagged Pricing. The live source of truth for models and prices is GET /v1/chat/models.

Something not working right now? Check live status and incidents. This page is the record of planned changes — announcements, new models and pricing.

  1. Platform

    The typed-answers endpoint is `POST /v1/jev` (was `/v1/decisions`)

    We renamed it during the beta. `POST /v1/jev` replaces POST /v1/decisions, and the skill document is at /skills/jev. Nothing else changed — same model (jev-1.13.0), same request body, same price, same llm scope. Why: a *decision* is something a person makes — heavy, discrete, final. What comes back here is a probability: a number per option, an expected level, a 0–1 strength. The old name invited people to read 0.92 as "settled". It is not settled until you pick a threshold, and picking it is your call. The old path is gone, not deprecated: it had no callers outside AICU, so we cut it rather than leave a second name in place. If you wrote code against it during the beta, change the URL and the response object field ("decisions" → "jev").

    Skill document ↗
  2. Pricing

    JPY prices for AP packs move to the invoice lane rate on Sep 30; USD prices are unchanged

    AP packs can be bought from two sellers: AICU Inc. in USD, and AICU Japan K.K. in JPY. The JPY lane exists for customers who need a Japanese qualified invoice (適格請求書), pay by bank transfer against an invoice, or commit an annual budget in advance. From 2026-09-30 05:00 UTC the JPY prices are set from a billing rate of 1 USD = ¥180: Pack S ¥1,700 → ¥1,800 (100,000 AP), Pack M ¥8,400 → ¥9,000 (500,000 AP), Pack L ¥33,500 → ¥36,000 (2,000,000 AP). USD prices do not change: $10.00 / $50.00 / $200.00. Validity stays 24 months. From now on the JPY price is always at least 10% above the USD price converted at the market rate, and we check it every week against the published USD/JPY table. That gap is the price of the invoice lane, not a currency surcharge — if you do not need a qualified invoice, the USD lane is the cheaper one and stays open. Announced nine days ahead, per our price-revision rule (increases get 7 days’ notice and take effect on a Wednesday at 05:00 UTC). The same notice retires two old JPY-only routes that priced the same AP *below* the USD lane: the legacy AP packages (POST /v1/stripe/checkout, ¥100 / ¥980 / ¥9,500) and the monthly JPY subscription plans (POST /v1/stripe/subscribe). Both now return HTTP 410 and point at the pack catalogue. No account was subscribed to either.

    AP packs ↗
  3. Model

    New image models: GPT Image 2.5 Flare and Sunburst are serving now

    gpt-image-2.5-flare is the one to reach for by default. It is better than the 2 series and takes about half the time (typically 85 s against 170 s), and it keeps reference-image support, so a character stays the same character across images. gpt-image-2.5-sunburst trades time for control: finer editing, typically 200 s. Sizes are 1024×1024, 1536×1024 and 1024×1536. Flare at 1024×1024 is 700 AP on low, 1,600 AP on medium and 6,400 AP on high; the wide and tall sizes are 1,300 / 2,300 / 5,100 AP. Every price is in GET /v1/models — read it from there rather than keeping a copy. Both models regularly run past 100 seconds, which is longer than the edge will hold a connection: if the HTTP call ends in 524, the image is not lost. Keep the id from the response (or the X-AICU-Image-Id header when streaming) and collect it with GET /v1/images/status/{id}. Use the id, not GET /v1/images/recent, whenever more than one generation can be in flight on the same key.

    Image API ↗
  4. Feature

    Jev: get a typed answer back with a confidence value, not prose

    POST /v1/jev sends your state and your questions to Jev (TypeSafe AI) and returns typed answers with a confidence value — a boolean, a number or a choice, in the shape you asked for. No prompt engineering to keep the model from replying in paragraphs, and no parsing of free text at your end. It is the piece you want when a decision sits inside your control flow: should this be escalated, is this record complete, which of these three queues. It uses your existing llm scope, so keys that already work need no change, and it is fast enough to sit on a request path. Billing is per call and small — a two-question decision costs single-digit AP.

    API reference ↗
  5. Program

    Beta extended by four weeks: beta keys are issued until Oct 14 and work until Oct 31; production starts Oct 21

    We are moving the beta calendar published on Sep 9 back by four weeks. Nothing changes for you before Oct 14. The revised dates, each at 09:00 JST: Sep 23 — Generative AI Expo Vol.6 (Hamamatsucho). Public beta demo at the AICU booth; keys are not switched off. Oct 7 — AP pack purchase opens in USD, and we email every account holder the production guide (commercial declaration, auto-recharge, card registration). Oct 14 — no new beta keys are issued; keys issued from here are production aicu_live_ keys. Oct 21 — production operation begins. New keys require a registered card. Oct 31 — beta keys stop working (the same day the beta and the Kumamoto relief programme end). Why: the paid purchase path is not yet open, and the week of Sep 19–24 is taken by the book launch, the expo and the Award first-round deadline. Starting production without a way to pay would be a label, not a launch. Existing aicu_alpha_ and aicu_beta_ keys keep working until Oct 31; expiry never deletes anything — usage history and balance stay, the key simply returns 401. We email the owner a week before any key lapses. Prices are unchanged by this notice.

    Issue a key ↗
  6. Pricing

    X (Twitter) gateway: from Sep 23, profile lookups and post reads are 100 AP per upstream fetch (post reads now include attached media)

    The X gateway (GET /v1/x/*, scope x) lets you read X posts, profiles, timelines and search results through your AICU key. From 2026-09-23 05:00 UTC: GET /v1/x/user/:username is billed 100 AP (about $0.01) per upstream lookup, and GET /v1/x/status/:id is also 100 AP (about $0.01) per upstream fetch. Post reads now return the text, author, created_at and the attached images/videos (media[] with URLs and sizes), so a quote card can be built from one call without screenshotting the page — that is why the earlier 10 AP figure (announced 2026-09-13) is revised to 100 AP. Both are cached (profiles 12 hours, posts 10 minutes) and cache hits stay free; failed upstream fetches are not billed. Account-ownership verification stays free. timeline and search are unchanged at 2,250 AP per successful call. Revised notice published 2026-09-14, 9 days ahead, per our price-revision rule (increases get 7 days’ notice; only HTTP 200 is billed). Agent-readable spec: https://api.aicu.ai/skills/x

    X gateway skill →
  7. Platform

    GET /v1/models now lists every model we serve — chat, image, speech, transcription — with prices, docs and update dates

    GET /v1/models is now the single source of truth for everything on the platform. Until today it listed LLMs and images only, and GET /v1/chat/models was the only place with prices. Every entry now carries the same fields: type (chat / image / tts / stt), the endpoint to call, stage, pricing with an explicit unit (per 1K tokens, per image by size and quality, per 1,000 characters, per 10 seconds), docs_url, pricing_url and updated_at. The list envelope carries docs, pricing_url and the latest updated_at. GET /v1/models/:id returns one model, OpenAI-style, and accepts image aliases such as flare. GET /v1/chat/models keeps working with the same fields as before; it is now a type=chat filter of the same array. One label change: stage no longer says alpha — the alpha programme ended on Sep 9, so models are beta until production starts on Sep 23. Also from today: when an account runs out of AP, the owner gets one email per day explaining the pause, sent in Japanese from our Japan reseller (info@aicu.jp) for accounts in Japan, and in English from assistant@aicu.ai everywhere else. Nothing is charged while paused.

    Try GET /v1/models ↗
  8. Program

    Key programme: alpha ends Sep 9, beta keys run to Sep 30

    Revised on 2026-09-14 — the four dates below were moved back by four weeks (issuance stop Oct 14, production Oct 21, beta keys expire Oct 31). See the Sep 14 notice. The free alpha evaluation closes and metered usage begins on 2026-09-09 (09:00 JST). Four dates follow, each at 09:00 JST inside the Wednesday maintenance window: Sep 9 — keys issued from here carry the aicu_beta_ prefix; existing aicu_alpha_ keys keep working. Sep 16 — no new beta keys are issued. Sep 23 — production operation begins. Sep 30 — beta keys stop working. The week between Sep 23 and Sep 30 is deliberate overlap: production starts at an event, and we are not switching everyone off on the morning we first run it in public. Move to a production key during that week and keep both live until the new one is confirmed. Once you are on card billing, the next key you issue is a production aicu_live_ key and is not bound to these dates. Expiry never deletes anything — usage history and balance are untouched; the key simply returns 401. We now email the owner a week before any key lapses.

    Issue a key ↗
  9. Model

    Fixed: four models failed unless you set reasoning_effort yourself

    Calls to gpt-5.6-terra, gpt-5.6-luna, gpt-5.4-mini and gpt-5.4-nano returned 400 — “reasoning_effort does not support ‘minimal’ with this model” — whenever the request left reasoning_effort unset. Setting it explicitly always worked, so the models themselves were fine; the default we supply was not. That default exists so reasoning cannot silently consume your whole max_tokens, and it was sending minimal, a value the newer dotted generations (gpt-5.4, gpt-5.6) renamed to none. The default is now chosen by generation instead of by model name, so the next derivative is covered before it ships. gpt-5.6-sol was unaffected and keeps its low default; the cheaper, faster models now default to none. Passing reasoning_effort yourself still overrides, as before.

    Model list ↗
  10. Platform

    Provider errors now keep their `param` and `code`

    When an upstream provider rejects a request, we passed its message through but replaced type and code with our own values and dropped param entirely — so a 400 told you something was wrong without telling you which field. Provider errors are now returned in OpenAI’s own shape, param and code included, with provider_status added alongside so you can still tell our layer from theirs. Nothing about status codes or Retry-After changes.

  11. Model

    Image input confirmed on the gpt-5.6 family — and what tool-calling clients need to know

    You can pass images to gpt-5.6-sol, gpt-5.6-terra and gpt-5.6-luna through POST /v1/chat/completions using the standard OpenAI image_url content part, verified end to end against production. Two things to know if you drive these from a coding tool such as Cline or OpenCode. First, tool calling and reasoning cannot be combined on this generation, so the gateway sends reasoning_effort: “none” whenever your request carries tools — you do not need to set it yourself. Second, reasoning tokens come out of max_tokens: with the default effort a short image question spent 74 of 89 completion tokens on reasoning, and a request capped at 40 tokens failed outright. Give image requests room, or send reasoning_effort: “none”. Note that GET /v1/chat/models still describes capabilities only in prose, so tools that look for a machine-readable field will assume no image support until you enable it in their config.

    OpenCode / Cline setup ↗
  12. Model

    Video generation arrives Sep 2 — Gemini Omni 1.1 Flash, from text, a still, or two keyframes

    From 2026-09-02 you can generate video: text-to-video, image-to-video, first-and-last-frame interpolation, and reference-to-video that holds a character across the shot. Ten seconds per call at 360p, 720p (default) or 1080p, returned synchronously — we measured 35.8s on average for 720p, so there is no job queue to poll. A seed is supported, and the same seed returns the same shot (PSNR 41.3 dB across our runs). Two things to know before you build on it: if you write that someone is speaking but do not supply the line, the model invents dialogue — specify the line, or ask for silence; and 4K is not in this release, because a 4K clip takes about 210 seconds to generate and will not return inside a single request. Free like everything else through Sep 8; billed from Sep 9.

    Measured benchmark →
  13. Platform

    Transcription moves to a new backend — more accurate, and it stops rate-limiting you

    The default backend for POST /v1/audio/transcriptions is now OpenAI gpt-4o-mini-transcribe, replacing the previous provider. Two reasons: the old backend returned 429 under any sustained use — we could not complete twelve retries over fourteen minutes on our own benchmark run — and it placed last in our Japanese STT benchmark at 64.97% CER, because roughly 40% of the body text was dropped at the seams where long audio is split internally. The new default measured 23.40% CER on the same audio, and it costs us slightly less. One behavioural change: the per-request size limit drops from 30 MB to 25 MB, which is the new backend’s own limit; the 30-minute duration limit is unchanged. The previous backend stays available with provider=sakura, and model=whisper-large-v3-turbo routes there too. Prices are unchanged by this switch.

    Japanese STT benchmark →
  14. Pricing

    Text-to-speech moves to per-character pricing on Sep 9 — 200 AP per 1,000 characters

    Speech synthesis is billed per character from 2026-09-09 05:00 UTC, matching how our upstream (ElevenLabs) meters it: 200 AP per 1,000 characters, about $0.02 — roughly one-eighth of ElevenLabs’ own list price for the same voices, and with the character licensing already cleared. Until now we billed on estimated audio seconds, and the estimate ran about 2x the audio we actually produced; that is fixed in the same release. Announced 13 days ahead per our price-revision rule (increases get 7-day notice; cuts are immediate).

    TTS pricing →
  15. Pricing

    Free alpha evaluation extended to Sep 8 — paid usage starts Sep 9

    Everything stays free through 2026-09-08, one week longer than previously planned: we are holding rates flat across the Sep 1–4 student-lecture week. From 2026-09-09 05:00 UTC, usage is metered against your AP balance. Nothing is charged automatically — the platform is prepaid, so if your balance runs out the API simply returns 402, and no invoice or card charge follows. Free tiers, the signup bonus and the first-use bonus all continue to work after that date.

    What "billing starts" means →
  16. Pricing

    Pricing notice: transcription 10 s = 3 AP from Sep 9 (05:00 UTC)

    Speech-to-text (POST /v1/audio/transcriptions) moves from 1 AP to 3 AP per 10 seconds of audio on 2026-09-09 05:00 UTC — about $0.11 per hour of audio, following our cost×3 policy (the current rate has been below provider cost). Update (Aug 26): effective date moved from Sep 2 to Sep 9 — we are keeping current rates through the Sep 1–4 student-lecture week. Announced well past our 7-day-notice rule for increases. Timestamps (verbose_json), srt/vtt output and the new /skills/stt spec shipped this week; a premium high-accuracy lane (ElevenLabs Scribe) is being evaluated — see the benchmark on the docs blog.

    Japanese STT benchmark →
  17. Platform

    Fixed: non-browser clients (e.g. Python-urllib) blocked with 403 on api.aicu.ai

    Cloudflare’s Browser Integrity Check was rejecting requests whose User-Agent looked non-browser-like — most visibly Python-urllib/*, the default UA of Python’s standard-library HTTP client — with 403 (error code: 1010), even for public endpoints like /health and /openapi.json. This is fixed for api.aicu.ai; if you were setting a fake User-Agent to work around it, that’s no longer necessary. (We kept this protection in place on our other, non-API domains.)

  18. Platform

    Fixed: GET /v1/usage was missing TTS and image requests

    The self-serve usage summary (GET /v1/usage) was silently scoped to LLM calls only, so text-to-speech and image-generation requests didn’t show up in summary.requests or by_model — even though they were billed correctly and appeared in GET /v1/usage/logs and the dashboard. Fixed: the summary now covers every billed service.

    View your usage →
  19. Model

    ElevenLabs v3 — emotional TTS, plus a dedicated compat endpoint

    POST /v1/audio/speech and the new POST /v1/el/tts now accept model: "eleven_v3" for ElevenLabs’ newest model — inline audio tags like [excited] or [whispers] shape delivery (5,000-char limit per request, vs 10,000 for the default eleven_multilingual_v2). Unrecognized model values now return a clear 400 instead of silently generating with the wrong model. Also fixed: POST /v1/tts/generate (the slug-based legacy path, e.g. slug: "luc4") had been pointing at a retired backend and always failing — it now runs on the same ElevenLabs gateway as everything else. Verified against production with real audio for both models.

    Audio API docs →
  20. Pricing

    Gemini pricing update

    The default gemini-flash now maps to Gemini 2.5 Flash-Lite at 3 / 12 AP per 1K tokens (input / output) — lighter and cheaper for high-volume work. A new gemini-3.6-flash is added at 45 / 225 AP for top-quality reasoning. All rates follow our cost×3 policy.

    See pricing tiers →
  21. Model

    Kimi-K2.7-Code (Powered by Sakura AI Engine)

    A coding-specialized model — kimi-k2.7-code. OpenAI-compatible and verified working in OpenCode (v1.18.9) through api.aicu.ai. Also works from Cline and Continue. Free during public preview; domestic inference, no training on your data.

    OpenCode setup →
  22. Platform

    API keys required · reasoning models improved

    All /v1 endpoints now require a Bearer API key (issue one in the dashboard; public listing endpoints stay open). Reasoning models (gpt-5*) now honor reasoning_effort (minimal/low/medium/high) and default to minimal when unset, so short prompts return content instead of spending the budget on reasoning tokens.

    Get an API key →
  23. Program

    Reiwa 8 Kumamoto Earthquake — free API for relief apps

    Free access to our AI API infrastructure (OpenAI-compatible LLMs, multilingual text-to-speech, Whisper speech-to-text, image generation) for developers building disaster-support apps, through 2026-10-31.

    AICU Call ↗
  24. Feature

    Audio API — multilingual TTS + Whisper STT

    Text-to-speech via ElevenLabs voices (elena / mei / mina / nao / saki) at POST /v1/audio/speech, and speech-to-text via Sakura whisper-large-v3-turbo at POST /v1/audio/transcriptions. Both are OpenAI-SDK compatible.

    Audio API docs →
  25. Feature

    Image API with C2PA provenance

    Generate images at POST /v1/images/generations (GPT-Image-2 and Nano Banana 2), and detect embedded C2PA content-credentials at POST /v1/images/inspect.

    Image API docs →