Skip to main content

AICU Audio API

OpenAI-compatible text-to-speech (TTS) and speech-to-text (STT). One key, the same endpoint host as LLM and image.

Text-to-speech (TTS)​

POST /v1/audio/speech (OpenAI-compatible)​

curl -X POST https://api.aicu.ai/v1/audio/speech \
-H "Authorization: Bearer aicu_live_xxx" \
-H "Content-Type: application/json" \
-d '{"input": "こんにちは、AICU です", "voice": "nao"}' \
--output speech.mp3
ParameterDescription
inputThe text to speak
voiceAn AICU voice ID (see the list below), or a raw ElevenLabs voice_id (20 alphanumeric characters). Defaults to nao
modeleleven_multilingual_v2 (default) or eleven_v3. See "Models" below
seedRandom seed for generation. If omitted, each character's default seed is used (see Character voices and seeds)

The ElevenLabs voices are multilingual, so Japanese, English, and other languages synthesize as-is — useful for things like multilingual emergency announcements.

:::info API key scopes Speech synthesis and transcription require a key with the tts scope. If you get a 403 with required_scope: tts, the key is missing that scope — add it with 権限変更 (Change permissions) at https://api.aicu.ai/dashboard/keys. You do not need to reissue the key. :::

POST /v1/el/tts (ElevenLabs-compatible path)​

The body is identical to /v1/audio/speech. voice accepts not only AICU-registered voices (elena and friends) but also a raw ElevenLabs voice_id — including voices you created or cloned in your own ElevenLabs account. TTS here is deliberately just a gateway (keys, billing, caching); how good a voice sounds is ElevenLabs' department. Character-specific credit requirements and royalties live in a different layer (/v1/char).

# AICU-registered voice
curl -X POST https://api.aicu.ai/v1/el/tts \
-H "Authorization: Bearer aicu_live_xxx" \
-H "Content-Type: application/json" \
-d '{"input": "こんにちは、AICU です", "voice": "nao", "model": "eleven_v3"}' \
--output speech.mp3

# Raw voice_id (a voice from your own ElevenLabs account)
curl -X POST https://api.aicu.ai/v1/el/tts \
-H "Authorization: Bearer aicu_live_xxx" \
-H "Content-Type: application/json" \
-d '{"input": "Hello from AICU", "voice": "pqHfZKP75CvOlQylNhV4", "model": "eleven_v3"}' \
--output speech.mp3

Models: Multilingual v2 vs. v3​

ModelCharacteristicsCharacters per request
eleven_multilingual_v2 (default)Stable and proven. Synthesis as you already know it10,000 characters
eleven_v3Write square-bracket audio tags such as [excited], [whispers], or [laughs] inside text, and that emotion or delivery is reflected in the audio5,000 characters

To narrate long text with v3, chunk it so each request fits under the limit (going over returns an explicit 400). An unknown model value is also a 400 — we never silently fall back to v2.

:::info Long-input stability (2026-08-16) Long inputs (roughly 3,000+ Japanese characters) used to fail partway through with 502/524. We made the internals streaming, and that's fixed. You can now generate reliably right up to the limits in the table (v2: 10,000 characters / v3: 5,000 characters), and time-to-first-byte is shorter too. Audio starts arriving in chunks, so the receiving side just needs to write the response body straight to a file to get a complete MP3. Regenerating the same input is served from cache, free, every time. :::

curl -X POST https://api.aicu.ai/v1/el/tts \
-H "Authorization: Bearer aicu_live_xxx" \
-H "Content-Type: application/json" \
-d '{
"input": "[SHOUTING][excited] GOOOAL! What a strike!!",
"voice": "nao",
"model": "eleven_v3"
}' \
--output goal.mp3

That's the exact pattern our live football commentary service RoboFoot runs in production — it stacks [SHOUTING] and [laughing] by tier to crank up the intensity.

Official ElevenLabs references:

GET /v1/tts/voices and GET /v1/el/voices (API key required)​

curl https://api.aicu.ai/v1/tts/voices \
-H "Authorization: Bearer aicu_live_xxx"

As of 2026-08-17 these require an API key (any scope). Unauthenticated requests get a 401.

VoiceDisplay nameCharacterSample
elenaElena BloomCalm, easy-to-follow female narration in standard Japanese
meiMei SoleilBright young girl — energetic and easy for children to follow
minaMina AzurComposed, intelligent narration that doesn't tire you out over long sessions
naoNao VerdeSoft-spoken, gentle science-nerd guy
sakiSaki NoirWhispery voice, gentle with a hint of mystery
marshaMarsha AranciaFresh, upbeat, easy-listening female voice
luc4Luca (ルカ / LuC4, the all-affirming boyfriend)Cheerful, outgoing young man with a Kansai flavor

Every sample is that member's self-introduction generated with the official seed (the same take as AiCuty-all-intro).

All of them are ElevenLabs voices and work with both v2 and v3. The character description is returned in the tagline field.

Each character's voice is pinned. Pass the slug as voice and you get the same voice with the same delivery, call after call. See Character voices and seeds for how it works.

For voices that require attribution, the required credit line is returned in the x-aicu-credit response header. The model actually used is returned in x-aicu-model.

Narrating long text (a full chapter, a whole book)​

Text longer than the per-request limit (10,000 characters on v2, 5,000 on v3) needs to be split, synthesized piece by piece, and concatenated. Here's how to turn a manuscript of tens or hundreds of thousands of characters into a single MP3.

Split on sentence boundaries​

Cutting mechanically by character count chops words in half and mangles pronunciation. Prefer sentence-ending punctuation, and avoid splitting across paragraphs, and the seams sound natural.

def split_chunks(text: str, limit: int = 4800):
chunks, buf = [], ""
for para in text.split("\n"):
para = para.strip()
if not para:
continue
if len(buf) + len(para) + 1 <= limit:
buf = f"{buf}\n{para}" if buf else para
continue
if buf:
chunks.append(buf); buf = ""
while len(para) > limit: # the paragraph itself is too long
cut = para.rfind("。", 0, limit)
cut = cut + 1 if cut > limit // 2 else limit
chunks.append(para[:cut]); para = para[cut:]
buf = para
if buf:
chunks.append(buf)
return chunks

Cache per chunk​

Long jobs always fail somewhere. Save each result under a filename derived from the hash of its text, and a re-run only synthesizes what's missing — and editing part of the manuscript only costs you the diff.

import hashlib, os, requests
from pathlib import Path

API = "https://api.aicu.ai/v1/tts/generate"
KEY = os.environ["AICU_API_KEY"] # requires scope: tts
VOICE, CACHE = "mina", Path(".cache")
CACHE.mkdir(exist_ok=True)

def synth(text: str, i: int) -> Path:
tag = hashlib.sha1((VOICE + text).encode()).hexdigest()[:8]
dest = CACHE / f"{i:03d}-{tag}.mp3"
if dest.exists():
return dest
r = requests.post(API,
headers={"Authorization": f"Bearer {KEY}"},
json={"text": text, "slug": VOICE, "format": "mp3"}, timeout=600)
r.raise_for_status()
dest.write_bytes(r.content)
print(f" {i}: {r.headers.get('x-credits-used')} AP")
return dest

Concatenate with the ffmpeg concat demuxer​

MP3s synthesized with the same settings can be joined without re-encoding (-c copy). It's fast and lossless.

import subprocess

def concat(parts, dest="out.mp3"):
listing = Path("concat.txt")
listing.write_text("".join(f"file '{p.resolve()}'\n" for p in parts))
subprocess.run(["ffmpeg", "-y", "-loglevel", "error",
"-f", "concat", "-safe", "0", "-i", str(listing),
"-c", "copy", dest], check=True)
listing.unlink()

Make chapters and headings seekable​

Measure each chunk's duration with ffprobe before concatenating and keep a running total, and you can implement "play from this heading." The player side just assigns audio.currentTime = startSeconds.

def duration(p) -> float:
out = subprocess.run(["ffprobe", "-v", "error", "-show_entries",
"format=duration", "-of", "csv=p=0", str(p)],
capture_output=True, text=True).stdout.strip()
return float(out or 0)

offsets, cursor = {}, 0.0
for i, ch in enumerate(chunks):
part = synth(ch, i)
offsets[i] = round(cursor, 2) # start second of this chunk
cursor += duration(part)

What it costs​

Billing is based on the number of characters you send, not on how long the audio turns out, so the same manuscript costs the same no matter which voice reads it. The actual charge comes back in the X-AICU-AP-Cost response header (x-credits-used carries the same number and is kept for older clients). For the current rate see Pricing.

Measured (the same 261-character Japanese text):

VoiceDurationExtrapolated to 120k characters
elena21.8sapprox. 2.8 hours
saki23.3sapprox. 3.0 hours
mei24.3sapprox. 3.1 hours
mina27.4sapprox. 3.5 hours
nao27.9sapprox. 3.6 hours

120,000 characters comes to a little over three hours of audio for roughly 1,100 AP (about ¥11). Cache hits are free, so editing the manuscript and regenerating only costs you the changed parts.

:::tip Clean up the text before narrating Feeding raw URLs, tables, and bullet markers straight in means the synthesizer reads the punctuation out loud, which is painful to listen to. A little preprocessing goes a long way: replace https?://\S+ with something like "(link omitted)", skip table markup, and strip leading # and - characters. :::

Speech-to-text (STT)​

POST /v1/audio/transcriptions (OpenAI-compatible / Whisper)​

Transcription runs on whisper-large-v3-turbo via Sakura AI Engine. Up to 30MB / 30 minutes per request. See https://api.aicu.ai/skills/stt for the machine-readable spec aimed at agents.

# Text only
curl -X POST https://api.aicu.ai/v1/audio/transcriptions \
-H "Authorization: Bearer aicu_live_xxx" \
-F file=@meeting.mp3 -F language=ja \
-F prompt="AICU, ComfyUI, しらいはかせ" # proper-noun hints (guards against misrecognition)
# → {"text": "the transcribed text…"}

# Segment-level timestamps (for subtitles and meeting minutes)
curl -X POST https://api.aicu.ai/v1/audio/transcriptions \
-H "Authorization: Bearer aicu_live_xxx" \
-F file=@meeting.mp3 -F language=ja \
-F response_format=verbose_json -F "timestamp_granularities[]=segment"
# → {"task":"transcribe","language":"ja","duration":60.0,"text":"…","segments":[{"id":0,"start":0.0,"end":4.32,"text":"…"}, …]}

# A subtitle file, directly
curl -X POST https://api.aicu.ai/v1/audio/transcriptions \
-H "Authorization: Bearer aicu_live_xxx" \
-F file=@meeting.mp3 -F language=ja -F response_format=srt -o meeting.srt
ParameterDescription
file (required)The audio file. Up to 30MB / 30 minutes (over that returns 413)
languageja, en, and so on. Specifying it improves both accuracy and speed
promptHints for proper nouns and domain terminology
response_formatjson (default) / verbose_json / text / srt / vtt
timestamp_granularities[]segment / word (only with verbose_json)

response_format and timestamp_granularities[] are passed straight through to the backend.

:::tip Long recordings (meetings, lectures) Don't upload the video. Extract the audio to mono 16kHz MP3 and split it into 30-minute pieces (65 minutes → 3 requests). Slicing into tiny 60-second pieces just multiplies your request count and makes you more likely to hit rate limits.

ffmpeg -y -i recording.mp4 -vn -map 0:a:0 -ac 1 -ar 16000 -b:a 64k audio.mp3 # 3GB → about 8MB
ffmpeg -y -i audio.mp3 -f segment -segment_time 1800 -c copy part_%02d.mp3 # 30 minutes each

:::

Errors and rate limits​

We return the real HTTP status (we never wrap a 429 in a 200 response body). curl -f, raise_for_status(), and the OpenAI SDK's automatic retries all just work.

HTTPMeaningWhat to do
413File is over 30MBExtract audio only, then split
429Rate limited (gateway: 60 req/min/key, or the backend's own quota)Wait Retry-After seconds and retry
502Backend failure (the upstream status is in error.upstream_status)Retry after a while. If it persists, contact us with the X-AICU-Request-Id

Response headers: X-AICU-AP-Cost (AP consumed), X-AICU-Request-Id, X-AICU-Latency-Ms, and X-RateLimit-Limit / -Remaining / -Reset.

Pricing​

10 seconds of audio = 3 AP (rounded up to 10-second units) since 2026-09-09. One hour ≈ 1,080 AP ≈ $0.11. With verbose_json we bill on the measured duration; other formats carry no duration, so we estimate from the audio we received and from the transcript length and bill on whichever is larger.

Drop-in with the OpenAI SDK​

from openai import OpenAI
client = OpenAI(base_url="https://api.aicu.ai/v1", api_key="aicu_live_xxx")

# TTS
speech = client.audio.speech.create(model="tts-1", voice="nao", input="こんにちは")
speech.stream_to_file("out.mp3")

# STT (with timestamps; on 429 the SDK reads Retry-After and retries for you)
with open("meeting.mp3", "rb") as f:
tr = client.audio.transcriptions.create(
model="whisper-large-v3-turbo", file=f, language="ja",
response_format="verbose_json", timestamp_granularities=["segment"],
)
for seg in tr.segments:
print(f"{seg.start:8.2f} --> {seg.end:8.2f} {seg.text}")

Swap the base URL and API key, and your existing OpenAI-compatible code works as-is.

© 2026 AICU Inc.