We Benchmarked Japanese STT on One Shared Audio Source — ElevenLabs Scribe Pulled Ahead
To choose the backend for AICU API transcription (POST /v1/audio/transcriptions), we measured the major STT APIs against the same Japanese audio and the same ground truth. The verdict: ElevenLabs Scribe was the most accurate, at a CER (character error rate) of 15.31%. We are now evaluating it as a high-accuracy lane in AICU API.
Method — use audio whose ground truth already exists
The hardest part of any STT benchmark is producing the ground truth. So we took the article announcing AiCuty's newest member — Marsha Arancia's appointment as AICU STUDY navigator (3,023 characters) — and turned it into audio by having Marsha herself read it aloud (AICU API TTS, voice: "marsha"). 8.2 minutes, 7.9 MB of mp3. The script is the ground truth, so CER can be computed mechanically.
- CER = character error rate (lower is better). Levenshtein distance after NFKC normalization and removal of punctuation and whitespace, divided by the ground truth character count
- Every provider received the same file, once. The proper-noun hint (
prompt) was identical as well - Quirks in the TTS reading (for example, how "AiCuty" gets pronounced) are a handicap shared by all providers — so read this as a relative comparison
Results
| Provider / model | CER | Processing time | List price (per hour of audio) |
|---|---|---|---|
| ElevenLabs Scribe (scribe_v1) | 15.31% | 16.7s | $0.22–0.40 |
| OpenAI whisper-1 | 21.89% | 29.4s | $0.36 |
| OpenAI gpt-4o-mini-transcribe | 23.40% | 16.9s | $0.18 |
| whisper-large-v3-turbo (via the current AICU API backend) | 64.97% | 9.9s | — |
The 64.97% via the current backend is not a reflection of model quality: roughly 40% of the body text was dropped at the seams where long audio is split internally (whisper-1, from the same model family, scored 21.89%). We are reviewing how that path should be handled, and as a first step we have documented the limits (30 MB / 30 minutes per request) and a chunking recipe in /skills/stt.
What Scribe got right
- Nothing dropped: no omissions across the full 8.2 minutes. It holds up on long audio (the spec allows up to 10 hours / 3 GB per file)
- Strong on spoken language: it picks up colloquial phrasing and hesitations accurately — lines like "……と名乗ったところで気づいたんだけど"
- Word- and character-level timestamps, speaker diarization (up to 32 speakers), and audio event tags such as laughter
For balance, one weakness: it likes to romanize personal names ("マーシャ・アランチャ" → Marcia Arancja). For subtitling, that is well within what a post-processing dictionary can absorb.
What AICU API will do
- No automatic fallback. The backend will be selectable explicitly via the
modelparameter (silently switching you to a different model is the scariest outcome of all — that is our design principle) - Evaluating a high-accuracy lane on Scribe (pricing to be announced separately)
- Pricing change, announced in advance: from 2026-09-09 05:00 UTC, transcription will be billed at 3 AP per 10 seconds of audio (≈ $0.11 per hour, under the cost×3 policy). Details on News
The reproduction script and the full logs live in the repository at scripts/stt-benchmark.py and docs/internal/stt-benchmark-2026-08.md (internal). Round two will be measured on real meeting audio.
Thanks to Marsha Arancia (AiCuty's editorial and publishing member / AICU STUDY navigator) for lending her voice to the audio source.
