Skip to main content

We Benchmarked Japanese STT on One Shared Audio Source — ElevenLabs Scribe Pulled Ahead

· 3 min read
AICU API Team
api.aicu.ai

To choose the backend for AICU API transcription (POST /v1/audio/transcriptions), we measured the major STT APIs against the same Japanese audio and the same ground truth. The verdict: ElevenLabs Scribe was the most accurate, at a CER (character error rate) of 15.31%. We are now evaluating it as a high-accuracy lane in AICU API.

Method — use audio whose ground truth already exists​

The hardest part of any STT benchmark is producing the ground truth. So we took the article announcing AiCuty's newest member — Marsha Arancia's appointment as AICU STUDY navigator (3,023 characters) — and turned it into audio by having Marsha herself read it aloud (AICU API TTS, voice: "marsha"). 8.2 minutes, 7.9 MB of mp3. The script is the ground truth, so CER can be computed mechanically.

  • CER = character error rate (lower is better). Levenshtein distance after NFKC normalization and removal of punctuation and whitespace, divided by the ground truth character count
  • Every provider received the same file, once. The proper-noun hint (prompt) was identical as well
  • Quirks in the TTS reading (for example, how "AiCuty" gets pronounced) are a handicap shared by all providers — so read this as a relative comparison

Results​

Provider / modelCERProcessing timeList price (per hour of audio)
ElevenLabs Scribe (scribe_v1)15.31%16.7s$0.22–0.40
OpenAI whisper-121.89%29.4s$0.36
OpenAI gpt-4o-mini-transcribe23.40%16.9s$0.18
whisper-large-v3-turbo (via the current AICU API backend)64.97%9.9s—

The 64.97% via the current backend is not a reflection of model quality: roughly 40% of the body text was dropped at the seams where long audio is split internally (whisper-1, from the same model family, scored 21.89%). We are reviewing how that path should be handled, and as a first step we have documented the limits (30 MB / 30 minutes per request) and a chunking recipe in /skills/stt.

What Scribe got right​

  • Nothing dropped: no omissions across the full 8.2 minutes. It holds up on long audio (the spec allows up to 10 hours / 3 GB per file)
  • Strong on spoken language: it picks up colloquial phrasing and hesitations accurately — lines like "……と名乗ったところで気づいたんだけど"
  • Word- and character-level timestamps, speaker diarization (up to 32 speakers), and audio event tags such as laughter

For balance, one weakness: it likes to romanize personal names ("マーシャ・アランチャ" → Marcia Arancja). For subtitling, that is well within what a post-processing dictionary can absorb.

What AICU API will do​

  1. No automatic fallback. The backend will be selectable explicitly via the model parameter (silently switching you to a different model is the scariest outcome of all — that is our design principle)
  2. Evaluating a high-accuracy lane on Scribe (pricing to be announced separately)
  3. Pricing change, announced in advance: from 2026-09-09 05:00 UTC, transcription will be billed at 3 AP per 10 seconds of audio (≈ $0.11 per hour, under the cost×3 policy). Details on News

The reproduction script and the full logs live in the repository at scripts/stt-benchmark.py and docs/internal/stt-benchmark-2026-08.md (internal). Round two will be measured on real meeting audio.

Thanks to Marsha Arancia (AiCuty's editorial and publishing member / AICU STUDY navigator) for lending her voice to the audio source.