Skip to main content

Mastering Japanese TTS

Japanese speech synthesis has two stumbling blocks that rarely come up in English.

  1. Wrong readings — 「入口」 comes out as 「にゅうこう」, and a company name gets read the way it is spelled
  2. Broken prosody — the words are correct, but a question flattens out

This page collects how to fix both, based on actual measurements — including per-voice differences in how tags land, which are not in the official documentation.

:::tip The short version

  • For readings, substitute kana before you send the text — the most reliable route, and steadier than a pronunciation dictionary
  • Prosody can only be touched with eleven_v3 audio tags — a pronunciation dictionary cannot do it
  • Tags behave differently per voice. Always measure on the voice you actually use :::

0. Try it in one call​

The gateway already fixes the most common readings for you. Text that contains Japanese goes through a server-side reading dictionary before synthesis (「AICU」→「アイキュー」, 「ChatGPT」→「チャットジーピーティー」, …), so the two curls below say the product names correctly without any client-side work.

# Multilingual v2 (default): readings fixed by the server-side dictionary
curl -X POST https://api.aicu.ai/v1/audio/speech \
-H "Authorization: Bearer $AICU_API_KEY" \
-H "Content-Type: application/json" \
-d '{"input": "AICU API の入口はこちらです", "voice": "nao"}' \
--output v2.mp3

# v3: same text, plus an audio tag for prosody (only eleven_v3 understands tags)
curl -X POST https://api.aicu.ai/v1/el/tts \
-H "Authorization: Bearer $AICU_API_KEY" \
-H "Content-Type: application/json" \
-d '{"text": "[excited] AICU API の入口はこちらです!", "voice": "nao", "model": "eleven_v3"}' \
--output v3.mp3

The dictionary is small and project-specific (the gateway's yomi dictionary). It only touches text that contains kana or kanji, and it replaces longer entries first. For names it does not know, substitute kana yourself before sending — the rest of this page explains why that beats a pronunciation dictionary.

1. Fixing readings​

The official route: the pronunciation dictionary​

ElevenLabs supports the W3C standard PLS (Pronunciation Lexicon Specification). Upload a .pls file, get a dictionary ID back, and reference it at synthesis time.

<?xml version="1.0" encoding="UTF-8"?>
<lexicon version="1.0" xmlns="http://www.w3.org/2005/01/pronunciation-lexicon"
alphabet="ipa" xml:lang="ja-JP">
<lexeme><grapheme>入口</grapheme><alias>いりぐち</alias></lexeme>
</lexicon>

There are three mechanisms for overriding a reading.

MechanismSupported modelsUsable in Japanese?
alias (substitute a different spelling)All models△ Does not always apply as expected
phoneme (phonetic symbols)eleven_flash_v2 only✗ Unavailable on multilingual models
IPA (written inline in the text)eleven_v3△ Officially documented as only 80–90% consistent

Measured: kana substitution was the most reliable in Japanese​

Registering an alias of 入口 → いりぐち and synthesizing produced the half-finished reading 「にゅうぐち」. The original misreading survived in the first half — the alias applied only partially.

No intervention (「にゅうこう」)

PLS alias applied (「にゅうぐち」)

Substituted with kana (「いりぐち」)

Substitute before you send and the model reads it plainly. Preprocessing is the most reliable route.

YOMI = {
"AICU media": "アイキューメディア", # register compound terms first
"AICU": "アイキュー",
"ComfyUI": "コンフィーユーアイ",
"入口": "いりぐち",
"curl": "しーゆーあーるえる",
}

def apply_yomi(text: str) -> str:
for src in sorted(YOMI, key=len, reverse=True): # replace longest terms first
text = text.replace(src, YOMI[src])
return text

:::caution Replace the longest terms first If AICU is processed first, AICU media becomes 「アイキュー media」. That is exactly what sorted(..., key=len, reverse=True) is there for. :::

AICU as-is

Substituted with アイキュー

2. Fixing prosody​

The pronunciation dictionary cannot handle prosody​

The W3C PLS specification states this explicitly in §1.4, "What PLS does not Support".

The most complex features have been postponed to a future revision of this specification.

PLS only has <grapheme>, <phoneme> and <alias>; there are no elements for prosody, accent or intonation anywhere in the specification.

For now, the only way to touch prosody in Japanese is eleven_v3 audio tags.

The audio tag list​

[laughs] [laughs harder] [starts laughing] [wheezing] [snorts]
[whispers] [sighs] [exhales] [sarcastic] [curious] [excited]
[crying] [mischievously]
[gunshot] [applause] [clapping] [explosion] [swallows] [gulps]
[strong X accent] [sings]

The official line is that this is not exhaustive and you should experiment — and indeed [softly], [warmly], [happy], [breathless] and [questioning] get a response too.

Measured: tags land differently on different voices​

Using the same sentence 「いりぐちはどちらですか?」 and swapping only the tag, we measured duration against the untagged take.

tagMinaSakiNotes
laughs3.0×1.7×Strongest response on Mina
happy2.2×1.6×
sarcastic2.2×1.1×Barely registers on Saki
breathless2.2×1.2×Same as above
questioning2.2×1.0×Same as above
excited2.1×1.6×Works on both
sighs1.3×2.0×Only Saki responds strongly (reversed)
cheerfully1.1×1.5×Barely registers on Mina
softly1.4×1.7×
curious / whispers1.3×1.2×Subtle

sighs is reversed: Mina 1.3× against Saki 2.0×. The official caveat that a tag working on one voice may not work on another shows up directly in the numbers.

Audio for all 16 tags × 2 voices is here.

  • Mina: /audio/ja-tts/tags/mina/00-none.mp3 through 15-questioning.mp3
  • Saki: /audio/ja-tts/tags/saki/00-none.mp3 through 15-questioning.mp3

:::note Duration is a proxy metric The multipliers in the table are playback duration relative to the untagged take — only a rough indicator of how far the tag landed. Always make the final call by ear. cheerfully (Mina 1.1×), for instance, was barely audible as a change at all. :::

Running a prosody dictionary​

Phrases whose reading is correct but whose prosody collapses can be corrected by inserting a tag immediately before the phrase.

Untagged, 「いりぐちはどちらですか?」 flattens out; add [softly] and it rises naturally as a question. Confirmed on both Mina and Saki.

PROSODY = {
"いりぐちはどちらですか": "softly",
}

def apply_prosody(text: str) -> str:
"""eleven_v3 only. v2 does not interpret tags and reads them out verbatim"""
for phrase in sorted(PROSODY, key=len, reverse=True):
text = text.replace(phrase, f"[{PROSODY[phrase]}] {phrase}")
return text

:::danger Never pass tags to v2 eleven_multilingual_v2 does not interpret audio tags. It reads the literal string [softly] aloud. Tags are v3-only. :::

A combined example​

Mina

Saki

[breathless] せいせいエーアイ時代に、[happy] つくる人をつくる!
[excited] げっかんアイキューのはいたつに来ました!
[curious] いりぐちは、どちらですか?

Readings (kana substitution) and prosody (tags) are easier to handle as two separate dictionaries, because their scope differs: readings apply to every model, prosody only to v3.

3. Summary​

What you wantHowSupported models
Readings of proper nouns and irregular kanji compoundsSubstitute kana before sendingAll models
Use a pronunciation dictionaryUpload PLS (.pls)All models (unreliable in Japanese)
Prosody and emotionaudio tageleven_v3 only
Precise accent and prosody controlNo option (outside the PLS spec)—

Misreadings and broken prosody are both project-specific. Add one line to the dictionary every time you find one and accuracy climbs with each pass.

See also​