Character voices and seed
With eleven_v3, synthesizing the same sentence with the same voice returns different audio every time.
Both timbre and energy change noticeably.
For a single narration track that hardly matters. But if it is meant to be this character's voice, used over and over, a different person on every call does not work.
This page is about how to stop that.
:::tip The short version
- AICU registered voices (
elenaand the rest) already have a fixed seed. Just pass the slug invoice - To pin it yourself, set
seedexplicitly. Same seed + same text = same audio - If you changed the seed and the audio did not change, suspect the cache (
force: true) :::
:::warning The parameter name changes with the endpoint
Which field carries the voice depends on which endpoint you call.
| Endpoint | Voice field | Text field |
|---|---|---|
/v1/audio/speech, /v1/tts/speech (OpenAI-compatible) | voice | input |
/v1/tts/generate (AICU's own) | slug | text |
Sending voice to /v1/tts/generate is not an error. It is silently ignored and you get the
default voice — 200, valid audio, no warning. That is a hard thing to notice.
Measured 2026-09-22, with seed fixed so the output is deterministic:
| Request | Response size (two runs) |
|---|---|
| no voice field | 22,196 / 22,196 bytes |
"voice": "mei" | 22,196 / 22,196 bytes — identical to no voice at all |
"slug": "mei" | 26,794 / 26,794 bytes |
If you batch a long script before noticing, you pay for all of it in the wrong voice. Generate one line first and check that the audio actually changed before you run the volume.
A wrong value does fail properly — {"slug": "no_such_character"} returns
400 {"error":"Unknown character: no_such_character. GET /v1/tts/voices を参照"}.
Only the wrong field name passes silently.
:::
What happens
Send the same request three times and compare the hashes of the returned MP3s.
for i in 1 2 3; do
curl -s -X POST https://api.aicu.ai/v1/audio/speech \
-H "Authorization: Bearer $AICU_API_KEY" -H "Content-Type: application/json" \
-d '{"voice": "mina", "model": "eleven_v3", "input": "こんにちは", "seed": '"$RANDOM"'}' \
| md5sum
done
Scatter the seed and all three hashes differ — and they sound like different people.
Fix seed and all three hashes match.
Default seeds for the AiCuty voices
AICU registered voices ship with a seed we have already picked. Omit seed and that one is used.
from openai import OpenAI
client = OpenAI(base_url="https://api.aicu.ai/v1", api_key="aicu_live_…")
# No need to write a seed. Same voice however many times you call
client.audio.speech.create(
model="eleven_v3", voice="mina", input="こんにちは"
).stream_to_file("mina.mp3")
| slug | Name | Character of the voice | v3 | v2 |
|---|---|---|---|---|
elena | Elena Bloom (エレナ・ブルーム) | Calm and easy to follow; standard-Japanese female narration | 330 | 330 |
mei | Mei Soleil (メイ・ソレイユ) | Bright young girl's voice, energetic and easy even for children to follow | 721 | 721 |
mina | Mina Azur (ミナ・アズール) | Calm, intelligent narration; no listening fatigue over long stretches | 101 | 105 |
nao | Nao Verde (ナオ・ヴェルデ) | Soft voice of a gentle, science-minded young man | 55 | 55 |
saki | Saki Noir (サキ・ノワール) | Whispering voice, gentle with a hint of mystery | 1033 | 1035 |
marsha | Marsha Arancia (マーシャ・アランチャ) | Fresh, energetic and easy-to-listen female voice | 418 | 414 |
luc4 | LuC4 (ルカ/全力肯定彼氏くん) | Pleasant low male voice | 1104 | 1104 |
You can hear the actual voices at AiCuty official voices (Japanese, English and French; Marsha also has Spanish and Italian).
Why v2 and v3 use different values
The same seed delivers a different voice on a different model. A take that sounded good on v3 is not necessarily good on v2, so we re-picked for each.
Mina is 101 on v3 and 105 on v2. Saki is 1033 on v3 and 1035 on v2.
How the values are chosen
We start from the character's birthday, month and day concatenated, then audition nearby values.
Mina was born on October 1, so 101 is the starting point, and from there we listen through 102, 103, 104, 105.
Birthday-derived values are easy to remember, easy to explain, and do not collide between characters.
Pinning your own voice
When you pass a raw ElevenLabs voice_id there is no default seed. Pick one yourself and state it explicitly.
curl -X POST https://api.aicu.ai/v1/audio/speech \
-H "Authorization: Bearer $AICU_API_KEY" -H "Content-Type: application/json" \
-d '{"voice": "pqHfZKP75CvOlQylNhV4", "model": "eleven_v3", "input": "…", "seed": 42}' \
--output out.mp3
There is no single right way to choose, but committing to a value and writing it down matters more than the value itself. The worst outcome is being unable to reproduce why the voice sounds the way it does.
Audition a few candidates before you decide. Even neighbouring seeds can give a very different impression.
When you change the seed and the audio does not change
Suspect the cache.
api.aicu.ai caches on a single key of "same voice × same model × same text × same seed".
Cache hits are not billed (x-aicu-ap-cost: 0).
curl -sS -D - -o /dev/null -X POST https://api.aicu.ai/v1/audio/speech \
-H "Authorization: Bearer $AICU_API_KEY" -H "Content-Type: application/json" \
-d '{"voice": "mina", "model": "eleven_v3", "input": "こんにちは"}' \
| grep -i 'x-aicu-ap-cost\|x-cache-hit'
If x-aicu-ap-cost: 0 comes back, the audio was served from cache.
When you really do need it regenerated, use force: true on /v1/tts/generate.
:::caution There was a period when the audio you got back was not what you asked for
The old gateway neither forwarded seed upstream nor included it in the cache key.
As a result, changing the seed still returned the previous audio — and since the response was a 200, there was no way to notice.
Both the forwarding and the cache key are fixed now.
:::
Lining several voices up together
If you are stitching audio from several characters, normalize loudness clip by clip before you mix. Levels differ between voices, and concatenating as-is leaves only the loud one standing out.
ffmpeg -i in.mp3 -af "loudnorm=I=-18:TP=-2:LRA=9" -ac 1 -ar 44100 out.mp3
A crossfade of about 0.1 seconds at the joins sounds natural. A plain concatenation chops off the tail of each breath.
ffmpeg -i a.mp3 -i b.mp3 -filter_complex "[0][1]acrossfade=d=0.1:c1=tri:c2=tri" out.mp3
:::note Do not use -c copy to concatenate mp3
With the concat demuxer plus -c copy, mp3 encoder delay rewinds the timestamps
(non monotonically increasing dts) and the joins break up. Rebuild them with the concat filter instead.
:::
See also
- Audio API — endpoints and handling long text
- Mastering Japanese TTS — fixing readings and controlling prosody on v3
- Character API — per-character usage records and royalty allocation