Skip to main content

Character voices and seed

With eleven_v3, synthesizing the same sentence with the same voice returns different audio every time. Both timbre and energy change noticeably.

For a single narration track that hardly matters. But if it is meant to be this character's voice, used over and over, a different person on every call does not work.

This page is about how to stop that.

:::tip The short version

  • AICU registered voices (elena and the rest) already have a fixed seed. Just pass the slug in voice
  • To pin it yourself, set seed explicitly. Same seed + same text = same audio
  • If you changed the seed and the audio did not change, suspect the cache (force: true) :::

:::warning The parameter name changes with the endpoint

Which field carries the voice depends on which endpoint you call.

EndpointVoice fieldText field
/v1/audio/speech, /v1/tts/speech (OpenAI-compatible)voiceinput
/v1/tts/generate (AICU's own)slugtext

Sending voice to /v1/tts/generate is not an error. It is silently ignored and you get the default voice — 200, valid audio, no warning. That is a hard thing to notice.

Measured 2026-09-22, with seed fixed so the output is deterministic:

RequestResponse size (two runs)
no voice field22,196 / 22,196 bytes
"voice": "mei"22,196 / 22,196 bytes — identical to no voice at all
"slug": "mei"26,794 / 26,794 bytes

If you batch a long script before noticing, you pay for all of it in the wrong voice. Generate one line first and check that the audio actually changed before you run the volume.

A wrong value does fail properly — {"slug": "no_such_character"} returns 400 {"error":"Unknown character: no_such_character. GET /v1/tts/voices を参照"}. Only the wrong field name passes silently.

:::

What happens​

Send the same request three times and compare the hashes of the returned MP3s.

for i in 1 2 3; do
curl -s -X POST https://api.aicu.ai/v1/audio/speech \
-H "Authorization: Bearer $AICU_API_KEY" -H "Content-Type: application/json" \
-d '{"voice": "mina", "model": "eleven_v3", "input": "こんにちは", "seed": '"$RANDOM"'}' \
| md5sum
done

Scatter the seed and all three hashes differ — and they sound like different people. Fix seed and all three hashes match.

Default seeds for the AiCuty voices​

AICU registered voices ship with a seed we have already picked. Omit seed and that one is used.

from openai import OpenAI

client = OpenAI(base_url="https://api.aicu.ai/v1", api_key="aicu_live_…")

# No need to write a seed. Same voice however many times you call
client.audio.speech.create(
model="eleven_v3", voice="mina", input="こんにちは"
).stream_to_file("mina.mp3")
slugNameCharacter of the voicev3v2
elenaElena Bloom (エレナ・ブルーム)Calm and easy to follow; standard-Japanese female narration330330
meiMei Soleil (メイ・ソレイユ)Bright young girl's voice, energetic and easy even for children to follow721721
minaMina Azur (ミナ・アズール)Calm, intelligent narration; no listening fatigue over long stretches101105
naoNao Verde (ナオ・ヴェルデ)Soft voice of a gentle, science-minded young man5555
sakiSaki Noir (サキ・ノワール)Whispering voice, gentle with a hint of mystery10331035
marshaMarsha Arancia (マーシャ・アランチャ)Fresh, energetic and easy-to-listen female voice418414
luc4LuC4 (ルカ/全力肯定彼氏くん)Pleasant low male voice11041104

You can hear the actual voices at AiCuty official voices (Japanese, English and French; Marsha also has Spanish and Italian).

Why v2 and v3 use different values​

The same seed delivers a different voice on a different model. A take that sounded good on v3 is not necessarily good on v2, so we re-picked for each.

Mina is 101 on v3 and 105 on v2. Saki is 1033 on v3 and 1035 on v2.

How the values are chosen​

We start from the character's birthday, month and day concatenated, then audition nearby values. Mina was born on October 1, so 101 is the starting point, and from there we listen through 102, 103, 104, 105.

Birthday-derived values are easy to remember, easy to explain, and do not collide between characters.

Pinning your own voice​

When you pass a raw ElevenLabs voice_id there is no default seed. Pick one yourself and state it explicitly.

curl -X POST https://api.aicu.ai/v1/audio/speech \
-H "Authorization: Bearer $AICU_API_KEY" -H "Content-Type: application/json" \
-d '{"voice": "pqHfZKP75CvOlQylNhV4", "model": "eleven_v3", "input": "…", "seed": 42}' \
--output out.mp3

There is no single right way to choose, but committing to a value and writing it down matters more than the value itself. The worst outcome is being unable to reproduce why the voice sounds the way it does.

Audition a few candidates before you decide. Even neighbouring seeds can give a very different impression.

When you change the seed and the audio does not change​

Suspect the cache.

api.aicu.ai caches on a single key of "same voice × same model × same text × same seed". Cache hits are not billed (x-aicu-ap-cost: 0).

curl -sS -D - -o /dev/null -X POST https://api.aicu.ai/v1/audio/speech \
-H "Authorization: Bearer $AICU_API_KEY" -H "Content-Type: application/json" \
-d '{"voice": "mina", "model": "eleven_v3", "input": "こんにちは"}' \
| grep -i 'x-aicu-ap-cost\|x-cache-hit'

If x-aicu-ap-cost: 0 comes back, the audio was served from cache. When you really do need it regenerated, use force: true on /v1/tts/generate.

:::caution There was a period when the audio you got back was not what you asked for The old gateway neither forwarded seed upstream nor included it in the cache key. As a result, changing the seed still returned the previous audio — and since the response was a 200, there was no way to notice. Both the forwarding and the cache key are fixed now. :::

Lining several voices up together​

If you are stitching audio from several characters, normalize loudness clip by clip before you mix. Levels differ between voices, and concatenating as-is leaves only the loud one standing out.

ffmpeg -i in.mp3 -af "loudnorm=I=-18:TP=-2:LRA=9" -ac 1 -ar 44100 out.mp3

A crossfade of about 0.1 seconds at the joins sounds natural. A plain concatenation chops off the tail of each breath.

ffmpeg -i a.mp3 -i b.mp3 -filter_complex "[0][1]acrossfade=d=0.1:c1=tri:c2=tri" out.mp3

:::note Do not use -c copy to concatenate mp3 With the concat demuxer plus -c copy, mp3 encoder delay rewinds the timestamps (non monotonically increasing dts) and the joins break up. Rebuild them with the concat filter instead. :::

See also​