Batch translation
Translating a documentation set is a few hundred calls to
/v1/chat/completions. The calls are the easy part. This page is about the three things that decide whether the result is usable: knowing the cost before you spend it, checking that the Markdown survived, and checking that the model actually translated anything.
Every number here was measured on 2026-09-23 against this site's own documentation tree — 17 files, 134,713 characters. The complete script is the one we run, and you can download it: translate-docs.py (MIT-style, no dependencies beyond the standard library).
What it costs, and what actually happened
We ran the whole set on 2026-09-23. This is the real output, not an estimate:
| Source | 17 files / 134,713 characters |
| Model | gpt-4o-mini |
| Measured rate | 18 AP per 2,500 characters |
| Whole set, one language | ≈ 970 AP (about US$0.10) |
| Files written | 15 |
| Files refused by the checks | 2 |
Two files out of seventeen came back damaged and were not written. Both looked completely normal: correct beginning, correct ending, finish_reason: stop, no error of any kind. That is the reason this page is mostly about verification rather than about the calls.
The call
One file, one request. Nothing streams, nothing chunks — a documentation page fits in the context window, and splitting it is how heading levels and table columns get lost at the seams.
curl -X POST https://api.aicu.ai/v1/chat/completions \
-H "Authorization: Bearer aicu_live_xxx" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"temperature": 0,
"max_tokens": 8000,
"messages": [
{"role": "system", "content": "You translate technical API documentation into Korean. Translate prose only. Reproduce byte-for-byte: YAML frontmatter, anything inside ``` fences, URLs, file paths, HTTP methods and status codes, and API names in backticks. Keep the Markdown structure identical. Output the translated document only."},
{"role": "user", "content": "<the whole .md file>"}
]
}'
temperature: 0 matters more here than in most uses. You want the same input to give the same output, so that re-running the job does not produce a diff made of synonyms.
The AP the call cost is on the response, so you never have to estimate after the fact:
X-AICU-AP-Cost: 18
X-AICU-AP-Balance: 6263161
Check that the structure survived
A bad translation announces itself as broken Markdown long before it announces itself as bad prose. A dropped fence swallows the rest of the page; a table that loses its separator row renders as a paragraph of pipes.
So compare the output against the input mechanically, and refuse to write the file if it fails:
def structure_check(src: str, out: str) -> list[str]:
bad = []
if src.count("```") != out.count("```"):
bad.append("code fences")
if src.count("|---") != out.count("|---"):
bad.append("tables")
for level in ("\n## ", "\n### "):
if src.count(level) != out.count(level):
bad.append(f"headings {level.strip()}")
if src.startswith("---") and not out.startswith("---"):
bad.append("frontmatter")
for url in set(re.findall(r"https?://[^\s)\"'`]+", src)):
if url not in out:
bad.append(f"lost URL {url}")
break
if out.strip().startswith("```"):
bad.append("whole document wrapped in a fence")
return bad
A file that fails is reported and left unwritten. Not translating a page is recoverable; publishing a broken one is not.
What it caught, in practice
The failures were not truncation. Both documents were translated to the last line and stopped normally. The model simply removed things from the middle.
A whole table, gone. pricing.md came back 172 lines shorter by one table — the final "how the estimate is built" table was absent, while the paragraph after it was translated and present. Separator rows 16 → 14, table rows 42 → 38.
A link stripped to plain text. opencode.md kept every table, every row and every heading. Inside one row, this:
| API key | `aicu_live_xxx` (issue one in the [dashboard](https://api.aicu.ai/dashboard/keys) — **required**) |
came back as this:
| API 키 | `aicu_live_xxx` (대시보드에서 발급 — **필수**) |
The sentence still tells the reader to get a key from the dashboard. The link to the dashboard is gone. Row counts, heading counts, fence counts and length are all unchanged — only the URL-survival check sees this one.
That is the argument for running several cheap checks rather than one clever one. Each of the three failures on this site was caught by a different check, and no single check caught more than one of them.
One of our own checks was wrong
Worth saying, because it is the other half of the same lesson. On the first run a third file failed with lost URL https://api.aicu.ai/dashboard/keys. — note the trailing period. The source had the URL at the end of a sentence, and the regex had swallowed the full stop as part of the URL. The translation ended the sentence with its own punctuation, so the literal string was absent and a perfectly good file was rejected.
url = url.rstrip(".,;:!?、。)」") # the sentence's punctuation is not part of the URL
A check that fails good output gets switched off. Tune for zero false positives and accept that you will miss things — the checks here were verified against 15 known-good translations before being trusted to reject anything.
What not to check
Length. It is the obvious idea and it does not work on Markdown: across the 15 good translations the line-count ratio ranged from 0.934 to 1.000, because hard-wrapped English paragraphs get rejoined into single lines in the target language. The dropped-table file sat at 0.820 — catchable — but the stripped-link file sat at 0.974, comfortably inside the normal range. A threshold tight enough to catch the second would reject several good files.
Count structural elements instead. Table rows, bullets and headings matched exactly on all 15 good translations, which makes any difference a real signal:
for label, count in (("table rows", lambda t: sum(1 for l in t.splitlines() if l.lstrip().startswith("|"))),
("bullets", lambda t: sum(1 for l in t.splitlines() if re.match(r"\s*[-*] ", l))),
("headings", lambda t: sum(1 for l in t.splitlines() if l.startswith("#")))):
if count(src) != count(out):
bad.append(f"{label} {count(src)}→{count(out)}")
Frontmatter comes back untouched — including the parts readers see
The system prompt reproduces YAML frontmatter byte-for-byte, which is right for sidebar_position, service_status and related, and wrong for title and description — those are rendered on the page and in search results. Translate those two afterwards, or exempt them in the prompt. Whichever you choose, check the rendered page in the target locale rather than the file.
One thing that is not a failure
The model translates human-language strings inside code blocks — prompt values, comments — while leaving identifiers, parameter names and model names alone:
- "prompt": "a red apple on a white table, product photo",
+ "prompt": "하얀 테이블 위에 빨간 사과, 제품 사진",
The system prompt asks for fenced blocks byte-for-byte, so strictly this is disobedience. It is also what a Korean reader wants from an example. We allow it, and check the things that would actually break: fence count, and that "character": "SakiNoir" and gpt-4o-mini come back untouched. Decide which side you want and write the check to match — but decide, rather than discovering it later in a rendered page.
Check that it was actually translated
This is the check that is easy to leave out, and the one that cost us the most to learn.
Benchmarking four models on English → Korean, gpt-5-mini scored 1.34 / 0.49 / 0.69 where the others scored around 1.85 / 1.55 / 0.27. That reads like bad prose. It was not. The output contained zero Hangul characters and was 2,499 characters long against a 2,500-character input: it had returned the English source unchanged.
structure_check passes that output perfectly — the structure is identical because the document is the original. Write it out and i18n/ko/ fills up with English wearing a Korean filename. Nothing downstream complains.
Two cheap tests catch it:
def translated_check(src: str, out: str, locale: str) -> list[str]:
bad = []
# 1. Is the prose simply the source? (language-independent)
a, b = strip_code_and_urls(src), strip_code_and_urls(out)
same = sum(1 for x, y in zip(a, b) if x == y) / max(len(a), len(b))
if same > 0.9:
bad.append(f"source returned unchanged ({same:.0%} identical)")
# 2. Does the target script actually appear?
rng = TARGET_SCRIPT.get(locale) # ko: Hangul, zh: Han, ja: kana
if rng and not bad:
lo, hi = rng
n = sum(1 for c in out if lo <= ord(c) <= hi)
if n / max(len(out), 1) < 0.03:
bad.append(f"almost no target-script characters ({n})")
return bad
The first test works for any language pair. The second only works where the target has its own script — French, Spanish and Portuguese share the Latin alphabet with the source, so there the identity test is all you get. It is enough: a model that fails to translate fails by echoing, and echoing is what test 1 detects.
The general form of the lesson is worth stating plainly, because it applies to every batch job you run against a language model:
A job that fails loudly is cheap. A job that succeeds falsely is expensive. Spend your verification effort on the outputs that look like success.
Do not pay twice for a file that has not changed
Store the SHA-256 of each source file next to the language you translated it into, and skip anything that matches:
digest = hashlib.sha256(src.encode()).hexdigest()
if manifest.get(str(path)) == digest and dest.exists():
continue # unchanged since the last run — skip
The second run of an unchanged tree costs 0 AP. This turns translation from an event into a habit: edit three pages, re-run, pay for three pages.
Keep the manifest in version control. It is the record of which translation corresponds to which revision of the source, which is the question you will actually want answered six months later.
Choosing the model
Measure, on your own documents. Four models, the same four files, English → Korean, scored by Jev:
| model | faithful | natural | untranslated | AP / 2,500 chars | s |
|---|---|---|---|---|---|
gpt-4o-mini | 1.83 | 1.55 | 0.26 | 18 | 6.9 |
gpt-4o | 1.84 | 1.56 | 0.30 | 278 | 9.9 |
gpt-5.6-sol | 1.87 | 1.66 | 0.25 | 849 | 9.2 |
gpt-5-mini | 1.34 | 0.49 | 0.69 | 46 | 9.6 |
gpt-5.6-sol is the best on all three axes. It is also 47× the price of gpt-4o-mini for margins of +0.04, +0.11 and −0.01 — 970 AP against 45,700 AP for the whole set.
The gpt-5-mini row is the one from the section above: those are not prose scores, they are the score for handing back the source.
We default to gpt-4o-mini and reach for gpt-5.6-sol on individual pages where the wording carries weight. The difference it buys is real but small, and it shows up as register rather than accuracy — supplying a subject the English left implicit, splitting a long dashed sentence into two, choosing the idiomatic particle.
Scoring the output
The three scores above come from /v1/jev, which answers typed questions rather than returning prose you then have to parse:
{
"state": {"source": "...", "translation": "..."},
"questions": {
"faithful": {
"type": "score",
"instructions": "Does the translation preserve the meaning of the source, including technical details?",
"criteria": ["Meaning is distorted", "Mostly faithful", "Faithful throughout"]
},
"natural": {
"type": "score",
"instructions": "Does the translation read naturally to a developer who speaks that language?",
"criteria": ["Stilted", "Readable but awkward", "Natural"]
},
"untranslated": {
"type": "noul",
"instructions": "Is there source-language prose left untranslated that should have been translated? (Code, URLs and parameter names are correctly left as-is.)"
}
}
}
About 8 AP per file, so scoring a whole set costs less than translating one page of it.
Use it to compare models and to spot outliers, not as the gate. The mechanical checks decide whether a file is written; a score is a number to look at afterwards. A model that returns the source unchanged is caught by four lines of character counting, not by asking another model whether it looks alright.
Running it
Download translate-docs.py — one file, standard library only.
export AICU_API_KEY=aicu_live_... # llm scope
python3 scripts/translate-docs.py --to ko # dry-run: cost estimate only
python3 scripts/translate-docs.py --to ko --apply # translate and write
python3 scripts/translate-docs.py --to ko --apply --judge # ... and score each file
python3 scripts/translate-docs.py --to ko --apply --only pricing.md
python3 scripts/translate-docs.py --to ko --apply --model gpt-5.6-sol --force
Dry-run is the default. It prints the file count, the character count and the estimated AP, and exits without calling anything. Every batch script you write against a paid API should open this way.
Output goes to i18n/<locale>/docusaurus-plugin-content-docs/current/, which is where Docusaurus expects it. Adapt the path and you have the same job for any tree of Markdown.
What carries over to any batch job
Translation is the example; the shape is general.
- Estimate before you spend. Dry-run by default,
--applyto commit. - Verify the output mechanically, and refuse to write what fails. Not doing the work beats doing it wrong.
- Test for false success, not just for errors. The output that passes every check and is still wrong is the one that reaches production.
- Make the work incremental. Hash the inputs; unchanged means free.
- Pick the model by measuring on your own data. The best model on all axes was 47× the price for a rounding error.
See also
- Chat (LLM) API — models, headers, limits
- Jev — typed scoring without parsing prose
- Pricing — how AP is charged and read off the response headers