Skip to main content

PDF & Document Processing Guide

Turn typeset PDFs into structured Markdown. Useful for proofreading, reusing existing material, and making illustrated manuals searchable. There is no dedicated endpoint — you post to the same POST /v1/chat/completions used by Vision (image input), with a key that carries the llm scope.

:::info About passing PDFs directly input_file from OpenAI's Responses API (which takes a PDF as-is) is not supported yet on api.aicu.ai — we're looking into it. For now, follow the steps on this page: convert the PDF to page images first, then pass them as image_url. :::

The idea: don't make the model transcribe​

If all you need is the text of a PDF, pdftotext already does the job — free, fast, and character-accurate. Only two things are missing:

  • Reading order — with multiple columns it jumps between them. The characters are correct, but the result doesn't read as prose
  • What's in the figures — you get exactly zero characters out of screenshots and sample images

So pass both the text layer and the page image, and leave the model just two jobs: restore the reading order, and describe the figures.

input = page image (PNG) + pdftotext text layer
output = structured Markdown (figures described in place as [Figure])

Because the body text comes from the text layer, accuracy doesn't degrade and the model emits fewer output tokens. Cheaper and more accurate than handing over page images alone and asking for a full transcription.

Setup​

brew install ghostscript poppler # gs (page rendering) + pdftotext/pdfinfo
pip install openai

Minimal example (one page)​

import base64, os, subprocess, tempfile
from pathlib import Path
from openai import OpenAI

client = OpenAI(
api_key=os.environ["AICU_API_KEY"],
base_url="https://api.aicu.ai/v1",
)

SYSTEM = """You are an expert at reading typeset pages and producing structured Markdown.
Your input is a page image plus text mechanically extracted with pdftotext.
The text is character-accurate, but its reading order is broken because the page is multi-column.

1. Take the body text from the supplied text and put it back in the correct reading order. Never rewrite the characters.
2. Mark headings with ## / ### according to their level.
3. Where a figure or screenshot appears, insert the following at that position:
> **[Figure]** What it shows (1-2 sentences). Include any important text or UI elements visible in it.
4. Turn numbered callouts in procedures into an ordered list in the body text.
5. Render boxed notes and asides as `> **[Note]**`.
6. Do not output typesetting artifacts such as page numbers, running heads, or crop marks.
7. No preamble — return the Markdown itself and nothing else."""


def page_to_markdown(pdf: str, page: int, model: str = "gemini-3.6-flash") -> str:
# 1) Render the page to PNG (around 120 dpi is plenty)
with tempfile.TemporaryDirectory() as tmp:
png = Path(tmp) / "page.png"
subprocess.run([
"gs", "-sDEVICE=png16m", "-r120",
f"-dFirstPage={page}", f"-dLastPage={page}",
"-dNOPAUSE", "-dQUIET", "-dBATCH", f"-sOutputFile={png}", pdf,
], check=True)
b64 = base64.b64encode(png.read_bytes()).decode()

# 2) Text layer for the same page
text = subprocess.run(
["pdftotext", "-layout", "-f", str(page), "-l", str(page), pdf, "-"],
check=True, capture_output=True,
).stdout.decode("utf-8", "replace").strip()

# 3) Send both
res = client.chat.completions.create(
model=model,
messages=[
{"role": "system", "content": SYSTEM},
{"role": "user", "content": [
{"type": "text",
"text": f"Text layer:\n```\n{text}\n```\n\n"
f"Look at the page image and return Markdown following the rules above."},
{"type": "image_url", "image_url": {
"url": f"data:image/png;base64,{b64}", "detail": "high",
}},
]},
],
)
return res.choices[0].message.content


print(page_to_markdown("manual.pdf", 1))

Processing every page (parallel + cache + cost tally)​

For long PDFs, cache per page so that a failure halfway through is cheap to retry. Read the X-AICU-AP-Cost header and you can tally the real cost as you go.

import base64, concurrent.futures as futures, os, re, subprocess, tempfile, sys
from pathlib import Path
from openai import OpenAI

MODEL = "gemini-3.6-flash"
SYSTEM = "..." # same system prompt as "Minimal example" above
client = OpenAI(api_key=os.environ["AICU_API_KEY"], base_url="https://api.aicu.ai/v1")


def page_count(pdf: str) -> int:
out = subprocess.run(["pdfinfo", pdf], check=True, capture_output=True)
m = re.search(r"^Pages:\s+(\d+)", out.stdout.decode(), re.M)
return int(m.group(1)) if m else 0


def one_page(pdf: str, page: int, cache: Path):
hit = cache / f"p{page:03d}.md"
if hit.exists(): # don't hit the API on a re-run
return page, hit.read_text(encoding="utf-8"), 0

with tempfile.TemporaryDirectory() as tmp:
png = Path(tmp) / "p.png"
subprocess.run(["gs", "-sDEVICE=png16m", "-r120",
f"-dFirstPage={page}", f"-dLastPage={page}",
"-dNOPAUSE", "-dQUIET", "-dBATCH",
f"-sOutputFile={png}", pdf], check=True)
b64 = base64.b64encode(png.read_bytes()).decode()

text = subprocess.run(["pdftotext", "-layout", "-f", str(page), "-l", str(page), pdf, "-"],
check=True, capture_output=True).stdout.decode("utf-8", "replace")

# go through raw_response so we can read the actual AP charged
raw = client.chat.completions.with_raw_response.create(
model=MODEL,
messages=[
{"role": "system", "content": SYSTEM},
{"role": "user", "content": [
{"type": "text", "text": f"Text layer:\n```\n{text}\n```"},
{"type": "image_url", "image_url": {
"url": f"data:image/png;base64,{b64}", "detail": "high"}},
]},
],
)
ap = int(raw.headers.get("x-aicu-ap-cost", 0) or 0)
md = raw.parse().choices[0].message.content
hit.write_text(md, encoding="utf-8")
return page, md, ap


def pdf_to_markdown(pdf: str, workers: int = 4) -> str:
total = page_count(pdf)
cache = Path(".cache") / Path(pdf).stem
cache.mkdir(parents=True, exist_ok=True)

results, ap_total = {}, 0
with futures.ThreadPoolExecutor(max_workers=workers) as pool:
jobs = {pool.submit(one_page, pdf, p, cache): p for p in range(1, total + 1)}
for f in futures.as_completed(jobs):
try:
page, md, ap = f.result()
except Exception as e:
print(f" p{jobs[f]} failed: {e}", file=sys.stderr)
continue
results[page] = md
ap_total += ap
print(f" p{page} done")

# 10,000 AP = $1 (approximate; the AP rate is subject to revision)
print(f"total {ap_total:,} AP ~= ${ap_total / 10000:,.4f}")
return "\n\n".join(results[p] for p in sorted(results))


Path("out.md").write_text(pdf_to_markdown("manual.pdf"), encoding="utf-8")

Choosing a model​

Measured on one spread page from a technical book (1535x1114 px, two columns, 8 callouts, 5 screenshots, detail: high).

ModelInput tokOutput tokSecAP/pageNotes
gemini-3.6-flash1,7801,2257.6301Recommended. Reads text inside screenshots too
gpt-5.6-sol2,7846848.91,017Just as faithful. Expensive
gpt-51,7106427.1203Tends to invent headings that aren't on the page
gpt-5-mini2,78475311.669Turns callout numbers into headings
gpt-4o1,9063346.5222Silently corrects typos in the original
gpt-4o-mini37,6364156.3191About 20x the image tokens

(Token counts and AP were measured in separate runs. In production, with a tuned prompt, gemini-3.6-flash came out at roughly 761 AP ~= $0.076 per page.)

:::caution Prices are approximate and AP rates get revised The AP figures in the table and the dollar amounts on this page are estimates as of the time of measurement. AP is denominated in USD at 10,000 AP = $1 (1 AP = $0.0001), but the AP rate is subject to price revisions. When estimating, always check the current values from GET /v1/chat/models and the LLM API pricing. :::

Caveat 1: gpt-4o-mini is not cheap for images​

For the same image it burns about 20x the tokens of the other models. That's by design — a high token multiplier offset by a low unit price, so what you pay ends up about the same. But those tokens are really consumed from the context window, so a design that batches several pages into one request will overflow.

Caveat 2: For proofreading, verify that the original is preserved​

The test page contained Colaub (a typo for Colab). gpt-4o "fixed" it to Colab in its output.

Helpful behavior, but fatal if proofreading is the goal: the process meant to find typos erases them. It happened even with an explicit "never rewrite the characters" in the prompt, so if your use case is proofreading, test on a page with known typos before picking a model.

For the record, gemini-3.6-flash, gpt-5.6-sol, gpt-5 and gpt-5-mini all preserved the typo.

Tips​

  • 120 dpi or so is enough. Going higher just adds tiles and cost; accuracy barely moves
  • Use detail: high. If you want the model to read what's inside the figures, low won't cut it
  • Cache per page. A long PDF will fail partway through
  • Don't split spreads. Given the whole spread, the model can follow body text that runs across both pages
  • If you don't need the figures and only want the body text, pdftotext -layout alone is enough. No API required

Troubleshooting​

SymptomCause and fix
403 ForbiddenThe key lacks the llm scope -> ask whoever issued it to add it
404 Not Found on /v1/responsesThe Responses API isn't supported. Use /v1/chat/completions
Cloudflare error 1010Your HTTP client's default User-Agent was flagged as a bot. Set an explicit UA
Output wrapped in ```markdownState in the prompt: no preamble, return only the Markdown itself
Thin figure descriptionsCheck that detail: "high" is set, and that the dpi isn't too low

See also​

© 2026 AICU Inc.