From direct provider calls to api.aicu.ai
A field report from migrating two things we actually run in-house: our Slack bot (Aicuty Bot) and our article-generation tool.
Why bother rewriting
Before the migration, our image-generation code looked like this:
- The Slack bot held an
OPENAI_API_KEYand calledapi.openai.comdirectly - The article tool held a
GEMINI_API_KEYand calledgenerativelanguage.googleapis.comdirectly - When either one failed, the error only ever landed in the runtime's stdout or a Slack message
The practical pain here wasn't the number of keys. It was that nobody could tell what was being used, or how much. When we tallied the production logs before migrating, 1,325 LLM requests had gone through with not a single entry in the ledger. No basis for billing anyone.
The principle: one place to call out from
api.aicu.ai is OpenAI-compatible, so swapping the endpoint and the key is all it takes.
| Before | After | |
|---|---|---|
| Endpoint | One per provider | https://api.aicu.ai/v1 |
| Keys | OPENAI_API_KEY + GEMINI_API_KEY + … | One AICU_API_KEY |
| Model selection | Whatever each SDK wants | The model argument |
| Usage records | Roll your own | Automatic, in the dashboard |
Step 1: Slip a thin client in front
Rewriting everything at once was not the safe move. Writing one thin module and dropping it in was. It's toggled by an environment variable, so if anything goes wrong you can fall straight back to the original path.
# aicu_images.py
import base64, os, requests
def is_enabled() -> bool:
return bool(os.environ.get("AICU_API_KEY"))
def generate(prompt, *, model="gpt-image-2", character=None,
reference_image=None, mask=None, size=None, quality=None,
aspect_ratio=None, timeout=240):
payload = {"model": model, "prompt": prompt}
if character:
payload["character"] = character
if reference_image:
payload["reference_image"] = base64.b64encode(reference_image).decode()
if mask:
payload["mask"] = base64.b64encode(mask).decode()
for k, v in (("size", size), ("quality", quality), ("aspect_ratio", aspect_ratio)):
if v:
payload[k] = v
res = requests.post(
"https://api.aicu.ai/v1/images/generations",
headers={"Authorization": f"Bearer {os.environ['AICU_API_KEY']}"},
json=payload, timeout=timeout,
)
res.raise_for_status()
data = res.json()
return base64.b64decode(data["data"][0]["b64_json"]), bool(data.get("cached"))
Step 2: Swap the call site (with a fallback)
+ if aicu_images.is_enabled():
+ try:
+ image_data, cached = aicu_images.generate(
+ prompt, model="gpt-image-2", character=member,
+ reference_image=ref_bytes, mask=mask_bytes,
+ size="1536x1024", quality="high",
+ )
+ except aicu_images.AicuImageError as exc:
+ print(f"[gpt-image2] AICU gateway failed: {exc}")
+
+ if image_data is None:
oai = OpenAI(api_key=OPENAI_API_KEY)
response = oai.images.edit(model="gpt-image-2", image=..., prompt=prompt)
image_data = base64.b64decode(response.data[0].b64_json)
Until you set AICU_API_KEY, behavior doesn't change by a millimeter. Only once it's set do calls route through the gateway. Shipping that "installed but inert" state to production first takes the fear out of the cutover.
Step 3: Migrating off Gemini
Gemini's SDK has a rather different shape, but the replacement actually comes out shorter.
- client = genai.Client(api_key=GEMINI_API_KEY)
- response = client.models.generate_content(
- model="gemini-3.1-flash-image-preview",
- contents=prompt,
- config=types.GenerateContentConfig(
- response_modalities=["IMAGE", "TEXT"],
- image_config=types.ImageConfig(aspect_ratio="16:9"),
- ),
- )
- for part in response.candidates[0].content.parts:
- if part.inline_data:
- output_path.write_bytes(part.inline_data.data)
+ data, cached = aicu_images.generate(
+ prompt, model="nano-banana", aspect_ratio="16:9"
+ )
+ output_path.write_bytes(data)
You can pass the original model name straight through and it still works (gemini-3.1-flash-image-preview resolves as an alias to nano-banana). Get it running first, shorten the name later.
Where we tripped
How to pass reference images. The first design had you upload to R2 ahead of time — but the tools being migrated already had the reference image in hand. So we made reference_image accept base64 directly. No pre-registration, and the migration collapsed down to swapping a single function.
Don't drop mask editing. The Slack bot could attach a mask image to edit part of a picture. Because the gateway passes mask through, we migrated without cutting the feature. When a migration takes features away, the people using it go back.
Key scopes. The keys we'd already issued only carried ["tts"], so image calls came back 403. Alpha keys now come with llm / tts / images from the start.
What we got out of it
- Failures are visible. The dashboard shows which model failed and how many times. Things we'd have missed before — like "707 failures because OpenRouter was out of credit" — surface immediately.
- The cache pays off. Regenerating under the same conditions returns from R2 and isn't billed (
cached: true). - We know our costs. Per-model margin and break-even are computable, so pricing arguments can run on numbers.
- Character royalties can be apportioned. → Character API
Support
- Dashboard: https://api.aicu.ai/dashboard
- Usage history: https://api.aicu.ai/dashboard/history
© 2026 AICU Inc.