Jev
:::warning What comes back is a probability, not a decision
Jev returns likelihoods and similarities. 0.92 means "strongly leaning that way", not "settled". The name choice returns is the option that scored highest, not a decision. Nothing here decides anything — turning a probability into a decision by picking a threshold is your job, as the caller, and so is what happens on either side of it.
That is why this endpoint is called /v1/jev and not /v1/decisions. A decision is a heavy, discrete, final thing that a person makes. A name heavier than the thing it names invites readers to treat 0.92 as a verdict, so we changed the name. POST /v1/decisions was removed on 2026-09-23 and now returns 404 — there is no alias.
:::
This endpoint is TypeSafe's Jev. Same model, same answers —
"model": "jev-1.13.0"comes back in every response. No TypeSafe account, no Cloudflare account, no waitlist: theaicu_live_key you already have is enough. Cloudflare's published examples run here with only the URL and the key changed.
Ask a question, get back a typed value your code can branch on — a choice, a score, or a likelihood — each with the probabilities behind it. No JSON to parse out of prose, no retries when a model wraps its answer in a code fence.
The problem this replaces
The usual way to get a decision out of a language model is to ask for JSON and hope. We measured how well that actually holds up — 10 cases, three runs each (n=30), all through this same API, on 2026-09-22.
| Method | Parsed cleanly | Value was one we declared | p50 | p95 |
|---|---|---|---|---|
POST /v1/jev | 30 / 30 | 30 / 30 | 439 ms | 525 ms |
gpt-5.4-nano, asked for JSON | 30 / 30 | 30 / 30 | 949 ms | 1,451 ms |
deepseek-v3, asked for JSON | 20 / 30 | 20 / 30 | 1,697 ms | 2,402 ms |
A capable model returns clean JSON. An earlier run of this page reported gpt-5.4-nano at 5/6 and
built an argument on it; at n=30 it is 30/30. We were reading noise. If your model is good and your
prompt is careful, "it cannot return JSON" is not the reason to use this endpoint.
Three differences survive the larger sample, and they are the honest case for it.
- Speed, especially the tail. 525 ms at p95 against 1,451 ms — 2.8x. Median is 2.2x. The short tail is what lets you put a decision on a request path.
- The same input gives the same answer. Across the 10 cases,
/v1/jevagreed with itself 3 times out of 3 every time.gpt-5.4-nanosplit 2/3 on one of them. - You get probabilities and confidence. Not just the answer, but how strongly — which is what lets you route the uncertain cases to a human instead of guessing a threshold.
deepseek-v3 is a different story: it failed 10 times out of 30, with no code-fence rescue attempted. It returned
"route":"tech|billing" repeatedly — copying the pipe-separated option list out of the prompt as if it
were a value. Weaker models do fail this way, and a decision endpoint removes the category entirely:
you declare the options, the answer is one of them.
The three question types
| Type | You supply | You get back |
|---|---|---|
choice | named options | one of your option keys, plus a probability for each |
score | ordered labels, low to high | a continuous score across your scale, plus a probability per step |
noul | just the question | a number from 0 to 1 — how strongly yes |
noul is not a boolean. It is the strength of a yes, so you pick the threshold rather than the model picking it for you. A confirmation dialog might fire at 0.5; deleting something might want 0.9.
Your first call
curl -X POST https://api.aicu.ai/v1/jev \
-H "Authorization: Bearer aicu_live_xxx" \
-H "Content-Type: application/json" \
-d '{
"state": "Ticket: \"Since yesterday /v1/chat/completions returns 401. The key was issued last week. I have not received an invoice either.\"",
"questions": {
"route": {
"type": "choice",
"instructions": "Pick the team that should handle this first",
"criteria": {
"tech": "Technical support",
"billing": "Billing",
"sales": "Sales",
"abuse": "Abuse investigation"
}
},
"urgency": {
"type": "score",
"instructions": "How urgent is this",
"criteria": ["next business day", "today", "within hours", "immediately"]
},
"needs_human": {
"type": "noul",
"instructions": "Should a human operator take over"
}
}
}'
The response, unedited:
{
"object": "jev",
"model": "jev-1.13.0",
"answers": {
"route": {
"type": "choice",
"choice": "tech",
"probabilities": { "tech": 0.89, "billing": 0.1, "abuse": 0.01, "sales": 0 },
"confidence": 0.85
},
"urgency": {
"type": "score",
"score": 1.57,
"legend": { "0": "next business day", "1": "today", "2": "within hours", "3": "immediately" },
"probabilities": { "0": 0.04, "1": 0.4, "2": 0.5, "3": 0.06 },
"confidence": 0.45
},
"needs_human": { "type": "noul", "noul": 0.8 }
},
"usage": { "input_tokens": 492, "output_tokens": 79 },
"ap_cost": 2
}
answers.route.choice is one of the keys you declared. It cannot be anything else — so the branch below never needs a default: that logs an alert and pages someone.
route = res["answers"]["route"]["choice"] # "tech" | "billing" | "sales" | "abuse"
urgency = res["answers"]["urgency"]["score"] # 0.0 .. 3.0
if res["answers"]["needs_human"]["noul"] > 0.7:
assign_to_human(route, urgency)
Read the confidence, not just the answer
Every question comes back with its own confidence, and they differ inside a single response. In the example above route is 0.85 while urgency is 0.45 — the model is clear about who should take the ticket and genuinely unsure how fast. The probabilities say why: urgency is split 0.40 / 0.50 between "today" and "within hours".
That is the useful part. Treat low confidence as a routing signal of its own:
r = res["answers"]["route"]
if r["confidence"] < 0.6:
queue_for_review(ticket) # let a person look, rather than guess
else:
assign(r["choice"])
Prose answers do not give you this. A model that says "this is a billing issue" says it in exactly the same tone whether it is sure or not.
Ask several questions at once
The three questions in the example were answered in one call, at one price. Group everything you need about one piece of state into a single request: it is cheaper, it is one round trip, and every answer is made against the same context.
Scoring: put your scale in the labels
criteria for a score is an ordered array, lowest first. The labels do real work — they are how the model knows what the scale means, so write them the way you would write them for a new colleague.
{
"risk": {
"type": "score",
"instructions": "How far does this submission fall outside the rules",
"criteria": ["fine", "minor", "serious", "clear violation"]
}
}
Given a contest entry whose author wrote that they had trained on a real person's photos without permission, this returned 2.94 with confidence 0.94 and probabilities {"2": 0.05, "3": 0.95} — pinned to the top of the scale rather than hedged in the middle.
Limits
| Limit | Value |
|---|---|
| Questions per request | 20 |
state | 20,000 characters |
instructions, per question | 2,000 characters |
criteria items, per question | 20 |
Each criteria item | 500 characters |
questions total, as JSON | 64,000 bytes |
| Upstream timeout | 20 seconds |
state is required. model is optional and accepts jev-latest (the default) or jev-1.13.0; anything else is rejected with the accepted list in allowed_models.
Pricing
2 AP per request, plus 1 AP for every 1,000 characters of input, rounded down — so a typical question costs 2 AP. The number of questions in a request does not change the price; the size of what you send does.
Input here means the state plus the questions, since both are sent upstream. A request at the limits costs about 27 AP.
The charge is reserved before the upstream call and refunded in full if the call fails — a failed decision costs nothing. Every response carries X-AICU-AP-Cost, X-AICU-AP-Balance and X-Credits-Remaining.
Authentication
Bearer key with the llm scope. If your key already calls /v1/chat/completions, it already works here — no new scope, no reissue.
Errors
| Status | code | When |
|---|---|---|
| 400 | invalid_parameter | missing state, a malformed question, an unknown model, or a limit exceeded |
| 400 | invalid_json | the body is not JSON |
| 402 | insufficient_credits | not enough AP |
| 403 | — | the key lacks the llm scope |
| 502 | generation_failed | the provider failed or returned no answers — refunded |
| 503 | internal_error | the decision provider is unavailable — refunded |
When not to use it
This returns decisions, not text. If you need a sentence for a human to read, call the LLM API — and if you need both, it is normal to do both: decide with this, then write the reply with a model, only for the tickets that turned out to need one.
Under the hood
Requests are served by TypeSafe's Jev through Cloudflare Workers AI. You do not need a TypeSafe account, a Cloudflare account, or a key for either — the same aicu_live_ key you already have is enough.
Jev is a System One model: instead of generating a string you then parse, it evaluates one state against typed questions and returns the answers in parallel, with calibrated probabilities. That is why a request with three questions costs the same and takes about the same time as a request with one.
state accepts a string, an object or an array — Cloudflare's published examples, including the ones that pass nested objects, run here unchanged. Numbers, booleans and null are rejected with 400 invalid_parameter, because upstream Jev does not accept them either.