Skip to main content

Jev

:::warning What comes back is a probability, not a decision

Jev returns likelihoods and similarities. 0.92 means "strongly leaning that way", not "settled". The name choice returns is the option that scored highest, not a decision. Nothing here decides anything — turning a probability into a decision by picking a threshold is your job, as the caller, and so is what happens on either side of it.

That is why this endpoint is called /v1/jev and not /v1/decisions. A decision is a heavy, discrete, final thing that a person makes. A name heavier than the thing it names invites readers to treat 0.92 as a verdict, so we changed the name. POST /v1/decisions was removed on 2026-09-23 and now returns 404 — there is no alias.

:::

This endpoint is TypeSafe's Jev. Same model, same answers — "model": "jev-1.13.0" comes back in every response. No TypeSafe account, no Cloudflare account, no waitlist: the aicu_live_ key you already have is enough. Cloudflare's published examples run here with only the URL and the key changed.

Ask a question, get back a typed value your code can branch on — a choice, a score, or a likelihood — each with the probabilities behind it. No JSON to parse out of prose, no retries when a model wraps its answer in a code fence.

The problem this replaces​

The usual way to get a decision out of a language model is to ask for JSON and hope. We measured how well that actually holds up — 10 cases, three runs each (n=30), all through this same API, on 2026-09-22.

MethodParsed cleanlyValue was one we declaredp50p95
POST /v1/jev30 / 3030 / 30439 ms525 ms
gpt-5.4-nano, asked for JSON30 / 3030 / 30949 ms1,451 ms
deepseek-v3, asked for JSON20 / 3020 / 301,697 ms2,402 ms

A capable model returns clean JSON. An earlier run of this page reported gpt-5.4-nano at 5/6 and built an argument on it; at n=30 it is 30/30. We were reading noise. If your model is good and your prompt is careful, "it cannot return JSON" is not the reason to use this endpoint.

Three differences survive the larger sample, and they are the honest case for it.

  1. Speed, especially the tail. 525 ms at p95 against 1,451 ms — 2.8x. Median is 2.2x. The short tail is what lets you put a decision on a request path.
  2. The same input gives the same answer. Across the 10 cases, /v1/jev agreed with itself 3 times out of 3 every time. gpt-5.4-nano split 2/3 on one of them.
  3. You get probabilities and confidence. Not just the answer, but how strongly — which is what lets you route the uncertain cases to a human instead of guessing a threshold.

deepseek-v3 is a different story: it failed 10 times out of 30, with no code-fence rescue attempted. It returned "route":"tech|billing" repeatedly — copying the pipe-separated option list out of the prompt as if it were a value. Weaker models do fail this way, and a decision endpoint removes the category entirely: you declare the options, the answer is one of them.

The three question types​

TypeYou supplyYou get back
choicenamed optionsone of your option keys, plus a probability for each
scoreordered labels, low to higha continuous score across your scale, plus a probability per step
nouljust the questiona number from 0 to 1 — how strongly yes

noul is not a boolean. It is the strength of a yes, so you pick the threshold rather than the model picking it for you. A confirmation dialog might fire at 0.5; deleting something might want 0.9.

Your first call​

curl -X POST https://api.aicu.ai/v1/jev \
-H "Authorization: Bearer aicu_live_xxx" \
-H "Content-Type: application/json" \
-d '{
"state": "Ticket: \"Since yesterday /v1/chat/completions returns 401. The key was issued last week. I have not received an invoice either.\"",
"questions": {
"route": {
"type": "choice",
"instructions": "Pick the team that should handle this first",
"criteria": {
"tech": "Technical support",
"billing": "Billing",
"sales": "Sales",
"abuse": "Abuse investigation"
}
},
"urgency": {
"type": "score",
"instructions": "How urgent is this",
"criteria": ["next business day", "today", "within hours", "immediately"]
},
"needs_human": {
"type": "noul",
"instructions": "Should a human operator take over"
}
}
}'

The response, unedited:

{
"object": "jev",
"model": "jev-1.13.0",
"answers": {
"route": {
"type": "choice",
"choice": "tech",
"probabilities": { "tech": 0.89, "billing": 0.1, "abuse": 0.01, "sales": 0 },
"confidence": 0.85
},
"urgency": {
"type": "score",
"score": 1.57,
"legend": { "0": "next business day", "1": "today", "2": "within hours", "3": "immediately" },
"probabilities": { "0": 0.04, "1": 0.4, "2": 0.5, "3": 0.06 },
"confidence": 0.45
},
"needs_human": { "type": "noul", "noul": 0.8 }
},
"usage": { "input_tokens": 492, "output_tokens": 79 },
"ap_cost": 2
}

answers.route.choice is one of the keys you declared. It cannot be anything else — so the branch below never needs a default: that logs an alert and pages someone.

route = res["answers"]["route"]["choice"] # "tech" | "billing" | "sales" | "abuse"
urgency = res["answers"]["urgency"]["score"] # 0.0 .. 3.0
if res["answers"]["needs_human"]["noul"] > 0.7:
assign_to_human(route, urgency)

Read the confidence, not just the answer​

Every question comes back with its own confidence, and they differ inside a single response. In the example above route is 0.85 while urgency is 0.45 — the model is clear about who should take the ticket and genuinely unsure how fast. The probabilities say why: urgency is split 0.40 / 0.50 between "today" and "within hours".

That is the useful part. Treat low confidence as a routing signal of its own:

r = res["answers"]["route"]
if r["confidence"] < 0.6:
queue_for_review(ticket) # let a person look, rather than guess
else:
assign(r["choice"])

Prose answers do not give you this. A model that says "this is a billing issue" says it in exactly the same tone whether it is sure or not.

Ask several questions at once​

The three questions in the example were answered in one call, at one price. Group everything you need about one piece of state into a single request: it is cheaper, it is one round trip, and every answer is made against the same context.

Scoring: put your scale in the labels​

criteria for a score is an ordered array, lowest first. The labels do real work — they are how the model knows what the scale means, so write them the way you would write them for a new colleague.

{
"risk": {
"type": "score",
"instructions": "How far does this submission fall outside the rules",
"criteria": ["fine", "minor", "serious", "clear violation"]
}
}

Given a contest entry whose author wrote that they had trained on a real person's photos without permission, this returned 2.94 with confidence 0.94 and probabilities {"2": 0.05, "3": 0.95} — pinned to the top of the scale rather than hedged in the middle.

Limits​

LimitValue
Questions per request20
state20,000 characters
instructions, per question2,000 characters
criteria items, per question20
Each criteria item500 characters
questions total, as JSON64,000 bytes
Upstream timeout20 seconds

state is required. model is optional and accepts jev-latest (the default) or jev-1.13.0; anything else is rejected with the accepted list in allowed_models.

Pricing​

2 AP per request, plus 1 AP for every 1,000 characters of input, rounded down — so a typical question costs 2 AP. The number of questions in a request does not change the price; the size of what you send does.

Input here means the state plus the questions, since both are sent upstream. A request at the limits costs about 27 AP.

The charge is reserved before the upstream call and refunded in full if the call fails — a failed decision costs nothing. Every response carries X-AICU-AP-Cost, X-AICU-AP-Balance and X-Credits-Remaining.

Authentication​

Bearer key with the llm scope. If your key already calls /v1/chat/completions, it already works here — no new scope, no reissue.

Errors​

StatuscodeWhen
400invalid_parametermissing state, a malformed question, an unknown model, or a limit exceeded
400invalid_jsonthe body is not JSON
402insufficient_creditsnot enough AP
403—the key lacks the llm scope
502generation_failedthe provider failed or returned no answers — refunded
503internal_errorthe decision provider is unavailable — refunded

When not to use it​

This returns decisions, not text. If you need a sentence for a human to read, call the LLM API — and if you need both, it is normal to do both: decide with this, then write the reply with a model, only for the tickets that turned out to need one.

Under the hood​

Requests are served by TypeSafe's Jev through Cloudflare Workers AI. You do not need a TypeSafe account, a Cloudflare account, or a key for either — the same aicu_live_ key you already have is enough.

Jev is a System One model: instead of generating a string you then parse, it evaluates one state against typed questions and returns the answers in parallel, with calibrated probabilities. That is why a request with three questions costs the same and takes about the same time as a request with one.

state accepts a string, an object or an array — Cloudflare's published examples, including the ones that pass nested objects, run here unchanged. Numbers, booleans and null are rejected with 400 invalid_parameter, because upstream Jev does not accept them either.