Launch offerup to +40% credits on every packends inClaim →
jagent.
Get a free key

Measured

Decisions API vs Jev: thirteen decision models on the same items

OpenAI's Decisions API went into public beta on October 6, 2026, three weeks after Jev opened the category. By 2026-10-08 OpenRouter served thirteen decision models from ten makers. We sent all of them the same 633 items and measured what came back.

Quick answer · verified 2026-10-08

Is the Decisions API better than Jev?

Neither. On simple choices GPT-6 Luna Decisions scored 94.1% and Jev 94.0%, inside a field that ran from 90.4% to 96.5%. With 59 and 77 options both trailed Clef Flash (92.0%). Both were among the fastest. The differences that matter are how often a confident answer was wrong, how many options a model takes, and the cost per decision.

Simple choices90.4% to 96.5% — every model, news topics and support tickets
59–77 optionsClef Flash, 92.0% — GPT-6 Luna 79.8%, Jev 79.0%
Best calibratedd1 — confidence off by 5.0 points on average
FastestTev1 4B Experimental, 136 ms — median from US East; GPT-6 Luna 138 ms

What the Decisions API is

OpenAI's Decisions API is POST /v1/decisions: you send an input (text, or user messages with images) and an array of questions, and get typed answers back instead of text. There are three question types — predicate (the probability a statement is true), choice (one of up to 255 options, with a probability for each) and score (a level on an ordered rubric) — and a question the model declines comes back as a refusal. It runs one model, GPT-6 Luna, at $0.10 per million input tokens with output free.

Jev, launched by TypeSafe on September 15, 2026, does the same job on its System One contract: a state and named choice, score and noul questions (noul is the predicate). Eight other makers serve models on that contract, and OpenRouter carries them all, GPT-6 Luna Decisions included, behind one endpoint.

Every decision model on the Decisions API

ModelMaker$ / M inputImagesQuestionsReleased
Jev 1.13TypeSafe0.042noall three2026-09-18
GPT-6 Luna DecisionsOpenAI0.100yesall three2026-10-06
ClefCloudflare0.240yesall three2026-10-01
Clef FlashCloudflare0.090yesall three2026-10-01
Decider V1.1 27BPerplexity0.020yesall three2026-10-07
Solar DecideUpstage0.050noall; ≤26 options2026-09-28
Solar Decide FlashUpstage0.050noall; ≤26 options2026-10-08
d1Liquid AI0.040noall three2026-10-01
Mercury Decide (free)Inception0.000noall three2026-09-30
Tev1 4B ExperimentalTogether AI0.042noall; ≤20 options2026-09-30
Kev 4BJared Palmer0.042noall three2026-09-25
Span-01Respan0.020noyes/no only2026-09-26
Span-01 Lite (free)Respan0.000noyes/no only2026-09-26

Prices are OpenRouter's list prices; none of these models bills output. Clef, Clef Flash and Kev publish their weights. Respan's Span-01 is a different kind of model — it scores whether behaviours you describe occur in a conversation — so it appears only in the yes/no results below.

Accuracy on the same items

Six sets of 100 items come from BTZSC, the zero-shot classification benchmark, sampled with a fixed seed: news topics (4 options), tweet emotions (6), voice-assistant intents (59), bank support intents (77), toxic comments and product-review sentiment as yes/no questions. The seventh is our own 27 support tickets, written with one decisive signal each before any model ran. Every item was one question, sent to every model in the same words.

ModelNews + tickets59–77 optionsYes/noOur tickets
d196.5%86.0%96.5%100.0%
Decider V1.1 27B96.0%85.0%95.5%100.0%
Mercury Decide (free)95.6%85.0%94.0%96.3%
Clef95.0%89.0%97.0%100.0%
Kev 4B95.0%78.5%91.0%100.0%
Solar Decide94.6%refused93.0%96.3%
GPT-6 Luna Decisions94.1%79.8%93.5%96.3%
Solar Decide Flash94.1%refused91.5%96.3%
Jev 1.1394.0%79.0%95.0%100.0%
Clef Flash92.6%92.0%97.0%96.3%
Tev1 4B Experimental90.4%refused93.5%88.9%
Span-01 Lite (free)——93.5%—
Span-01——93.5%—

“News + tickets” averages the two choice sets every model accepted. Solar Decide and Tev1 refused the two large sets: Solar takes at most 26 options and Tev1 at most 20. With 100 items a set, models a few points apart are not reliably different: read tiers, not ranks. Tweet emotions are left out of the average because Mercury Decide scored 92.0% there against 56.0%–64.0% for everyone else, which suggests it saw that dataset in training.

When a confident answer is wrong

Confidence is what gates an action: act above a threshold, send the rest to a person. So the useful number is how often an answer at 90% confidence or more was wrong. It ranged from 2.3% for Clef to 13.7% for Solar Decide; Jev was at 11.5% and GPT-6 Luna at 9.2%.

ModelWrong at ≥90%Calibration errorConfidence: clear / ambiguous tickets
Clef2.3%0.0570.84 / 0.59
Kev 4B2.6%0.1230.66 / 0.39
d13.2%0.0500.95 / 0.52
Clef Flash4.2%0.0590.80 / 0.68
Decider V1.1 27B4.6%0.0690.99 / 0.81
Tev1 4B Experimental5.1%0.0620.89 / 0.86
Mercury Decide (free)5.4%0.0650.99 / 0.65
GPT-6 Luna Decisions9.2%0.0880.92 / 0.79
Solar Decide Flash9.5%0.1160.81 / 0.71
Jev 1.1311.5%0.1090.98 / 0.82
Solar Decide13.7%0.1290.89 / 0.96

Calibration error is the average gap between stated confidence and accuracy; lower is better. The last column compares our 27 clear tickets with six written to pull two ways: confidence that drops on those is usable.

Speed and cost per decision

ModelMedian90th pctTokens / item$ per 1,000 decisions
Span-01128 ms159 ms118$0.0029
Tev1 4B Experimental136 ms186 ms164$0.0072
GPT-6 Luna Decisions138 ms187 ms218$0.0401
Jev 1.13184 ms287 ms393$0.0304
Decider V1.1 27B269 ms390 ms158$0.0072
Clef313 ms695 ms222$0.1541
d1322 ms496 ms104$0.0078
Kev 4B359 ms789 ms92$0.0115
Clef Flash369 ms558 ms222$0.0578
Solar Decide559 ms1048 ms412$0.0216
Solar Decide Flash597 ms982 ms412$0.0216

Latency was measured from Washington, D.C. through OpenRouter, 20 calls each, Jev included. The two free-tier models are rate-limited upstream and were not timed. Cost is what the gateway billed. Price per token misleads: the same item cost 393 tokens on Jev and 218 on GPT-6 Luna, so at under half the list price Jev billed 76% of Luna's cost per decision. The cheapest paid model for full choice, score and yes/no questions was Decider V1.1 27B.

Three things that break quietly

Images on OpenRouter. OpenAI's format sends an image as an input_image part. OpenRouter's Decisions API does not read that shape — it answers anyway, from the base64 read as text. A plain red square came back “blue” from all three image models we tried; sent as a chat-style image_url part, all three said red.

Option caps. Solar Decide rejects a choice with more than 26 options and Tev1 with more than 20. Neither cap is in the catalogue. Respan answers yes/no questions only and wants text or a conversation, not a JSON object.

Call any of them with the OpenAI SDK

POST https://jev-agent.com/api/v1/decisions takes OpenAI's request and returns OpenAI's response, so the official SDK works with the base URL and key changed — and model can be any of the thirteen. Images are rewritten to the shape that is read, and a model that cannot take a request returns a 400 instead of guessing.

ts
import OpenAI from "openai"; // 7.30 or later

const client = new OpenAI({ baseURL: "https://jev-agent.com/api/v1", apiKey: process.env.JAGENT_KEY });

const d = await client.decisions.create({
  model: "jev-latest", // or "gpt-6-luna", "cloudflare/clef", "liquid/d1" …
  input: "I was charged twice for order 5512. Please refund the duplicate.",
  questions: [
    { type: "predicate", name: "refund", instructions: "Does the customer ask for money back?" },
    { type: "choice", name: "queue", instructions: "Which queue?", choices: [{ value: "billing" }, { value: "technical" }, { value: "sales" }] },
  ],
});

Billing is in credits: one per 1,000 input tokens for Jev and every model priced at or below it, scaled by price above that (GPT-6 Luna 2.4, Clef 5.8). The full list with limits is GET /api/v1/models, the same models answer Jev-shaped requests at /api/v1/systemone, and the playground has a model picker. Keys are on API access.

Which decision model to use

For yes/no gates the spread was small, so pick on latency and price. For dozens of intents, pick a model that takes them (Clef Flash led at 77 options). For thresholds that trigger actions, weigh the confident-and-wrong column above accuracy. Then run your own items: these are public datasets.

Questions

Is OpenAI's Decisions API the same as Jev?

The same idea, not the same model. Both answer typed questions with probabilities instead of writing text. OpenAI's Decisions API runs GPT-6 Luna at POST /v1/decisions; Jev is TypeSafe's model, where the yes/no type is called noul rather than predicate.

How much does the Decisions API cost?

OpenAI charges $0.10 per million input tokens and nothing for output. Jev lists $0.042. Per decision, our run billed $0.000040 for GPT-6 Luna Decisions and $0.000030 for Jev on the same items, because the models count tokens differently.

Which decision model is most accurate?

On simple choices they tie: every model scored between 90.4% and 96.5% on news topics and support tickets. With 59 and 77 options the field spreads out, and Clef Flash led at 92.0%.

Can I call other decision models with the OpenAI SDK?

Yes, through Jagent: point the SDK's baseURL at https://jev-agent.com/api/v1 and use a Jagent key. decisions.create then accepts gpt-6-luna, jev-latest or any model listed at /api/v1/models.