Measured
Decisions API vs Jev: thirteen decision models on the same items
OpenAI's Decisions API went into public beta on October 6, 2026, three weeks after Jev opened the category. By 2026-10-08 OpenRouter served thirteen decision models from ten makers. We sent all of them the same 633 items and measured what came back.
Quick answer · verified 2026-10-08
Is the Decisions API better than Jev?
Neither. On simple choices GPT-6 Luna Decisions scored 94.1% and Jev 94.0%, inside a field that ran from 90.4% to 96.5%. With 59 and 77 options both trailed Clef Flash (92.0%). Both were among the fastest. The differences that matter are how often a confident answer was wrong, how many options a model takes, and the cost per decision.
| Simple choices | 90.4% to 96.5% — every model, news topics and support tickets |
|---|---|
| 59–77 options | Clef Flash, 92.0% — GPT-6 Luna 79.8%, Jev 79.0% |
| Best calibrated | d1 — confidence off by 5.0 points on average |
| Fastest | Tev1 4B Experimental, 136 ms — median from US East; GPT-6 Luna 138 ms |
What the Decisions API is
OpenAI's Decisions API is POST /v1/decisions: you send an input (text, or user messages with images) and an array of questions, and get typed answers back instead of text. There are three question types — predicate (the probability a statement is true), choice (one of up to 255 options, with a probability for each) and score (a level on an ordered rubric) — and a question the model declines comes back as a refusal. It runs one model, GPT-6 Luna, at $0.10 per million input tokens with output free.
Jev, launched by TypeSafe on September 15, 2026, does the same job on its System One contract: a state and named choice, score and noul questions (noul is the predicate). Eight other makers serve models on that contract, and OpenRouter carries them all, GPT-6 Luna Decisions included, behind one endpoint.
Every decision model on the Decisions API
| Model | Maker | $ / M input | Images | Questions | Released |
|---|---|---|---|---|---|
| Jev 1.13 | TypeSafe | 0.042 | no | all three | 2026-09-18 |
| GPT-6 Luna Decisions | OpenAI | 0.100 | yes | all three | 2026-10-06 |
| Clef | Cloudflare | 0.240 | yes | all three | 2026-10-01 |
| Clef Flash | Cloudflare | 0.090 | yes | all three | 2026-10-01 |
| Decider V1.1 27B | Perplexity | 0.020 | yes | all three | 2026-10-07 |
| Solar Decide | Upstage | 0.050 | no | all; ≤26 options | 2026-09-28 |
| Solar Decide Flash | Upstage | 0.050 | no | all; ≤26 options | 2026-10-08 |
| d1 | Liquid AI | 0.040 | no | all three | 2026-10-01 |
| Mercury Decide (free) | Inception | 0.000 | no | all three | 2026-09-30 |
| Tev1 4B Experimental | Together AI | 0.042 | no | all; ≤20 options | 2026-09-30 |
| Kev 4B | Jared Palmer | 0.042 | no | all three | 2026-09-25 |
| Span-01 | Respan | 0.020 | no | yes/no only | 2026-09-26 |
| Span-01 Lite (free) | Respan | 0.000 | no | yes/no only | 2026-09-26 |
Prices are OpenRouter's list prices; none of these models bills output. Clef, Clef Flash and Kev publish their weights. Respan's Span-01 is a different kind of model — it scores whether behaviours you describe occur in a conversation — so it appears only in the yes/no results below.
Accuracy on the same items
Six sets of 100 items come from BTZSC, the zero-shot classification benchmark, sampled with a fixed seed: news topics (4 options), tweet emotions (6), voice-assistant intents (59), bank support intents (77), toxic comments and product-review sentiment as yes/no questions. The seventh is our own 27 support tickets, written with one decisive signal each before any model ran. Every item was one question, sent to every model in the same words.
| Model | News + tickets | 59–77 options | Yes/no | Our tickets |
|---|---|---|---|---|
| d1 | 96.5% | 86.0% | 96.5% | 100.0% |
| Decider V1.1 27B | 96.0% | 85.0% | 95.5% | 100.0% |
| Mercury Decide (free) | 95.6% | 85.0% | 94.0% | 96.3% |
| Clef | 95.0% | 89.0% | 97.0% | 100.0% |
| Kev 4B | 95.0% | 78.5% | 91.0% | 100.0% |
| Solar Decide | 94.6% | refused | 93.0% | 96.3% |
| GPT-6 Luna Decisions | 94.1% | 79.8% | 93.5% | 96.3% |
| Solar Decide Flash | 94.1% | refused | 91.5% | 96.3% |
| Jev 1.13 | 94.0% | 79.0% | 95.0% | 100.0% |
| Clef Flash | 92.6% | 92.0% | 97.0% | 96.3% |
| Tev1 4B Experimental | 90.4% | refused | 93.5% | 88.9% |
| Span-01 Lite (free) | — | — | 93.5% | — |
| Span-01 | — | — | 93.5% | — |
“News + tickets” averages the two choice sets every model accepted. Solar Decide and Tev1 refused the two large sets: Solar takes at most 26 options and Tev1 at most 20. With 100 items a set, models a few points apart are not reliably different: read tiers, not ranks. Tweet emotions are left out of the average because Mercury Decide scored 92.0% there against 56.0%–64.0% for everyone else, which suggests it saw that dataset in training.
When a confident answer is wrong
Confidence is what gates an action: act above a threshold, send the rest to a person. So the useful number is how often an answer at 90% confidence or more was wrong. It ranged from 2.3% for Clef to 13.7% for Solar Decide; Jev was at 11.5% and GPT-6 Luna at 9.2%.
| Model | Wrong at ≥90% | Calibration error | Confidence: clear / ambiguous tickets |
|---|---|---|---|
| Clef | 2.3% | 0.057 | 0.84 / 0.59 |
| Kev 4B | 2.6% | 0.123 | 0.66 / 0.39 |
| d1 | 3.2% | 0.050 | 0.95 / 0.52 |
| Clef Flash | 4.2% | 0.059 | 0.80 / 0.68 |
| Decider V1.1 27B | 4.6% | 0.069 | 0.99 / 0.81 |
| Tev1 4B Experimental | 5.1% | 0.062 | 0.89 / 0.86 |
| Mercury Decide (free) | 5.4% | 0.065 | 0.99 / 0.65 |
| GPT-6 Luna Decisions | 9.2% | 0.088 | 0.92 / 0.79 |
| Solar Decide Flash | 9.5% | 0.116 | 0.81 / 0.71 |
| Jev 1.13 | 11.5% | 0.109 | 0.98 / 0.82 |
| Solar Decide | 13.7% | 0.129 | 0.89 / 0.96 |
Calibration error is the average gap between stated confidence and accuracy; lower is better. The last column compares our 27 clear tickets with six written to pull two ways: confidence that drops on those is usable.
Speed and cost per decision
| Model | Median | 90th pct | Tokens / item | $ per 1,000 decisions |
|---|---|---|---|---|
| Span-01 | 128 ms | 159 ms | 118 | $0.0029 |
| Tev1 4B Experimental | 136 ms | 186 ms | 164 | $0.0072 |
| GPT-6 Luna Decisions | 138 ms | 187 ms | 218 | $0.0401 |
| Jev 1.13 | 184 ms | 287 ms | 393 | $0.0304 |
| Decider V1.1 27B | 269 ms | 390 ms | 158 | $0.0072 |
| Clef | 313 ms | 695 ms | 222 | $0.1541 |
| d1 | 322 ms | 496 ms | 104 | $0.0078 |
| Kev 4B | 359 ms | 789 ms | 92 | $0.0115 |
| Clef Flash | 369 ms | 558 ms | 222 | $0.0578 |
| Solar Decide | 559 ms | 1048 ms | 412 | $0.0216 |
| Solar Decide Flash | 597 ms | 982 ms | 412 | $0.0216 |
Latency was measured from Washington, D.C. through OpenRouter, 20 calls each, Jev included. The two free-tier models are rate-limited upstream and were not timed. Cost is what the gateway billed. Price per token misleads: the same item cost 393 tokens on Jev and 218 on GPT-6 Luna, so at under half the list price Jev billed 76% of Luna's cost per decision. The cheapest paid model for full choice, score and yes/no questions was Decider V1.1 27B.
Three things that break quietly
Images on OpenRouter. OpenAI's format sends an image as an input_image part. OpenRouter's Decisions API does not read that shape — it answers anyway, from the base64 read as text. A plain red square came back “blue” from all three image models we tried; sent as a chat-style image_url part, all three said red.
Option caps. Solar Decide rejects a choice with more than 26 options and Tev1 with more than 20. Neither cap is in the catalogue. Respan answers yes/no questions only and wants text or a conversation, not a JSON object.
Call any of them with the OpenAI SDK
POST https://jev-agent.com/api/v1/decisions takes OpenAI's request and returns OpenAI's response, so the official SDK works with the base URL and key changed — and model can be any of the thirteen. Images are rewritten to the shape that is read, and a model that cannot take a request returns a 400 instead of guessing.
import OpenAI from "openai"; // 7.30 or later
const client = new OpenAI({ baseURL: "https://jev-agent.com/api/v1", apiKey: process.env.JAGENT_KEY });
const d = await client.decisions.create({
model: "jev-latest", // or "gpt-6-luna", "cloudflare/clef", "liquid/d1" …
input: "I was charged twice for order 5512. Please refund the duplicate.",
questions: [
{ type: "predicate", name: "refund", instructions: "Does the customer ask for money back?" },
{ type: "choice", name: "queue", instructions: "Which queue?", choices: [{ value: "billing" }, { value: "technical" }, { value: "sales" }] },
],
});Billing is in credits: one per 1,000 input tokens for Jev and every model priced at or below it, scaled by price above that (GPT-6 Luna 2.4, Clef 5.8). The full list with limits is GET /api/v1/models, the same models answer Jev-shaped requests at /api/v1/systemone, and the playground has a model picker. Keys are on API access.
Which decision model to use
For yes/no gates the spread was small, so pick on latency and price. For dozens of intents, pick a model that takes them (Clef Flash led at 77 options). For thresholds that trigger actions, weigh the confident-and-wrong column above accuracy. Then run your own items: these are public datasets.
Questions
Is OpenAI's Decisions API the same as Jev?
The same idea, not the same model. Both answer typed questions with probabilities instead of writing text. OpenAI's Decisions API runs GPT-6 Luna at POST /v1/decisions; Jev is TypeSafe's model, where the yes/no type is called noul rather than predicate.
How much does the Decisions API cost?
OpenAI charges $0.10 per million input tokens and nothing for output. Jev lists $0.042. Per decision, our run billed $0.000040 for GPT-6 Luna Decisions and $0.000030 for Jev on the same items, because the models count tokens differently.
Which decision model is most accurate?
On simple choices they tie: every model scored between 90.4% and 96.5% on news topics and support tickets. With 59 and 77 options the field spreads out, and Clef Flash led at 92.0%.
Can I call other decision models with the OpenAI SDK?
Yes, through Jagent: point the SDK's baseURL at https://jev-agent.com/api/v1 and use a Jagent key. decisions.create then accepts gpt-6-luna, jev-latest or any model listed at /api/v1/models.