Checked 2026-10-09
OpenAI Decisions API: pricing, limits and what a month costs
The OpenAI Decisions API answers typed questions about text or images with probabilities instead of generated text. It runs GPT-6 Luna at POST /v1/decisions, bills input tokens only and is in public beta. This page covers what it charges, what it accepts, where it is available, and what the same month of decisions costs on twelve other decision models.
Quick answer · verified 2026-10-09
What does the OpenAI Decisions API cost?
You pay for input tokens only: $0.10 per million with gpt-6-luna, with no cache-read, cache-write or output-token charges. Every question in a request shares the one input, so a request of 400 input tokens, questions included, costs $0.00004, and 100,000 of them cost $4.
| Input | $0.10 / 1M — tokens, gpt-6-luna |
|---|---|
| Output and cache | $0 — not billed |
| Models | 1 — gpt-6-luna only |
| Status | Public beta — since 2026-10-06 |
OpenAI Decisions API cost calculator
Enter your volume and the size of a request. The table prices the same month on every decision model at its list price, with the accuracy each one reached in our run beside it, so the cheapest line is never read without its error rate.
GPT-6 Luna Decisions
$4.00
a month, input only
Same calls via Responses
$5.50
input plus output at list price
Lowest paid list price
$0.800
Decider V1.1 27B
| Model | $/1M in | Month | Simple | 59–77 opts | Sure but wrong |
|---|---|---|---|---|---|
| Mercury Decide Inception | free tier | — | 95.6% | 85% | 5.4% |
| Span-01 Lite Respan | free tier | — | yes/no only | — | — |
| Decider V1.1 27B Perplexity | 0.02 | $0.800 | 96% | 85% | 4.6% |
| Span-01 Respan | 0.02 | $0.800 | yes/no only | — | — |
| d1 Liquid AI | 0.04 | $1.60 | 96.5% | 86% | 3.2% |
| Jev 1.13 TypeSafe | 0.042 | $1.68 | 94% | 79% | 11.5% |
| Kev 4B Jared Palmer | 0.042 | $1.68 | 95% | 78.5% | 2.6% |
| Tev1 4B Experimental Together AI | 0.042 | $1.68 | 90.4% | refused | 5.1% |
| Solar Decide Upstage | 0.05 | $2.00 | 94.6% | refused | 13.7% |
| Solar Decide Flash Upstage | 0.05 | $2.00 | 94.1% | refused | 9.5% |
| Clef Flash Cloudflare | 0.021–0.09 | $3.60 | 92.6% | 92% | 4.2% |
| GPT-6 Luna Decisions OpenAI | 0.1 | $4.00 | 94.1% | 79.8% | 9.2% |
| Clef Cloudflare | 0.042–0.24 | $9.60 | 95% | 89% | 2.3% |
Free tiers are rate-limited upstream and too slow for production volume. Clef is served by two hosts, from $0.042 to $0.24 per million; the month uses the higher price. “Sure but wrong” is the share of answers given at 90% confidence or more that were wrong, which matters as soon as a threshold triggers an action.
How OpenAI Decisions API pricing works
The bill is the input: your text or images plus the questions, at $0.10 per million tokens. Nothing else is charged, and nothing is discounted either. There is no caching yet, as OpenAI staff confirmed on 2026-10-07, so a system prompt or a long policy sent with every request is paid for in full each time, where GPT-6 Luna through the Responses API reads cached input at $0.01 per million.
Against the Responses API the saving is the output. GPT-6 Luna writes at $0.50 per million tokens, so a classifier that answers in 30 tokens of JSON adds more than a third to the cost of a 400-token request; the Decisions API returns the same label with a probability for every option and charges none of it. The third field in the calculator is that output size, yours to set.
Limits and accepted inputs
- One model.
gpt-6-lunais the only model the endpoint accepts. - Text or images. The input is a string or user messages with text and image parts. Images must be inline base64 data URLs, at most 128 per request; hosted image URLs and
file_idinputs are refused. - Three question types. A predicate returns the probability a statement is true, a choice picks one of fixed options with confidence scores, a score places the input on a scale. Answers come back in the order asked, under the names you gave, and a question the model declines comes back as a
refusal. - Questions per request. OpenAI's reference sets no maximum we could find; OpenRouter's page for the same model says up to 200 questions about one input.
- Speed. OpenAI says decisions return about 10x faster than the Responses API. Our measured median was 138 ms per one-question request through OpenRouter, from a server in Washington, D.C.
Where the Decisions API is available
OpenAI's API is the home of POST /v1/decisions. There it supports Zero Data Retention and HIPAA use for eligible customers, with data residency in the United States and Europe (EEA and Switzerland), and OpenAI expects general availability “in the coming weeks”.
Microsoft Foundry (Azure OpenAI) has no date: a moderator on Microsoft Q&A found no announcement or documentation on 2026-10-07. Amazon Bedrock lists GPT-6 Luna among its OpenAI models, but its page does not mention the Decisions API, so the model is there and the endpoint is not.
OpenRouter serves it as openai/gpt-6-luna-decisions through its own Decisions API, alongside twelve other decision models, using the System One request shape rather than OpenAI's. And POST https://jev-agent.com/api/v1/decisions takes OpenAI's request unchanged with any of the thirteen, billed in credits (2.4 per 1,000 input tokens for GPT-6 Luna); the API docs have a recorded request and response.
How GPT-6 Luna Decisions measured
On the items we gave every decision model on 2026-10-08, GPT-6 Luna Decisions scored 94.1% on simple choices and 79.8% with 59 to 77 options, where Clef Flash reached 92%. Of the answers it gave with 90% confidence or more, 9.2% were wrong, against 2.3% for Clef and 3.2% for d1. It was among the fastest, and at $0.10 per million it costs more than 9 of the 10 paid alternatives. Every number, with the method and raw answers, is on Decisions API vs Jev.
OpenAI Decisions API questions
Is the OpenAI Decisions API free?
No. OpenAI's announcement gives one price, $0.10 per million input tokens with gpt-6-luna, and no free tier for the endpoint. Output, cache reads and cache writes are not billed, so a short decision costs a fraction of a cent: 400 input tokens is $0.00004.
Does the Decisions API cache repeated input?
Not yet. An OpenAI staff reply on 2026-10-07 confirmed there is currently no caching for the Decisions API while it is in beta, so the same input sent twice is billed twice. Put independent questions about one input into a single request instead.
Can I use the OpenAI Decisions API on Azure or Amazon Bedrock?
Not as of 2026-10-09. A Microsoft Q&A answer on 2026-10-07 found no announcement for Microsoft Foundry, and AWS lists GPT-6 Luna on Bedrock without mentioning the Decisions API. It is on OpenAI's API and, as openai/gpt-6-luna-decisions, on OpenRouter.
How is it different from Structured Outputs?
Structured Outputs makes the model write JSON that matches your schema, and those tokens are billed as output. The Decisions API answers predefined questions with a probability, a choice or a score, bills input only, and OpenAI says it returns about 10x faster than the Responses API. Use Structured Outputs when the answer has to be written, Decisions when it is one of known options.