OpenAI opens the Decisions API to all developers in public beta: gpt-6-luna typed answers at $0.10 per 1M input tokens, output free
On Oct 6, 2026, a week after announcing it at DevDay, OpenAI released the Decisions API (POST /v1/decisions) in public beta. It runs gpt-6-luna and returns typed answers (a probability, a choice from fixed options, or a rubric score) about 10x faster than the Responses API. It costs $0.10 per 1M input tokens with no output or caching charges, and OpenAI expects general availability "in the coming weeks". It is OpenAI's answer to TypeSafe's Jev and the decision models that Cloudflare and Perplexity shipped in the meantime.
Key facts
- API changelog (Oct 6): 'Released the Decisions API in beta with gpt-6-luna. Turn text and images into typed answers 10x faster than the Responses API.'
- Endpoint: POST https://api.openai.com/v1/decisions; gpt-6-luna is the only supported model; Playground at platform.openai.com/decisions
- Three question types: predicate (probability 0–1 that a condition is true), choice (one of the supplied values, with a probability distribution and a confidence field) and score (probability-weighted average over ordered rubric levels); several independent questions can share one input
- Input: text, or user messages with text and images
- Price: $0.10 per 1M input tokens; no output-token, cache-read or cache-write charges; regional-processing premiums and long-context multipliers apply
- Supports Zero Data Retention and HIPAA for eligible customers; data residency in the US and Europe (EEA + Switzerland)
- Status: public beta, GA expected 'in the coming weeks'; SDK support from Python 3.26.0, JavaScript 7.30.0, Go 3.73.0, Ruby 0.101.0, Java 4.78.0
- Comparison (Simon Willison): Jev charges 4.2 cents per 1M input tokens, OpenAI 10 cents; Perplexity's Decisions API is $0.04; unlike Jev, gpt-6-luna takes images
- Same day, OpenAI also cut API usage tiers from five to three (Build, Launch, Grow); on Oct 5 it added an in-product HIPAA BAA flow for API organizations
What happened
OpenAI announced the Decisions API at DevDay on Sept 29, 2026, in limited preview, and promised a broad release "in the coming days"
without giving a price. On Oct 6 the API changelog listed the Decisions API as released in public beta. The new guide explains how it works.
A request carries shared input (text and/or images) and an array of named questions, each of type predicate, choice or score.
The response returns an answers array with probabilities rather than generated text, plus a refusal type. OpenAI says this is about
10x faster than the Responses API. OpenAI recommends setting thresholds with labelled examples from your own application. Only input
tokens are billed, at $0.10 per 1M, which is gpt-6-luna's normal input rate.
Simon Willison released an llm-openai-decisions plugin the same day. He noted that the API shape is "very similar to Jev" (TypeSafe's decision model)
and that OpenAI's input price is about 2.4x Jev's.
Why it matters
Decision models (one forward pass that scores a fixed set of answers) went from one startup's product (Jev, Sept 15) to a category with OpenAI, Cloudflare and Perplexity offerings within three weeks. OpenAI's version is the most expensive per input token of the four, but it accepts images and supports ZDR, HIPAA and EU data residency, which matter to the enterprise moderation, routing and triage work these models target. No accuracy benchmarks or maximum number of options were published.
Changelog
- 2026-10-07: created (resolves upcoming item 2026-10-02-openai-decisions-api-broad-release)
- 2026-10-07: video pass: linked openai-introducing-the-decisions-api
Videos (1)
Introducing the Decisions API
OpenAI · 2026-10-06 · officialDescription by Gemini, which watched the video:
Summary
Romain Huet (Head of Developer Experience at OpenAI) and Charlie Guo (Developer Experience at OpenAI) introduce OpenAI’s Decisions API. The API is powered by GPT-6 Luna and designed to return structured categorical decisions from text and visual inputs with sub-second latency.
What is shown
- Text extraction & lead routing UI: Demonstrating unstructured order text extraction into form fields [00:48] and inbound sales lead qualification [01:03] classifying company type, size, requirements, and next steps in 81 ms across six decisions [01:14].
- Vision lane-driving simulation: A top-down 2D driving game where camera frames of the road and obstacles are processed in real time by the Decisions API to select lane 1, 2, or 3 to steer the car clear of roadblocks [01:31].
- Voice avatar facial expressions: A conversational web application pairing GPT-Live 1 for voice with the Decisions API, which dynamically selects facial expressions for an animated frog character corresponding to conversational sentiment [02:17].
- Robotics object tracking with "Lavender": A tabletop programmable robot (a preview unit of Hugging Face’s Microdot robot) running GPT-Live and the Decisions API. The camera feed evaluates video frames to control head orientation to track a green apple [03:39], follow "the fruit" when asked ambiguously [04:12], and choose to track an Xbox controller over an apple when asked "what's more fun to play with?" [04:31].
- Post-credits blooper: Romain asks the robot if it is excited about the Decisions API, and the robot shakes its head horizontally [05:27].
Claims & numbers
- The Decisions API runs on GPT-6 Luna (presenter says at [00:19]).
- Focusing the model on constrained multiple choices makes it nearly 10 times faster than full generative output while maintaining image understanding, broad language support, and safety protection (presenter says at [00:28]).
- In the inbound sales demo, the server-side Decisions API processes and returns six decisions in 81 ms / less than 100 ms (presenter and UI show at [01:13]).
Notable quotes
- "By focusing the model on just a few choices, we can make it nearly 10 times faster..." — Romain Huet [00:28]
- "On the server side, the Decisions API here responded in less than 100 milliseconds, very impressive." — Romain Huet [01:11]
- "GPT-Live 1 handles the voice, while the Decisions API is choosing which expression the character should use as I talk." — Charlie Guo [02:20]
Assessment
This is an official OpenAI launch video demonstrating real-time interactive developer demos across web apps, vision navigation, animated agents, and hardware robotics. The displayed latencies (such as the 81 ms server processing time and real-time vision-based camera tracking) are live functional proof-of-concepts designed to highlight low-latency classification rather than benchmark evaluations.
Described by gemini-3.8-flash on 2026-10-07 from the video's audio and frames.
People
Related events
- OpenAI DevDay 2026: dots agents, GPT-6.1 Sol, Ultrafast, a $500 Pro plan and 20+ launches ★★★★
- TypeSafe AI releases Jev, a 'System One' decision model that returns typed probabilities instead of text ★★★
- Cloudflare releases Clef and Clef-flash, open-weight multimodal decision models that are Jev-API compatible ★★★
- Perplexity launches a Decisions API and open-sources pplx-decider-v1-27b, a decision model it says edges Jev (85.71% vs 84.51%) ★★
Sources (4)
- officialOpenAI API changelog (Oct 6: Decisions API beta; usage tiers)
- docsOpenAI docs: Decisions guide (question types, pricing and availability)
- officialOpenAI: DevDay 2026 recap (Decisions API announced Sept 29)
- discussionSimon Willison: llm-openai-decisions 0.1a0 (Oct 6)
id: 2026-10-06-openai-decisions-api-public-beta · updated 2026-10-07 · open in the interactive timeline