Home / Blog / OpenAI Decisions API Returns Model Classifications in 150…
Tech News

OpenAI Decisions API Returns Model Classifications in 150 Milliseconds

OpenAI's Decisions API constrains GPT-6 Luna to predefined options and answers in about 150 milliseconds — near 10x faster than a standard Luna API call.

By Dillip Chowdary • Sep 30, 2026 • Source: OpenAI

OpenAI Decisions API Returns Model Classifications in 150 Milliseconds

OpenAI launched the Decisions API at DevDay 2026, a purpose-built endpoint for classification and single-decision tasks that returns answers in roughly 150 milliseconds — against about 1.6 seconds for an equivalent standard call — by constraining the model to a predefined set of answer options. It runs on GPT-6 Luna, accepts text or image context, and ships as a limited preview with broader rollout expected shortly.

This piece explains what the Decisions API changes architecturally, where a ten-fold latency reduction matters, how it compares to the workarounds developers currently use for fast classification, and what to test during the preview. It is written for engineers building routing layers, moderation pipelines, and agent control flow.

The Decisions API: what OpenAI launched

The mechanics are deliberately narrow: a developer defines a question and its allowed answers up front, supplies context — text or images — and the API returns which option applies. By refusing open-ended generation entirely, the serving path skips most of what makes LLM calls slow; OpenAI quotes approximately 150 milliseconds per decision versus 1.6 seconds for asking GPT-6 Luna the same thing through the standard API — roughly a ten-fold acceleration.

OpenAI names two target uses: content routing and agent decision-making. Both share a shape — high call volume, small answer space, and position on the critical path where every millisecond of model latency multiplies through the system.

Why constrained decisions deserve their own endpoint

OpenAI Decisions API Returns Model Classifications in 150 Milliseconds
Illustration · Pexels

A large share of production LLM calls are not conversations; they are predicates. Which queue does this ticket belong in, is this input safe, did this step succeed, which tool should run next — developers have been answering these with full chat-completion calls, prompt-engineered to output one word, paying full generation latency for a categorical answer. The Decisions API formalizes that pattern into a contract: options in, option out.

The constraint is also a correctness feature. A model restricted to predefined options cannot hallucinate an out-of-vocabulary answer, drift into explanation, or return unparseable output — failure modes every production classification prompt currently defends against with retries and parsing shims. Removing that defensive layer simplifies the calling code as much as the latency win speeds it up.

Advertisement

Tech Pulse Daily

Get tomorrow's pulse first

Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.

Where 150 milliseconds changes the architecture

At 1.6 seconds, model-backed decisions live off the hot path: developers batch them, cache them, or accept sluggishness. At 150 milliseconds, a model decision fits inside a web request, a chat turn, or an agent step without dominating it. Moderation checks can gate content synchronously at post time; routers can classify per message; agent loops that make dozens of control-flow decisions per task stop paying seconds of pure decision overhead per run.

The agent case is the strategic one for OpenAI. The same keynote shipped always-on dots, computer use in the Agents API, and multi-agent features — architectures whose inner loops are dense with small decisions. A fast, cheap, constrained decision primitive is the connective tissue that makes those loops responsive, and its arrival on GPT-6 Luna — OpenAI's small-model line — signals the intended cost profile.

Who should adopt it, and against what alternatives

Teams currently running fine-tuned small classifiers get a genuine build-versus-buy question: a hosted 150-millisecond frontier-family classifier with zero training pipeline competes directly with maintaining an in-house model, especially for teams whose categories change often — updating a list of options is a config change, not a retraining run. Teams using function-calling or logit tricks to force categorical outputs from chat endpoints get the same behavior with a cleaner contract and most of their latency back.

It is not a fit for decisions needing visible reasoning, multi-step deliberation, or answers outside a known set — those remain standard-API work. The image-context support widens the surface: visual moderation and screenshot-based routing fall inside the same fast path.

Preview status and what to watch

Access is limited-preview now, with OpenAI saying broader availability is imminent; pricing was not published at the keynote, and per-decision cost is the number that determines whether high-volume routing workloads move over. Also unpublished: option-count limits, context-size ceilings, and latency distribution under load — the 150-millisecond figure needs a p99 before it anchors an SLA.

During the preview, the test that matters is accuracy against your existing classifier on your real distribution, not the latency demo. Speed only pays if decision quality holds; run the shadow comparison first, and watch for the pricing announcement to complete the cost math against fine-tuned alternatives.

Developer Action Items

  • ☐ Diff the official changelog for OpenAI 1.6 before you bump — APIs, defaults, and removed flags only.
  • ☐ Install through the vendor's documented channel in staging; keep a one-command rollback and time-box the canary.
  • ☐ Grep your repo for old flag names, lockfile pins, and plugin versions that the notes mark as breaking.
  • ☐ Prefer the first patch cut over the day-zero tag unless you have a reason to be on the leading edge.
  • ☐ If OpenAI did not name a region, plan, or SKU, screenshot the official availability line before you promise it to users.
Dillip Chowdary

Author

Dillip Chowdary

Writes Tech Bytes coverage of AI, engineering, and the tools that actually ship. Editor of Tech Pulse Daily.

Related on Tech Bytes

Advertisement

5-min tech signal

Weekday briefing for engineers who skip the noise.

No spam · Unsubscribe anytime

Advertisement

✈️ CareerPilot

Your AI job-search copilot

Match your resume against live Ashby, Greenhouse & Lever openings — fit scores, job-specific resume optimization and email alerts.

Find matching jobs →

Free Tools

Browse all tools →