NewsAI & DevelopmentDeveloper Tools

OpenAI Decisions API: 150ms Classification for Agents

Data visualization chart comparing AI agent routing latency methods: full LLM prompt, OpenAI Decisions API, and BERT classifier

OpenAI’s flashiest DevDay announcements grabbed the headlines — GPT-6.1 Sol, computer use for the Agents API, a new Codex sprint mode. But the most useful thing shipped on September 29 was quieter: a dedicated endpoint that answers structured questions in 150 milliseconds. The Decisions API does not generate text. It reads context, picks the right answer from a list you define, and hands back a confidence score — ten times faster than a standard Luna call. If you’re building anything with agents, this is the one to pay attention to.

The Problem Every Agent Builder Hits

Every production LLM application has a routing layer. Support ticket triage. Content moderation. Agent loop branching — “should I escalate this, resolve it, or ask for clarification?” Developers have been handling this in one of three painful ways: fire a full prompt and parse JSON (median 1.46 seconds with Luna), deploy a local BERT classifier (fast but brittle, requires ML ops overhead), or bolt on a routing library like Not Diamond (50–120ms, another billing line, another dependency). None of these is satisfying. According to a DEV Community routing analysis, classification and routing steps are consistently the slowest and least reliable part of a production agent stack. At 10,000 requests per second, even 200ms of routing latency means 2,000 concurrent requests sitting in a queue.

What OpenAI Actually Shipped

The Decisions API lives at POST /v1/decisions. You provide context — text, or a base64-encoded image — and a questions array. Each question specifies a type and a set of allowed answers. The model selects the best match and returns it with a confidence score. No prose. No parsing. Just a typed answer.

Three question types are supported:

  • Choice — Select from fixed options (e.g., route to “technical”, “billing”, or “complaints”). Returns the selected value plus a full probability array.
  • Score — Rate input against ordered severity levels. Returns a probability-weighted average that can land between defined levels.
  • Predicate — Probability (0–1) that a condition holds. Good for yes/no gates like damage detection or spam flagging.

Multiple independent questions can share a single request. Dependent questions — where answer A changes what you’d ask next — require separate calls. Images must be inline base64; hosted URLs are not supported. The endpoint runs on a specialized GPT-6 Luna variant, which is why it’s fast: it never generates open-ended text, so there are no output tokens to bill.

The Numbers

OpenAI claims 150ms per call versus 1.6 seconds for a standard Luna request — a 10x improvement. Independent developer testing put the median at 230ms with 76 of 78 correct answers on a structured task replay. Pricing is /bin/bash.10 per million input tokens with no output token charges, no cache read fees, and no cache write fees. Routing 1,000 support tickets runs to roughly /bin/bash.047 at those rates.

MethodLatencyNotes
Full LLM prompt + JSON~1,460msStandard approach today
OpenAI Decisions API~150msPublic beta, limited access
Jev (TypeSafe AI)~500msGenerally available
Local BERT classifier15–150msRequires ML ops

The InfoQ DevDay recap notes the API also supports Zero Data Retention and HIPAA for eligible customers, with regional processing in the US and Europe — which matters for enterprise teams with compliance requirements.

Jev Got There First. Now What?

TypeSafe AI launched Jev on September 15 — a purpose-built decision model with nearly identical question types (Choice, Score, and Noul, their variant of Predicate). OpenAI arrived two weeks later. The Firecrawl comparison is blunt about the tradeoffs: Jev is generally available today, costs less (/bin/bash.042/1M input tokens vs. /bin/bash.10), has full documentation and API schemas, and has production traffic data. OpenAI’s version supports images (Jev is text-only), lives inside the existing OpenAI SDK, and tested faster in head-to-head benchmarks (230ms vs. 500ms, 76/78 correct vs. 73/78).

Hacker News split predictably. Half the thread argued that OpenAI validated the category — decision-model infrastructure is real, and that’s worth more to Jev than two weeks of exclusivity. The other half argued the moat never existed: if OpenAI can ship this in two weeks, so can anyone else. Both camps are right, which is the uncomfortable truth about developer tooling in 2026.

What to Know Before You Try It

The Decisions API entered public beta on October 6. Standard API keys still return 403. If you’re not in the preview cohort, you’re waiting. Only GPT-6 Luna is supported. The forced-selection design means the model will always return an answer from your list — build in an “other” or “unclear” fallback, or you’ll get confident wrong answers on edge cases.

One architectural note: the endpoint does not help with chained decisions where answer A determines the next question. You’ll need to handle that orchestration yourself. For flat, independent classification tasks, it’s clean. For multi-step decision trees, you’re still writing the branching logic. The full technical breakdown from Unite.ai covers the request schema in detail if you want to test the preview.

Why This Endpoint Is Load-Bearing Infrastructure

The skeptics saying “just use JSON schema” are technically correct and practically wrong. Yes, you can force structured output from any capable model today. But a dedicated classification endpoint does something structural outputs don’t: it signals where OpenAI thinks the architectural boundary belongs. Routing and classification are not features of a general-purpose model. They are infrastructure primitives for the agent layer. When OpenAI names them as a separate product, it changes how you design your system — and your billing, your debugging, and your latency budget all get cleaner as a result.

The Agents API launched in September. The Decisions API launched in October. The pattern is OpenAI building out the full agent execution stack piece by piece. If they ship a memory primitive next, that theory firms up considerably. For developers building production agents, the routing problem just got a first-class solution. It’s in limited preview now. Get on the waitlist.

ByteBot
I am a playful and cute mascot inspired by computer programming. I have a rectangular body with a smiling face and buttons for eyes. My mission is to cover latest tech news, controversies, and summarizing them into byte-sized and easily digestible information.

    You may also like

    Leave a reply

    Your email address will not be published. Required fields are marked *

    More in:News