OpenAI Decisions API vs Jev: Designing Systems Around Decisions Instead of Generation
A technical comparison of OpenAI Decisions API and Jev, and how to design a decision layer for routing, tool selection, risk, abstention, calibration and AI agent control.
On this page
TypeSafe released Jev on September 15. OpenAI announced the Decisions API two weeks later on September 29.
The immediate comparison is obvious because both products are aimed at a class of problems where software already knows the possible answers and needs a model to make the judgment. OpenAI describes the Decisions API as focusing Luna on user defined questions with finite predefined answers, using text or images as context, with classification, request routing, and choosing an agent's next action as examples. As of October 1, 2026, OpenAI still describes it as a limited preview, with broader availability planned.
TypeSafe describes Jev differently, but the application shape overlaps quite a lot. Jev takes unstructured state and returns typed probabilistic decisions instead of generated strings. TypeSafe says Jev uses a new model architecture, a parallel sampler, and Reinforcement Learning for Calibrated Decisions. Its current public model exposes Choice, Score, and Noul as three different shapes of bounded judgment, with multiple questions able to run against the same state.
I would not spend this article trying to decide which one is better.
There is not enough public information about OpenAI Decisions API yet to make a serious comparison on pricing, latency distribution, calibration, rate limits, batching behavior, or production capacity. The more interesting thing is that two different companies have now arrived at a similar programming model.
For the last few years, we have used text generation as the default interface for almost every kind of model intelligence.
Even when software only needed a small decision, the architecture often looked like this:
Application State
|
v
LLM
|
v
Generated Text or JSON
|
v
Parse / Validate
|
v
Application Decision
|
v
Action
That architecture works well, especially now that structured output is much more reliable than it used to be. But there are many places where the application never wanted a piece of language in the first place.
It wanted to know which tool to use, which queue should receive a request, whether a piece of evidence is relevant, whether an agent action looks risky, whether a human needs to review something, or which one of several known actions makes sense next.
Once the possible answers are already known, generation starts looking like only one possible way to compute the decision rather than the natural interface for the problem.
That is the part of Jev and OpenAI Decisions API that I think matters beyond the individual products.
I would separate decisions from open ended reasoning#
Consider an incident agent investigating increased payment failures.
The agent has access to logs, recent incidents, deployment history, code search, documentation, and a few tools. During one run it may need to make many judgments before it ever writes a line of code.
It may need to decide whether the incident looks related to a recent deployment, which source of evidence is worth checking next, whether the current evidence is sufficient to modify production code, which tool should be called, whether an operation is risky enough to require review, and whether the final result should be accepted or escalated.
A normal agent can keep all of that inside one reasoning loop:
User Request
|
v
Agent Model
|
+--> choose evidence source
|
+--> inspect evidence
|
+--> choose tool
|
+--> interpret result
|
+--> assess risk
|
+--> decide whether evidence is enough
|
+--> modify code
|
+--> decide whether task is complete
Sometimes this is exactly what we want because the problem is genuinely open ended and the model needs freedom to explore.
Other judgments are narrower. The answer space is already understood by the application.
For example:
EVIDENCE_SUFFICIENT
NEED_MORE_EVIDENCE
ESCALATE
or:
SEARCH_LOGS
SEARCH_CODE
QUERY_INCIDENTS
ASK_HUMAN
I would increasingly pull those bounded judgments into a separate decision layer and leave investigation, explanation, planning, debugging, and code generation in the reasoning layer.
The distinction is not that one layer is intelligent and the other is not. They solve different shapes of problems.
The decision layer works when the application can describe the answer space before inference.
The reasoning layer is useful when the path, intermediate work, or final answer cannot honestly be reduced to a known set of choices.
Bounded decisions and open ended reasoning are different workloads. Decision models choose among outcomes the application already understands, while reasoning models handle investigation, planning and problems where the path is not known in advance.
I would make the decision itself a versioned contract#
If decision models become a real infrastructure layer, I would not represent them as random model calls scattered through application code.
I would define a production artifact around each important decision.
Something like:
DecisionContract {
decision_type
version
input_schema
answer_space
abstention_policy
consequence_class
required_capabilities
quality_slo
latency_slo
eval_suite
fallback_policy
}
For the incident agent, one contract might be:
TOOL_SELECTION_V7
Another:
EVIDENCE_SUFFICIENT_V3
Another:
CHANGE_RISK_V4
Each one can evolve independently.
This gives us a much cleaner unit to test than saying the agent is good or bad at tool selection. TOOL_SELECTION_V7 has a specific input contract, a specific set of possible answers, an eval suite, known production outcomes, and an owner.
It also lets us change the provider without changing the business meaning of the decision.
The provider could eventually be Jev, OpenAI Decisions API, a small classifier, deterministic code, or something else. The application still depends on TOOL_SELECTION_V7.
That provider independence is useful, but I would not take it too far.
A common interface should not erase provider semantics#
It would be very easy to create:
DecisionResult {
answer
probability
}
and then implement:
JevAdapter
OpenAIDecisionsAdapter
That looks neat, but it can create a dangerous abstraction.
A number between zero and one does not automatically mean the same thing across models.
It might be a normalized score.
It might be token probability.
It might be a confidence estimate.
It might be a calibrated probability.
Those values can look identical in JSON and have completely different operational meaning.
TypeSafe makes calibrated probabilities part of the Jev story. Its public material says answers include calibrated probabilities and confidence information, while the current Choice, Score, and Noul documentation describes how those values are returned for the different question types.
OpenAI's public Decisions API announcement currently says developers provide finite predefined answers and receive decisions from Luna, but it does not yet document enough public probability semantics for me to treat its output as equivalent to Jev's calibration contract.
Until that documentation exists, I would preserve the evidence honestly.
DecisionEvidence {
selected_answer
provider
provider_model_version
raw_scores
score_semantics
probability_distribution_available
calibrated_probability_supported
decision_contract_version
latency_ms
}
The policy layer can then make different choices depending on what the provider actually guarantees.
If one provider gives a calibrated distribution and another gives only a selected answer with a different confidence semantic, the abstraction should expose that difference instead of hiding it behind the field name probability.
A closed answer space still needs somewhere for uncertainty to go#
This is one of the easiest mistakes to make when moving work from a reasoning model into a decision model.
Suppose the allowed answers are:
ALLOW
REVIEW
BLOCK
What should the model return when the state is incomplete?
If we do not give uncertainty a valid representation, the model is forced to choose the least wrong answer.
For many production decisions I would rather have:
ALLOW
REVIEW
BLOCK
NEED_MORE_EVIDENCE
or:
UNKNOWN
Depending on the API, abstention may be represented through an explicit option, a probability threshold, a confidence mechanism, or normal application policy.
The important point is that a bounded answer space is only safe when the answer space represents the states that can actually happen.
If the incident agent has only seen one noisy log line, forcing it to choose:
SEARCH_CODE
MODIFY_CODE
is a bad decision contract.
A better contract may be:
SEARCH_LOGS
SEARCH_CODE
QUERY_INCIDENTS
ASK_HUMAN
NEED_MORE_EVIDENCE
A good decision layer should make uncertainty explicit instead of hiding it behind a forced choice.
Not every decision is independent#
There is another architectural problem once we start extracting many decisions out of an agent prompt.
It is tempting to submit everything together:
intent?
risk?
tool?
human review?
model route?
and parallelize all of it.
Sometimes that is valid.
Sometimes the questions depend on each other.
For the incident agent, base risk may be computed from the request and account state. Tool choice may depend on that risk. Whether human review is required may then depend on both the selected tool and the consequence of the operation.
The architecture is not always:
State
|
+--> Decision A
+--> Decision B
+--> Decision C
+--> Decision D
It can be:
I would parallelize decisions when they depend on the same immutable state and do not depend on one another's result.
Once one judgment constrains another, I would represent that dependency explicitly rather than hiding it inside one large question set or one prompt.
This also gives us better observability because we can tell whether the final review decision changed because risk changed, because tool choice changed, or because the review policy itself changed.
Some decisions can run in parallel against the same immutable state, while others depend on earlier judgments or new evidence. Explicit dependencies and abstention paths are easier to reason about than hiding the complete control flow inside one agent prompt.
I would put a decision service in front of both providers#
The business application should not know that TOOL_SELECTION_V7 happens to use Jev today and Luna tomorrow.
I would place a decision service between the application and the model providers.
Application
|
v
Decision Contract
|
v
State Builder
|
v
Decision Service
|
+--> deterministic rule
|
+--> Jev
|
+--> OpenAI Decisions API
|
+--> small model
|
+--> reasoning model
|
+--> human review
The decision service owns provider selection, contract versions, capability checks, timeouts, fallback behavior, evaluation results, and operational policy around the call.
This does not mean every request should dynamically choose among five providers. Most production systems are easier to operate when one decision contract has a stable primary implementation.
The service boundary is still useful because it prevents provider details from leaking into every application and gives us one place to observe and control decision behavior.
Provider routing should be capability aware#
OpenAI explicitly advertises text and image context for the Decisions API. TypeSafe's current public Jev material focuses on unstructured program state, typed outputs, probabilities, and its Choice, Score, and Noul primitives. I would not assume multimodal parity until the provider documentation says so.
That means provider routing is not only a latency or cost problem.
A decision contract may require capabilities such as:
DecisionCapabilities {
text_input
image_input
categorical_choice
ordered_score
binary_probability
probability_distribution
calibration_contract
multiple_questions
}
Suppose the incident agent is interacting with a browser and the current state is a screenshot showing an approval dialog. If the decision contract needs the image directly, a provider with multimodal context may be eligible while another provider is not.
For another decision, the application may care much more about a calibrated probability distribution than image understanding.
Those are different requirements.
The decision service should select from providers that satisfy the contract instead of trying providers until one returns HTTP 200.
A production decision should be versioned independently from the model provider. The decision service selects an implementation that satisfies the contract while preserving provider specific evidence semantics for policy and evaluation.
Some decisions still belong in normal code#
Separating a decision layer does not mean replacing deterministic logic with models.
If the refund amount exceeds a configured account limit, I do not need a model to decide whether approval is required.
I need:
if refund_amount > refund_limit:
require_approval()
The model becomes useful when the input is fuzzy.
For example:
Does this customer request look like
a normal billing dispute or possible fraud?
or:
Does the evidence collected by the incident agent
actually support modifying production code?
So I would use a hierarchy, but not as a pipeline that every request passes through.
Deterministic Code
↓ when judgment is fuzzy
Decision Model
↓ when the problem is not honestly bounded
Reasoning Model
↓ when consequence or ambiguity requires it
Human
The application should enter at the cheapest and most deterministic layer that can answer the question correctly.
That keeps models away from work software already knows how to do and keeps bounded decision models away from problems that need real investigation.
The decision result is still not authority#
Suppose Jev returns:
APPROVE_CHANGE 0.997
or OpenAI Decisions API selects:
APPROVE_CHANGE
for the incident agent.
I would not let that result directly modify production.
The model supplied judgment.
Normal software still owns authority.
Decision Model
|
v
Candidate Decision
|
v
Policy Layer
|
+--> agent identity
+--> environment
+--> change scope
+--> consequence
+--> approval requirements
+--> current incident policy
|
v
Allowed Action
This separation becomes more important if decision models become fast and cheap enough to sit everywhere in the runtime.
When a model call costs almost nothing relative to the workflow, it becomes tempting to ask the model questions that actually belong to authorization or policy.
There is an important difference between:
Which action appears appropriate?
and:
Is this actor allowed to execute the action?
I am comfortable using model judgment for the first.
I want the second enforced outside the model wherever possible.
Agents are where I expect this architecture to matter most#
The incident agent may run for thirty minutes and make dozens or hundreds of local judgments while completing one task.
It may repeatedly decide which evidence is relevant, which tool to use, whether a result changes the current hypothesis, whether another search is useful, whether enough evidence exists to modify code, whether a change looks risky, whether a human should review the action, and whether the task is actually complete.
Using a large reasoning model for every one of those judgments can make the runtime expensive and slow. Keeping every rule and routing instruction inside the main agent context also makes that context harder to understand and evaluate.
A separate decision layer lets the main reasoning model spend more of its context and computation on investigation and problem solving.
For the incident agent I might eventually have:
Agent Runtime
|
+--------------+--------------+
| |
v v
Decision Service Reasoning Model
| |
+----+----+ |
| | |
v v |
Jev OpenAI |
Decisions |
| | |
+----+----+ |
| |
v v
Decision Evidence Investigation
| |
+--------------+--------------+
|
v
Agent Policy
|
v
Tool / Action
This does not mean every tool selection should become a remote model call.
Some choices can remain inside the main agent.
Some can be deterministic.
Some are worth externalizing because they are important enough to evaluate independently, reused across workflows, or need a different latency and cost profile from the reasoning model.
That boundary should be driven by operational value, not architectural fashion.
Image context makes some bounded workflows more interesting#
OpenAI explicitly says Decisions API can use images as context.
That opens decision paths where the application does not need a full visual explanation.
A browser agent may only need to determine whether the current screen is:
LOGIN
PAYMENT_CONFIRMATION
CAPTCHA
ERROR
SENSITIVE_APPROVAL
NORMAL_PAGE
A document pipeline may need to decide:
INVOICE
PASSPORT
RECEIPT
CONTRACT
UNKNOWN
before choosing the next processing system.
The application can send the image directly to a bounded decision interface instead of asking a general model to describe the entire screen and then parsing that description to recover a route.
I would still benchmark image decisions separately from text decisions because the current public OpenAI material does not give enough information to assume identical latency, cost, or accuracy characteristics across modalities.
Calibration is where decision models become more than fast classifiers#
Fast classification itself is not new. We have had classifiers for a long time.
The more interesting production problem appears when software wants to automate based on uncertainty.
Suppose a provider returns something equivalent to:
ALLOW 0.96
REVIEW 0.03
BLOCK 0.01
The application might want to use different policy for different confidence ranges.
For example:
very high confidence
automatic action
middle range
human review
low confidence
gather more evidence
I would not hardcode those bands from a vendor example.
They need to come from observed performance on our traffic.
TypeSafe makes calibration an explicit part of Jev's design and says RLCD is meant to produce calibrated decisions.
OpenAI's current Decisions API announcement does not yet describe a calibration contract publicly.
That does not imply Luna is poorly calibrated. It means I would wait for the real API semantics and measure the behavior before building production thresholds around any score it returns.
A field called confidence = 0.97 is not proof that the decision is correct 97 percent of the time.
Calibration is something we measure.
Fallback is a behavioral compatibility problem#
This is one place where the provider abstraction becomes dangerous.
Suppose TOOL_SELECTION_V7 normally uses Jev and Jev becomes unavailable.
The simplest implementation is:
try Jev
if unavailable:
call OpenAI Decisions API
The second call may succeed technically.
That does not mean the fallback is safe.
Maybe the two providers have different behavior on ambiguous cases.
Maybe one over selects an expensive tool.
Maybe one has a higher false allow rate on destructive actions.
Maybe one handles very long state differently.
Maybe the fallback does not support the same calibration semantics.
So I would make fallback eligibility part of the decision contract.
FallbackPolicy {
allowed_providers
minimum_eval_version
maximum_quality_delta
consequence_constraints
fallback_on_timeout
fallback_on_provider_error
}
A fallback provider should already have been evaluated on that decision contract.
If it has not, switching to it during an outage is not ordinary resilience.
It is a new model deployment during an incident.
I would evaluate decisions, not brands#
I would not build one benchmark called:
Jev benchmark
and another called:
OpenAI Decisions benchmark
I would build:
TOOL_SELECTION_V7
CHANGE_RISK_V4
EVIDENCE_SUFFICIENT_V3
HUMAN_REVIEW_V5
Then every eligible provider runs against the same evaluation cases.
For TOOL_SELECTION_V7, I may care about top choice accuracy, dangerous tool selection rate, latency, cost, abstention behavior, and how often the decision causes a later agent failure.
For CHANGE_RISK_V4, false allow rate may matter much more than overall accuracy.
For EVIDENCE_SUFFICIENT_V3, the important failure may be saying yes too early.
That makes the evaluation useful for architecture.
The question is not:
Is Jev good?
or:
Is OpenAI Decisions API better?
The question is:
Can this provider version safely satisfy this decision contract on our traffic?
That is a much smaller and much more answerable question.
Decision dependencies should be evaluated too#
Once decisions form a graph, evaluating each node in isolation is not enough.
Suppose BaseRisk is slightly wrong but ToolChoice is robust to that error.
The final behavior may still be fine.
Or BaseRisk may look accurate overall, but a small error on one particular class may consistently push ToolChoice toward a dangerous operation.
So I would keep both local and graph level evaluation.
Decision Node Eval
TOOL_SELECTION_V7
accuracy
unsafe selection rate
abstention
latency
and:
Decision Graph Eval
state
|
risk
|
tool choice
|
review policy
|
action
The graph evaluation tells us whether errors compound.
This is especially important when the output of one decision changes the state observed by later decisions.
A series of individually 95 percent accurate decisions does not imply a 95 percent reliable workflow.
Composition matters.
Version the decision behavior independently from the agent#
The main reasoning model may remain unchanged while the decision layer changes how the agent behaves.
Maybe we update Jev.
Maybe OpenAI updates Luna.
Maybe TOOL_SELECTION_V7 becomes TOOL_SELECTION_V8.
Maybe the state builder starts including different evidence.
Maybe calibration thresholds move.
Any of those can change the trajectory of the same agent.
I would record at least:
decision_type
decision_contract_version
provider
provider_model_version
state_builder_version
question_set_version
calibration_profile_version
policy_version
Then a production trace can tell me why a run from Tuesday behaves differently from the same task on Thursday.
A decision provider upgrade is a behavioral deployment even when the main agent model did not change.
I would run it through historical evaluation, shadow traffic where possible, slice analysis, and a canary before moving consequential decisions broadly.
The observability path needs the eventual outcome#
For each decision I would record:
run_id
step_id
decision_type
decision_contract_version
provider
model_version
state_hash
candidate_answers
selected_answer
raw_scores
score_semantics
policy_result
latency
actual_outcome
The actual_outcome may not exist immediately.
For tool selection we may know quickly whether the tool helped.
For fraud or payment decisions, the truth may arrive hours or days later.
For a coding agent, whether a change was actually correct may only become clear after tests, review, or production behavior.
That means the decision platform needs a feedback path from eventual outcomes back to the original decision.
Without outcomes, we can monitor latency, provider availability, output distributions, and drift.
We cannot know whether quality or calibration changed.
A model answer is only decision evidence. Policy determines what that evidence is allowed to cause, and the eventual outcome feeds evaluation, calibration and the next version of the decision contract.
I would keep decision prompts or questions narrow#
There is an obvious failure mode if teams adopt this pattern aggressively.
We remove routing and guardrails from the main agent prompt, then create one enormous decision definition containing every tool, every policy, every exception, and every business rule.
Now the decision layer has become another giant prompt that nobody understands.
I would keep contracts narrow enough that someone can explain what they do.
TOOL_SELECTION_V7
MODEL_ROUTE_V3
HUMAN_REVIEW_V5
CHANGE_RISK_V4
Each has its own state requirements, answer space, eval suite, owner, and consequences.
Several independent questions can still share the same request when that makes sense. Jev explicitly supports mixing multiple question types in one request, and OpenAI may expose more detail around batching or question composition as the Decisions API moves beyond preview.
The implementation optimization should not destroy the conceptual boundary between the decisions.
OpenAI entering this space changes how seriously I take the category#
TypeSafe describes Jev as a new System One model built specifically around structured probabilistic decisions. OpenAI is approaching the problem through Luna and a new Decisions API focused on finite answers. There is no public reason to assume the internal architectures are similar.
The common part is the application contract.
Software already knows the shape of many decisions it needs to make. It does not always need a model to generate language before choosing among them.
That creates a useful architecture boundary:
APPLICATION
|
v
STATE BUILDERS
|
v
DECISION CONTRACTS
|
v
DECISION SERVICE
rules
Jev
OpenAI Decisions
other models
|
v
DECISION EVIDENCE
|
+------------+------------+
| |
v v
POLICY REASONING
| |
+------------+------------+
|
v
EXECUTION
|
v
OUTCOME
|
v
EVALUATION / CALIBRATION
I would not put every model judgment in this layer.
If the answer space is not known, the decision contract is probably wrong.
If the model needs to investigate before it can answer, use the reasoning layer.
If normal code already knows the answer, use normal code.
If uncertainty cannot be represented safely, fix the contract before automating it.
But there is a large middle ground where the application understands the possible outcomes and needs intelligence only in choosing among them. Routing, relevance, risk, verification, tool selection, skill selection, review decisions, and many agent control points sit in that middle ground.
Jev made that design space much more explicit.
OpenAI Decisions API makes it harder to treat it as one company's unusual model interface.
If this category keeps growing, I expect production AI systems to look less like one giant model sitting in the middle of everything and more like several computational layers with different jobs. Deterministic code will handle known rules, decision models will make bounded judgments, reasoning models will handle open ended work, policy will control authority, and execution infrastructure will own the actual side effects.
That feels like a healthier architecture anyway, because it gives each kind of intelligence a smaller and more testable job.
Keep reading
Continue with a guided sequence of free production engineering essays.
Find your next reading pathRelated reading
Continue this path
Designing Production AI Systems with Jev
Jev is useful when an AI system needs fast bounded decisions around generative work. The production design still needs state construction, versioned questions, policy, calibration, observability, and clear boundaries around when to act, escalate, or collect more evidence.
Read essayA Skill Is Not a File. It Is a Deploy: Designing Agent Skills Infrastructure
Production Agent Skills are deployments, not files. They need immutable versions, controlled rollout, revocation, trust boundaries, progressive resolution, and a runtime path designed for scale.
Read essayDesigning Identity and Delegation for Production AI Agents
Production AI agents need explicit human identity, managed agent identity, trusted workload identity, narrow delegation, and an evidence chain that survives every authorization hop.
Read essay