TypeSafe released Jev on September 15. OpenAI announced the Decisions API two weeks later on September 29.

The immediate comparison is obvious because both products are aimed at a class of problems where software already knows the possible answers and needs a model to make the judgment. OpenAI describes the Decisions API as focusing Luna on user defined questions with finite predefined answers, using text or images as context, with classification, request routing, and choosing an agent's next action as examples. As of October 1, 2026, OpenAI still describes it as a limited preview, with broader availability planned.

TypeSafe describes Jev differently, but the application shape overlaps quite a lot. Jev takes unstructured state and returns typed probabilistic decisions instead of generated strings. TypeSafe says Jev uses a new model architecture, a parallel sampler, and Reinforcement Learning for Calibrated Decisions. Its current public model exposes Choice, Score, and Noul as three different shapes of bounded judgment, with multiple questions able to run against the same state.

I would not spend this article trying to decide which one is better.

There is not enough public information about OpenAI Decisions API yet to make a serious comparison on pricing, latency distribution, calibration, rate limits, batching behavior, or production capacity. The more interesting thing is that two different companies have now arrived at a similar programming model.

For the last few years, we have used text generation as the default interface for almost every kind of model intelligence.

Even when software only needed a small decision, the architecture often looked like this:

Application State
       |
       v
      LLM
       |
       v
Generated Text or JSON
       |
       v
Parse / Validate
       |
       v
Application Decision
       |
       v
Action

That architecture works well, especially now that structured output is much more reliable than it used to be. But there are many places where the application never wanted a piece of language in the first place.

It wanted to know which tool to use, which queue should receive a request, whether a piece of evidence is relevant, whether an agent action looks risky, whether a human needs to review something, or which one of several known actions makes sense next.

Once the possible answers are already known, generation starts looking like only one possible way to compute the decision rather than the natural interface for the problem.

That is the part of Jev and OpenAI Decisions API that I think matters beyond the individual products.

I would separate decisions from open ended reasoning#

Consider an incident agent investigating increased payment failures.

The agent has access to logs, recent incidents, deployment history, code search, documentation, and a few tools. During one run it may need to make many judgments before it ever writes a line of code.

It may need to decide whether the incident looks related to a recent deployment, which source of evidence is worth checking next, whether the current evidence is sufficient to modify production code, which tool should be called, whether an operation is risky enough to require review, and whether the final result should be accepted or escalated.

A normal agent can keep all of that inside one reasoning loop:

User Request
     |
     v
Agent Model
     |
     +--> choose evidence source
     |
     +--> inspect evidence
     |
     +--> choose tool
     |
     +--> interpret result
     |
     +--> assess risk
     |
     +--> decide whether evidence is enough
     |
     +--> modify code
     |
     +--> decide whether task is complete

Sometimes this is exactly what we want because the problem is genuinely open ended and the model needs freedom to explore.

Other judgments are narrower. The answer space is already understood by the application.

For example:

EVIDENCE_SUFFICIENT

NEED_MORE_EVIDENCE

ESCALATE

or:

SEARCH_LOGS

SEARCH_CODE

QUERY_INCIDENTS

ASK_HUMAN

I would increasingly pull those bounded judgments into a separate decision layer and leave investigation, explanation, planning, debugging, and code generation in the reasoning layer.

The distinction is not that one layer is intelligent and the other is not. They solve different shapes of problems.

The decision layer works when the application can describe the answer space before inference.

The reasoning layer is useful when the path, intermediate work, or final answer cannot honestly be reduced to a known set of choices.

Architecture diagram separating a production AI system into a decision layer and a reasoning layer. The decision layer handles bounded judgments such as routing, classification, tool choice, verification and risk, while the reasoning layer handles investigation, planning, debugging, explanation and code generation. Both feed into deterministic policy before execution.

Bounded decisions and open ended reasoning are different workloads. Decision models choose among outcomes the application already understands, while reasoning models handle investigation, planning and problems where the path is not known in advance.

I would make the decision itself a versioned contract#

If decision models become a real infrastructure layer, I would not represent them as random model calls scattered through application code.

I would define a production artifact around each important decision.

Something like:

DecisionContract {
    decision_type
    version

    input_schema
    answer_space

    abstention_policy

    consequence_class

    required_capabilities

    quality_slo
    latency_slo

    eval_suite

    fallback_policy
}

For the incident agent, one contract might be:

TOOL_SELECTION_V7

Another:

EVIDENCE_SUFFICIENT_V3

Another:

CHANGE_RISK_V4

Each one can evolve independently.

This gives us a much cleaner unit to test than saying the agent is good or bad at tool selection. TOOL_SELECTION_V7 has a specific input contract, a specific set of possible answers, an eval suite, known production outcomes, and an owner.

It also lets us change the provider without changing the business meaning of the decision.

The provider could eventually be Jev, OpenAI Decisions API, a small classifier, deterministic code, or something else. The application still depends on TOOL_SELECTION_V7.

That provider independence is useful, but I would not take it too far.

A common interface should not erase provider semantics#

It would be very easy to create:

DecisionResult {
    answer
    probability
}

and then implement:

JevAdapter
OpenAIDecisionsAdapter

That looks neat, but it can create a dangerous abstraction.

A number between zero and one does not automatically mean the same thing across models.

It might be a normalized score.

It might be token probability.

It might be a confidence estimate.

It might be a calibrated probability.

Those values can look identical in JSON and have completely different operational meaning.

TypeSafe makes calibrated probabilities part of the Jev story. Its public material says answers include calibrated probabilities and confidence information, while the current Choice, Score, and Noul documentation describes how those values are returned for the different question types.

OpenAI's public Decisions API announcement currently says developers provide finite predefined answers and receive decisions from Luna, but it does not yet document enough public probability semantics for me to treat its output as equivalent to Jev's calibration contract.

Until that documentation exists, I would preserve the evidence honestly.

DecisionEvidence {
    selected_answer

    provider
    provider_model_version

    raw_scores
    score_semantics

    probability_distribution_available
    calibrated_probability_supported

    decision_contract_version

    latency_ms
}

The policy layer can then make different choices depending on what the provider actually guarantees.

If one provider gives a calibrated distribution and another gives only a selected answer with a different confidence semantic, the abstraction should expose that difference instead of hiding it behind the field name probability.

A closed answer space still needs somewhere for uncertainty to go#

This is one of the easiest mistakes to make when moving work from a reasoning model into a decision model.

Suppose the allowed answers are:

ALLOW

REVIEW

BLOCK

What should the model return when the state is incomplete?

If we do not give uncertainty a valid representation, the model is forced to choose the least wrong answer.

For many production decisions I would rather have:

ALLOW

REVIEW

BLOCK

NEED_MORE_EVIDENCE

or:

UNKNOWN

Depending on the API, abstention may be represented through an explicit option, a probability threshold, a confidence mechanism, or normal application policy.

The important point is that a bounded answer space is only safe when the answer space represents the states that can actually happen.

If the incident agent has only seen one noisy log line, forcing it to choose:

SEARCH_CODE

MODIFY_CODE

is a bad decision contract.

A better contract may be:

SEARCH_LOGS

SEARCH_CODE

QUERY_INCIDENTS

ASK_HUMAN

NEED_MORE_EVIDENCE

A good decision layer should make uncertainty explicit instead of hiding it behind a forced choice.

Not every decision is independent#

There is another architectural problem once we start extracting many decisions out of an agent prompt.

It is tempting to submit everything together:

intent?

risk?

tool?

human review?

model route?

and parallelize all of it.

Sometimes that is valid.

Sometimes the questions depend on each other.

For the incident agent, base risk may be computed from the request and account state. Tool choice may depend on that risk. Whether human review is required may then depend on both the selected tool and the consequence of the operation.

The architecture is not always:

State
  |
  +--> Decision A
  +--> Decision B
  +--> Decision C
  +--> Decision D

It can be:

I would parallelize decisions when they depend on the same immutable state and do not depend on one another's result.

Once one judgment constrains another, I would represent that dependency explicitly rather than hiding it inside one large question set or one prompt.

This also gives us better observability because we can tell whether the final review decision changed because risk changed, because tool choice changed, or because the review policy itself changed.

Decision graph for an incident agent showing which AI decisions can run in parallel and which depend on earlier results. Intent and base risk can be evaluated from the same incident state, while tool selection depends on those judgments. Evidence sufficiency can result in continuing, gathering more evidence, or escalating to a reasoning model or human before policy allows execution or review.

Some decisions can run in parallel against the same immutable state, while others depend on earlier judgments or new evidence. Explicit dependencies and abstention paths are easier to reason about than hiding the complete control flow inside one agent prompt.

I would put a decision service in front of both providers#

The business application should not know that TOOL_SELECTION_V7 happens to use Jev today and Luna tomorrow.

I would place a decision service between the application and the model providers.

Application
     |
     v
Decision Contract
     |
     v
State Builder
     |
     v
Decision Service
     |
     +--> deterministic rule
     |
     +--> Jev
     |
     +--> OpenAI Decisions API
     |
     +--> small model
     |
     +--> reasoning model
     |
     +--> human review

The decision service owns provider selection, contract versions, capability checks, timeouts, fallback behavior, evaluation results, and operational policy around the call.

This does not mean every request should dynamically choose among five providers. Most production systems are easier to operate when one decision contract has a stable primary implementation.

The service boundary is still useful because it prevents provider details from leaking into every application and gives us one place to observe and control decision behavior.

Provider routing should be capability aware#

OpenAI explicitly advertises text and image context for the Decisions API. TypeSafe's current public Jev material focuses on unstructured program state, typed outputs, probabilities, and its Choice, Score, and Noul primitives. I would not assume multimodal parity until the provider documentation says so.

That means provider routing is not only a latency or cost problem.

A decision contract may require capabilities such as:

DecisionCapabilities {
    text_input
    image_input

    categorical_choice
    ordered_score
    binary_probability

    probability_distribution

    calibration_contract

    multiple_questions
}

Suppose the incident agent is interacting with a browser and the current state is a screenshot showing an approval dialog. If the decision contract needs the image directly, a provider with multimodal context may be eligible while another provider is not.

For another decision, the application may care much more about a calibrated probability distribution than image understanding.

Those are different requirements.

The decision service should select from providers that satisfy the contract instead of trying providers until one returns HTTP 200.

Production architecture for a versioned AI Decision Contract. An application sends a named decision such as TOOL_SELECTION_V7 through a state builder and decision service. Capability checks determine which implementations are eligible, including deterministic rules, Jev, OpenAI Decisions API, reasoning models or human review. Provider specific outputs remain Decision Evidence with their original score and calibration semantics before policy determines what can happen.

A production decision should be versioned independently from the model provider. The decision service selects an implementation that satisfies the contract while preserving provider specific evidence semantics for policy and evaluation.

Some decisions still belong in normal code#

Separating a decision layer does not mean replacing deterministic logic with models.

If the refund amount exceeds a configured account limit, I do not need a model to decide whether approval is required.

I need:

if refund_amount > refund_limit:
    require_approval()

The model becomes useful when the input is fuzzy.

For example:

Does this customer request look like
a normal billing dispute or possible fraud?

or:

Does the evidence collected by the incident agent
actually support modifying production code?

So I would use a hierarchy, but not as a pipeline that every request passes through.

Deterministic Code

      ↓ when judgment is fuzzy

Decision Model

      ↓ when the problem is not honestly bounded

Reasoning Model

      ↓ when consequence or ambiguity requires it

Human

The application should enter at the cheapest and most deterministic layer that can answer the question correctly.

That keeps models away from work software already knows how to do and keeps bounded decision models away from problems that need real investigation.

The decision result is still not authority#

Suppose Jev returns:

APPROVE_CHANGE  0.997

or OpenAI Decisions API selects:

APPROVE_CHANGE

for the incident agent.

I would not let that result directly modify production.

The model supplied judgment.

Normal software still owns authority.

Decision Model
      |
      v
Candidate Decision
      |
      v
Policy Layer
      |
      +--> agent identity
      +--> environment
      +--> change scope
      +--> consequence
      +--> approval requirements
      +--> current incident policy
      |
      v
Allowed Action

This separation becomes more important if decision models become fast and cheap enough to sit everywhere in the runtime.

When a model call costs almost nothing relative to the workflow, it becomes tempting to ask the model questions that actually belong to authorization or policy.

There is an important difference between:

Which action appears appropriate?

and:

Is this actor allowed to execute the action?

I am comfortable using model judgment for the first.

I want the second enforced outside the model wherever possible.

Agents are where I expect this architecture to matter most#

The incident agent may run for thirty minutes and make dozens or hundreds of local judgments while completing one task.

It may repeatedly decide which evidence is relevant, which tool to use, whether a result changes the current hypothesis, whether another search is useful, whether enough evidence exists to modify code, whether a change looks risky, whether a human should review the action, and whether the task is actually complete.

Using a large reasoning model for every one of those judgments can make the runtime expensive and slow. Keeping every rule and routing instruction inside the main agent context also makes that context harder to understand and evaluate.

A separate decision layer lets the main reasoning model spend more of its context and computation on investigation and problem solving.

For the incident agent I might eventually have:

                         Agent Runtime
                              |
               +--------------+--------------+
               |                             |
               v                             v
        Decision Service               Reasoning Model
               |                             |
          +----+----+                        |
          |         |                        |
          v         v                        |
        Jev      OpenAI                      |
                 Decisions                   |
          |         |                        |
          +----+----+                        |
               |                             |
               v                             v
        Decision Evidence                Investigation
               |                             |
               +--------------+--------------+
                              |
                              v
                          Agent Policy
                              |
                              v
                         Tool / Action

This does not mean every tool selection should become a remote model call.

Some choices can remain inside the main agent.

Some can be deterministic.

Some are worth externalizing because they are important enough to evaluate independently, reused across workflows, or need a different latency and cost profile from the reasoning model.

That boundary should be driven by operational value, not architectural fashion.

Image context makes some bounded workflows more interesting#

OpenAI explicitly says Decisions API can use images as context.

That opens decision paths where the application does not need a full visual explanation.

A browser agent may only need to determine whether the current screen is:

LOGIN

PAYMENT_CONFIRMATION

CAPTCHA

ERROR

SENSITIVE_APPROVAL

NORMAL_PAGE

A document pipeline may need to decide:

INVOICE

PASSPORT

RECEIPT

CONTRACT

UNKNOWN

before choosing the next processing system.

The application can send the image directly to a bounded decision interface instead of asking a general model to describe the entire screen and then parsing that description to recover a route.

I would still benchmark image decisions separately from text decisions because the current public OpenAI material does not give enough information to assume identical latency, cost, or accuracy characteristics across modalities.

Calibration is where decision models become more than fast classifiers#

Fast classification itself is not new. We have had classifiers for a long time.

The more interesting production problem appears when software wants to automate based on uncertainty.

Suppose a provider returns something equivalent to:

ALLOW     0.96

REVIEW    0.03

BLOCK     0.01

The application might want to use different policy for different confidence ranges.

For example:

very high confidence
automatic action

middle range
human review

low confidence
gather more evidence

I would not hardcode those bands from a vendor example.

They need to come from observed performance on our traffic.

TypeSafe makes calibration an explicit part of Jev's design and says RLCD is meant to produce calibrated decisions.

OpenAI's current Decisions API announcement does not yet describe a calibration contract publicly.

That does not imply Luna is poorly calibrated. It means I would wait for the real API semantics and measure the behavior before building production thresholds around any score it returns.

A field called confidence = 0.97 is not proof that the decision is correct 97 percent of the time.

Calibration is something we measure.

Fallback is a behavioral compatibility problem#

This is one place where the provider abstraction becomes dangerous.

Suppose TOOL_SELECTION_V7 normally uses Jev and Jev becomes unavailable.

The simplest implementation is:

try Jev

if unavailable:
    call OpenAI Decisions API

The second call may succeed technically.

That does not mean the fallback is safe.

Maybe the two providers have different behavior on ambiguous cases.

Maybe one over selects an expensive tool.

Maybe one has a higher false allow rate on destructive actions.

Maybe one handles very long state differently.

Maybe the fallback does not support the same calibration semantics.

So I would make fallback eligibility part of the decision contract.

FallbackPolicy {
    allowed_providers

    minimum_eval_version

    maximum_quality_delta

    consequence_constraints

    fallback_on_timeout

    fallback_on_provider_error
}

A fallback provider should already have been evaluated on that decision contract.

If it has not, switching to it during an outage is not ordinary resilience.

It is a new model deployment during an incident.

I would evaluate decisions, not brands#

I would not build one benchmark called:

Jev benchmark

and another called:

OpenAI Decisions benchmark

I would build:

TOOL_SELECTION_V7

CHANGE_RISK_V4

EVIDENCE_SUFFICIENT_V3

HUMAN_REVIEW_V5

Then every eligible provider runs against the same evaluation cases.

For TOOL_SELECTION_V7, I may care about top choice accuracy, dangerous tool selection rate, latency, cost, abstention behavior, and how often the decision causes a later agent failure.

For CHANGE_RISK_V4, false allow rate may matter much more than overall accuracy.

For EVIDENCE_SUFFICIENT_V3, the important failure may be saying yes too early.

That makes the evaluation useful for architecture.

The question is not:

Is Jev good?

or:

Is OpenAI Decisions API better?

The question is:

Can this provider version safely satisfy this decision contract on our traffic?

That is a much smaller and much more answerable question.

Decision dependencies should be evaluated too#

Once decisions form a graph, evaluating each node in isolation is not enough.

Suppose BaseRisk is slightly wrong but ToolChoice is robust to that error.

The final behavior may still be fine.

Or BaseRisk may look accurate overall, but a small error on one particular class may consistently push ToolChoice toward a dangerous operation.

So I would keep both local and graph level evaluation.

Decision Node Eval

TOOL_SELECTION_V7
accuracy
unsafe selection rate
abstention
latency

and:

Decision Graph Eval

state
  |
risk
  |
tool choice
  |
review policy
  |
action

The graph evaluation tells us whether errors compound.

This is especially important when the output of one decision changes the state observed by later decisions.

A series of individually 95 percent accurate decisions does not imply a 95 percent reliable workflow.

Composition matters.

Version the decision behavior independently from the agent#

The main reasoning model may remain unchanged while the decision layer changes how the agent behaves.

Maybe we update Jev.

Maybe OpenAI updates Luna.

Maybe TOOL_SELECTION_V7 becomes TOOL_SELECTION_V8.

Maybe the state builder starts including different evidence.

Maybe calibration thresholds move.

Any of those can change the trajectory of the same agent.

I would record at least:

decision_type

decision_contract_version

provider

provider_model_version

state_builder_version

question_set_version

calibration_profile_version

policy_version

Then a production trace can tell me why a run from Tuesday behaves differently from the same task on Thursday.

A decision provider upgrade is a behavioral deployment even when the main agent model did not change.

I would run it through historical evaluation, shadow traffic where possible, slice analysis, and a canary before moving consequential decisions broadly.

The observability path needs the eventual outcome#

For each decision I would record:

run_id

step_id

decision_type

decision_contract_version

provider

model_version

state_hash

candidate_answers

selected_answer

raw_scores

score_semantics

policy_result

latency

actual_outcome

The actual_outcome may not exist immediately.

For tool selection we may know quickly whether the tool helped.

For fraud or payment decisions, the truth may arrive hours or days later.

For a coding agent, whether a change was actually correct may only become clear after tests, review, or production behavior.

That means the decision platform needs a feedback path from eventual outcomes back to the original decision.

Without outcomes, we can monitor latency, provider availability, output distributions, and drift.

We cannot know whether quality or calibration changed.

Production feedback loop showing a Decision Contract flowing through a Decision Service into Decision Evidence and deterministic policy. Policy can act, request review, gather more evidence, or escalate to reasoning or a human. The resulting outcome is joined back to the original decision and used for decision evaluation, graph evaluation, calibration, policy updates and the next version of the Decision Contract.

A model answer is only decision evidence. Policy determines what that evidence is allowed to cause, and the eventual outcome feeds evaluation, calibration and the next version of the decision contract.

I would keep decision prompts or questions narrow#

There is an obvious failure mode if teams adopt this pattern aggressively.

We remove routing and guardrails from the main agent prompt, then create one enormous decision definition containing every tool, every policy, every exception, and every business rule.

Now the decision layer has become another giant prompt that nobody understands.

I would keep contracts narrow enough that someone can explain what they do.

TOOL_SELECTION_V7

MODEL_ROUTE_V3

HUMAN_REVIEW_V5

CHANGE_RISK_V4

Each has its own state requirements, answer space, eval suite, owner, and consequences.

Several independent questions can still share the same request when that makes sense. Jev explicitly supports mixing multiple question types in one request, and OpenAI may expose more detail around batching or question composition as the Decisions API moves beyond preview.

The implementation optimization should not destroy the conceptual boundary between the decisions.

OpenAI entering this space changes how seriously I take the category#

TypeSafe describes Jev as a new System One model built specifically around structured probabilistic decisions. OpenAI is approaching the problem through Luna and a new Decisions API focused on finite answers. There is no public reason to assume the internal architectures are similar.

The common part is the application contract.

Software already knows the shape of many decisions it needs to make. It does not always need a model to generate language before choosing among them.

That creates a useful architecture boundary:

                         APPLICATION
                              |
                              v
                        STATE BUILDERS
                              |
                              v
                     DECISION CONTRACTS
                              |
                              v
                       DECISION SERVICE

                rules
                Jev
                OpenAI Decisions
                other models

                              |
                              v
                      DECISION EVIDENCE
                              |
                 +------------+------------+
                 |                         |
                 v                         v
              POLICY                  REASONING
                 |                         |
                 +------------+------------+
                              |
                              v
                          EXECUTION
                              |
                              v
                            OUTCOME
                              |
                              v
                    EVALUATION / CALIBRATION

I would not put every model judgment in this layer.

If the answer space is not known, the decision contract is probably wrong.

If the model needs to investigate before it can answer, use the reasoning layer.

If normal code already knows the answer, use normal code.

If uncertainty cannot be represented safely, fix the contract before automating it.

But there is a large middle ground where the application understands the possible outcomes and needs intelligence only in choosing among them. Routing, relevance, risk, verification, tool selection, skill selection, review decisions, and many agent control points sit in that middle ground.

Jev made that design space much more explicit.

OpenAI Decisions API makes it harder to treat it as one company's unusual model interface.

If this category keeps growing, I expect production AI systems to look less like one giant model sitting in the middle of everything and more like several computational layers with different jobs. Deterministic code will handle known rules, decision models will make bounded judgments, reasoning models will handle open ended work, policy will control authority, and execution infrastructure will own the actual side effects.

That feels like a healthier architecture anyway, because it gives each kind of intelligence a smaller and more testable job.