Decision models bring structured choices, scores and probabilities to AI agents. Explore Jev, Strands Decider 2B, Clef and other 2026 releases, their open-weight and hosted options, and how they could reduce workflow latency and cost while still requiring calibration, validation and human oversight.
In the fast-moving world of artificial intelligence, most headlines still revolve around ever-larger language models that generate fluent text, write code, or produce images. Yet a quieter but potentially consequential shift has gained momentum in recent weeks. A growing group of specialized systems, known as decision models or System One models, has attracted attention. Classification and probability scoring are not new; the emerging development is their packaging into general-purpose, structured decision interfaces for software and agents. These models do not generate paragraphs. Instead, they return structured choices, scores, or probabilities drawn from options defined by the developer. Many are designed to be fast, inexpensive, and purpose-built for the repetitive micro-decisions that power modern AI agents.
Amazon’s recent open-sourcing of Strands Decider 2B stands as a clear marker of this transition. Released on 1 October 2026, the 2-billion-parameter model is based on Alibaba’s Qwen3.5 architecture and is designed to run locally on modest hardware. Its developers report competitive accuracy and calibration, with a median decision latency of approximately 115 milliseconds on an RTX 3090. The release arrives at a moment when developers are actively seeking reliable, low-latency components for agent workflows.
What follows is an examination of what decision models actually are, the notable releases of 2026, the balance between open and proprietary options, the practical beneficiaries, and some forward-looking considerations.
Understanding Decision Models: From Text Generation to Structured Choice
Traditional large language models excel at open-ended generation. They produce coherent prose, summarize documents, or reason step by step. That flexibility can become an unnecessary overhead when an application simply needs a clear yes-or-no answer, a routing decision among five support categories, or a confidence score for an agent’s next tool call. Parsing free-form text can introduce latency, cost, and the risk of malformed or incorrect outputs. Many LLMs already support schema-constrained outputs, however, so the distinction is not simply structured answers versus prose. Decision models aim to make bounded judgments more efficiently and expose probabilities directly.
Decision models invert this approach. The developer supplies context (text, structured data, or sometimes images) together with a closed set of possible answers. Many of these models evaluate the input without generating text token by token, often in a single forward pass, and return one of three common response types:
- A discrete choice from the provided options
- A numeric score on a defined scale
- An estimated probability that a statement is true (called a “Noul” in Jev-style interfaces)
Because the output space is constrained, software can consume the result without extracting an answer from generated prose. Several models also target calibrated probabilities, although calibration must be verified on the intended task. Applications still need thresholds, validation, and error handling: a valid output can contain an incorrect decision. Reported latency varies by model, input length, hardware, and deployment. TypeSafe reports 70–500 milliseconds for Jev, while some local models achieve lower latency on specific short-input benchmarks. Hosted inference can cost fractions of a cent per request, but comparisons with generative inference depend on the workload and provider pricing.
The System One terminology draws conceptual inspiration from psychologist Daniel Kahneman’s distinction between fast, intuitive System 1 thinking and slower, deliberative System 2 reasoning. Decision models aim to supply rapid judgments for agents’ routine gates, while reserving heavier language models for complex planning or explanation.
Notable Decision Models Released in 2026
Interest in the category grew rapidly after TypeSafe AI introduced Jev on 15 September 2026. Within weeks a wave of alternatives appeared, both commercial and open-weight. The following list captures notable model releases and related services through 2 October 2026, presented in approximate chronological order.
- Jev (TypeSafe AI, 15 September) :A commercial System One model that helped bring widespread attention to the category. Available through a hosted API, with strong early momentum among developers.
- Decision 1.0 family (vLLM Semantic Router team, 22 September) :Six open-weight models ranging from 0.6B to 9B parameters (Kai, Lex, Eos, Sol, Nox, Lux). Designed for routing, condition evaluation, and action selection.
- GLiNER2.5-Decide (Fastino Labs, 24 September) :Compact open-weight encoder, described by Fastino as a 340-million-parameter model, optimized for schema-defined decisions and local CPU inference.
- d1 (Liquid AI, late September) :Hosted decision model available through Liquid’s API as
d1:free. Liquid AI reported advantages over Jev on its Decision Index evaluation, multilingual tasks, and prompt-injection tests; these should be read as launch claims rather than universal performance guarantees. - OpenAI Decisions API (DevDay 2026, late September) : A Luna-powered service that answers developer-defined questions with finite, predefined answers, using text or image context. Announced in limited preview, with broader release planned.
- GLiDE (Fastino Labs, 30 September) : Hosted decision model that makes an initial assessment and allocates additional reasoning when a choice is uncertain. It shows that structured decision interfaces need not be limited to single-pass inference. Its launch benchmark results were reported by Fastino using the official Decision Index scorer.
- Databricks ai_decide (30 September) :Native AI function, launched in Beta, that applies a decision model directly to governed data inside the Databricks platform for routing, tagging, and agent evaluation. It is a managed platform function rather than a separately released model family.
- Strands Decider 2B (Amazon Web Services, 1 October) :An open-source release with an Apache 2.0 project license, built on Qwen3.5-2B. Its developers report competitive accuracy and calibration for its size class and median latency of approximately 115 milliseconds on an RTX 3090. The launch article’s latency graph used v18; the released reference checkpoint was v19. Training data and scripts were also released.
- Clef and Clef-flash (Cloudflare, 1 October) : Open models (27B and 9B) released under Apache 2.0, with image-input support. Cloudflare reported leading results on the Decision Index at launch. Hosted on Workers AI with a Jev-compatible interface.
Additional open efforts such as the Kev family, OpenDecider, and community projects like Kodiak further expanded the ecosystem. By early October, the growing number of checkpoints and families made a single model count difficult to interpret without specifying what was being counted.
How Many Decision Models Are Open-Sourced?
Several prominent releases identified above provide downloadable weights and permissive licensing. Strands Decider 2B, Clef and Clef-flash, and GLiNER2.5-Decide have Apache 2.0 releases. The Decision 1.0 team licenses its contributions under Apache 2.0 while retaining upstream terms documented in each repository. Weights, and in some cases training recipes, are available on Hugging Face and GitHub. Open weights and a reproducible training release are different levels of openness; each project’s materials and license should be checked separately.
Proprietary or hosted-only options remain important- Jev, OpenAI’s Decisions API, Liquid AI’s d1, Fastino’s GLiDE, and Databricks’ integrated function—but the open-source contingent is already large enough to support local experimentation, fine-tuning, and air-gapped deployment. This early mix of downloadable models and hosted services leaves room for both community experimentation and commercial development, although the category’s longer-term direction remains uncertain.
Who Benefits from Decision Models?
The primary beneficiaries are teams building agentic systems. In practice, an agent may make many small, repeated judgments: which tool to call next, whether a user request requires escalation, how to classify an incoming ticket, or whether a generated action complies with policy. Inserting a full language model for every such step inflates latency and cost. A suitable decision model can reduce both for bounded tasks while providing probabilities or confidence scores that an application can use to trigger human review when uncertainty is high.
Enterprises in customer support, content moderation, fraud detection, and workflow automation may gain particular value, subject to domain-specific evaluation. Support platforms can use these models to route tickets against predefined categories. Security teams can apply rapid policy checks before an agent executes sensitive operations. E-commerce systems can decide whether a product recommendation warrants further personalization. Because some smaller open models run on a laptop or modest GPU, smaller organizations and individual developers can experiment without continuous hosted-inference charges, although hardware and operating costs remain.
Regulated industries may also benefit from more consistent decision records. The constrained output space and explicit scores can make audit logs easier to analyze than free-form responses. For example, when a model estimates a probability of 0.92 that a claim requires specialist review, the input, model version, threshold, and subsequent human decision can be logged with clarity. That record does not, by itself, explain the model’s reasoning or establish regulatory compliance.
Practical Heads-Up for Practitioners
Several practical considerations deserve attention. First, decision models are not replacements for language models; they are complementary components. Many single-pass decision models are less suited to complex reasoning, long-horizon planning, and natural-language explanation. An effective architecture may combine both: a decision model for rapid gates and a language model for deeper work.
Second, calibration quality varies. Early evaluations show meaningful differences in how well stated probabilities match observed outcomes, including across different question types and tasks. Teams should evaluate models on their own data distributions rather than relying solely on public leaderboards.
Third, the rapid proliferation of models means interface compatibility is still evolving. Many adopt a Jev-style request format, which eases switching, yet subtle differences in question types and multimodal support remain. Standardization efforts will likely intensify in the coming months.
Finally, local open-source models introduce new operational responsibilities-model versioning, hardware provisioning, and monitoring for drift—while potentially reducing provider dependence and the need to transmit inputs to external inference services. Those benefits depend on the full deployment, including telemetry and any other hosted components.
The Future of Decision Models in Agentic AI
The speed with which decision models have appeared suggests that the industry has recognized a genuine missing layer in agent stacks. One direction already being explored is tighter integration with reinforcement learning and organization-specific feedback, allowing teams to specialize a base decision model on proprietary decision logs. Cloudflare’s Clef announcement, for example, included a reinforcement-learning fine-tuning initiative. Another is the development of hierarchical decision systems in which small, fast models handle routine choices and escalate only the ambiguous cases to larger models or humans.
There is also room for richer multimodal decision models that incorporate real-time sensor data, video frames, or structured enterprise databases without intermediate text conversion. As agents move into physical environments and complex business processes, the ability to decide quickly under partial observability may become increasingly valuable, provided accuracy and latency are validated for the operating conditions.
From a broader perspective, the rise of decision models may encourage a healthier division of labour in AI systems. Instead of asking a single model to be both eloquent and decisive, developers can assign each capability to the architecture best suited for it. That separation of concerns could produce more reliable, auditable, and cost-effective agentic applications when paired with appropriate evaluation and workflow controls.
Amazon’s decision to release Strands Decider 2B as fully open source, complete with training materials, reinforces this collaborative momentum. When a major cloud provider contributes a competitive model to the public commons, it broadens the opportunities for experimentation and practical integration, without establishing that every model or use case is ready for production. For developers and organizations building the next generation of autonomous systems, an expanding set of tools for faster, potentially cheaper, and more structured decision-making is becoming available. The task ahead is to integrate them thoughtfully into production workflows.

