System One vs. System Two AI Models: Inside Microsoft Decision-1, TypeSafe AI Jev, and the Latency Collapse
Why burning generative LLMs on boolean routing creates latency traps: exploring Daniel Kahneman's cognitive duality in AI, Microsoft's new Decision-1 model, TypeSafe AI Jev, and OpenAI Decisions API.

Series: ← The Always-On Agent Showdown: Architectural, Security, and Autonomy Comparison Across Meta Muse, OpenAI Dots, Google Gemini Spark, and Perplexity Computer (Previous) | Claude Haiku 5.5: Adaptive Thinking, Effort Scaling, and Frontier Agentic Work at Scale (Next) →
Summary
For the past four years, software engineering has suffered from a fundamental category error: we have used slow, multi-billion-parameter generative language models to answer simple binary and classification questions. Whenever an autonomous agent checks if an email requires escalation, determines whether a SQL query is read-only, or routes a ticket between billing and engineering, developers build 2,000-token prompts and wait seconds for an autoregressive LLM to emit a JSON object.
On October 9, 2026, Microsoft CEO Satya Nadella unveiled Microsoft-Decision-1, an open-weight post-trained decision model built on Qwen3.5-9B and served via Microsoft Foundry. Designed strictly to evaluate fixed candidate sets and return calibrated probability scores, Decision-1 runs 35 times faster than GPT-6 Sol and 4.5 times faster than Quyet-1.0-Large, priced at a microscopic $0.042 per 1M input tokens with zero output token fees. Alongside TypeSafe AI’s Jev and OpenAI’s Decisions API, Microsoft-Decision-1 crystallizes the System One vs. System Two bifurcation in AI: decoupling rapid, non-autoregressive reflex decisions from heavy, deliberative chain-of-thought generation.
Official Announcement & Model Card Comparison
| Attribute | Microsoft-Decision-1 | TypeSafe AI Jev | OpenAI Decisions API | Frontier LLMs (GPT-6 / Sonnet) |
|---|---|---|---|---|
| Provider | Microsoft Foundry | TypeSafe AI | OpenAI API | OpenAI / Anthropic |
| Release Date | October 9, 2026 | September 15, 2026 | September 29, 2026 | 2024 – 2026 |
| Foundational Base | Qwen3.5-9B (Post-Trained) | Custom Discriminative Core | GPT-6 Luna Specialized Head | Frontier Dense / MoE Transformers |
| Execution Paradigm | Single-pass forward propagation | Non-autoregressive classification | Single-pass score matrix | Autoregressive token-by-token decode |
| Output Type | Calibrated probability distribution | Typed Primitives (bool, enum, score) | Structured decision objects | Freeform text, JSON string |
| End-to-End Latency | ~50ms – 150ms | ~70ms – 180ms | ~120ms – 200ms | 1,500ms – 8,000ms+ |
| Input Token Pricing | $0.042 / 1M tokens | $0.042 / 1M tokens | $0.10 / 1M tokens | $2.00 – $15.00 / 1M tokens |
| Output Token Fee | $0 (Free / Unmetered) | $0 (Free / Unmetered) | $0 (Free / Unmetered) | $10.00 – $75.00 / 1M tokens |
| Structural Hallucination | 0% (Closed candidate bounds) | 0% (Enforced schema bounds) | 0% (Typed enum space) | Variable (JSON syntax corruption risk) |
Prior Reading Material
To trace how decision models and inference economics transform modern agent runtimes, explore our foundational deep dives:
- TypeSafe AI Jev: The Non-Autoregressive ‘System One’ Decision Engine and the Latency Collapse — The original deep dive on Jev, BAML schema integration, and non-autoregressive forward passes.
- OpenAI DevDay 2026: Dots Always-On Agents, GPT-6.1 Sol’s 80% Price Cut, the Decisions API — Plus Salesforce Koa and NVFP4 KV Cache — OpenAI’s typed Decisions API and ambient cloud agent infrastructure.
- The Always-On Agent Showdown: Architectural, Security, and Autonomy Comparison Across Meta Muse, OpenAI Dots, Google Gemini Spark, and Perplexity Computer — Comparative analysis of persistent virtualization layers and multi-model routing.
- Claude Haiku 5.5: Adaptive Thinking, Effort Scaling, and Frontier Agentic Work at Scale — The fast-tier reasoning frontier and subagent swarm orchestration.
1. The Sledgehammer Anti-Pattern: Why Chatbots Fail at Control Flow
Every software system is fundamentally a network of decision gates:
- “Does this incoming webhook require immediate paging or silent logging?”
- “Which subagent should handle this repository bug: Frontend, Backend, or Database?”
- “Does the user’s latest prompt attempt an unauthorized prompt injection or jailbreak?”
In standard deterministic code, these decisions are simple if/else checks or switch statements. But because user inputs and enterprise data are messy, unstructured natural language, developers turned to generative LLMs to make these choices.
The consequences have been catastrophic for agent latency and reliability:
flowchart TD
subgraph LegacyFlow["The Autoregressive Generative Bottleneck"]
direction TB
E["Inbound Event: Complex Customer Ticket"] --> P["Construct 2,500-Token Prompt + JSON Schema + Few-Shot Examples"]
P --> K["Frontier LLM Prefill Phase (GPU KV Cache Allocation)"]
K --> D["Autoregressive Decode Phase: Token 1 '{' ... Token 2 'category' ... Token 3 'billing'"]
D --> W["Wait 1,800ms - 3,500ms for Generation to Complete"]
W --> J["Regex / Pydantic JSON Parser Checks Output"]
J --> F{"Did JSON Parse Successfully?"}
F -- "Syntax Error / Hallucination" --> P
F -- "Valid Output" --> R["Route to Billing Handler"]
end
style LegacyFlow fill:#1a0f14,stroke:#ef4444,stroke-width:2px,color:#ffffff
style E fill:#27272a,stroke:#71717a,stroke-width:1px,color:#ffffff
style D fill:#450a0a,stroke:#f87171,stroke-width:1.5px,color:#ffffff
style W fill:#3f1414,stroke:#ef4444,stroke-width:1.5px,color:#ffffff
style F fill:#7f1d1d,stroke:#fca5a5,stroke-width:1.5px,color:#ffffff
style R fill:#064e3b,stroke:#10b981,stroke-width:1.5px,color:#ffffff
To decide between BILLING, TECH_SUPPORT, or SALES, an autoregressive model loads hundreds of gigabytes of weights, initializes key-value attention caches, and executes dozens of sequential forward passes. It spits out {, ", c, a, t, e, g, o, r, y, ", :, and so forth.
This approach introduces three structural flaws into agent pipelines:
- The Latency Floor: Autoregressive decoding is bounded by memory bandwidth. Even the fastest speculative decoding engines rarely exceed 250 tokens per second, guaranteeing a 1,000ms+ delay before control flow can branch.
- The JSON Tax: Developers pay input token charges for thousands of schema definition tokens, and pay expensive output token rates for boilerplates, brackets, and quotes.
- Parse Brittleness: When a generative model randomly prepends “Sure! Here is the JSON:” or hallucinates an unquoted key, the pipeline crashes.
2. Daniel Kahneman’s Paradigm in AI: Fast Reflexes vs. Slow Thought
In cognitive psychology, Daniel Kahneman documented in Thinking, Fast and Slow that human intelligence operates via two interconnected computational modes:
- System One (Fast, Instinctive, Reflexive): Operates automatically and quickly, with little or no effort and no sense of voluntary control. When a driver spots a red traffic light, their foot hits the brake in 200 milliseconds. They do not calculate stopping distances with kinetic friction equations; their brain triggers an automatic conditioned reflex.
- System Two (Slow, Deliberate, Analytical): Allocates attention to effortful mental operations, including complex computations, architectural designs, and formal logic. When that same driver plans a multi-state route navigating weather disruptions and toll tariffs, they engage System Two.
flowchart TD
subgraph HumanCognition["Dual-Process Human Cognition (Kahneman)"]
direction TB
Stimulus["Sensory Input (Visual, Auditory, Tactile)"] --> Router{"Cognitive Triage"}
Router -- "Immediate Danger / Pattern Match" --> S1["System One: Instinctive Reflex (~200ms)<br/>No conscious deliberation • Pattern recognition"]
Router -- "Complex Problem / Novel State" --> S2["System Two: Deliberative Reasoning (Seconds/Minutes)<br/>Attention-heavy • Multi-step logical chains"]
end
style HumanCognition fill:#0b0f19,stroke:#6366f1,stroke-width:2px,color:#ffffff
style Router fill:#1e1b4b,stroke:#818cf8,stroke-width:1.5px,color:#ffffff
style S1 fill:#064e3b,stroke:#10b981,stroke-width:1.5px,color:#ffffff
style S2 fill:#3b0764,stroke:#c084fc,stroke-width:1.5px,color:#ffffff
Modern artificial intelligence has heavily over-indexed on System Two. We trained foundation models to reason using extended thinking tokens, multi-turn tool loops, and massive chain-of-thought rollouts.
However, in autonomous agents, over 80% of all operations are pure System One tasks:
- Filtering safe vs. dangerous inputs (Guardrails)
- Classifying intent and sentiment (Triage)
- Choosing between 5 predefined tools (Tool Selection)
- Deciding whether a loop should terminate (Stopping Gates)
Using an autoregressive System Two engine for these tasks is the computational equivalent of solving a partial differential equation just to decide whether to step over a puddle.
3. Microsoft Decision-1 & The System One Frontier
On October 9, 2026, Microsoft officially legitimized System One decision engines with the release of Microsoft-Decision-1.
Announced by Satya Nadella and the Microsoft AI team, Decision-1 represents a fundamental architectural departure from standard generative models:
Figure 1: Architectural comparison between System One specialized decision models (Microsoft Decision-1, TypeSafe AI Jev, OpenAI Decisions API) and System Two generative frontier reasoners.
How Microsoft-Decision-1 Operates:
- Base Foundation: Rather than training from scratch, Microsoft post-trained the open-weight Qwen3.5-9B foundation model.
- Fixed Option Scoring: Instead of prompting the model to generate text, Decision-1 is fed a context prompt alongside a discrete set of candidate choices:
{ "model": "microsoft-decision-1", "prompt": "Evaluate patient vitals: Heart Rate 145 bpm, BP 85/50, SpO2 91%. Classify urgency level.", "options": ["ROUTINE", "URGENT", "EMERGENCY_STAT"] } - Calibrated Probabilities: In a single forward pass, the model evaluates the semantic representations of the input against the candidate options, returning calibrated probability distributions:
{ "decision": "EMERGENCY_STAT", "confidence": 0.9842, "probabilities": { "ROUTINE": 0.0004, "URGENT": 0.0154, "EMERGENCY_STAT": 0.9842 }, "latency_ms": 64 } - Performance Velocity: Across 36 internal evaluation benchmarks, Microsoft reported that Decision-1 is 35 times faster than OpenAI GPT-6 Sol and 4.5 times faster than Quyet-1.0-Large.
- Radical Tokenomics: Because the model emits a single categorical vector rather than generating dozens of output tokens, output tokens are completely unmetered ($0 output cost). The input cost is fixed at $0.042 per 1,000,000 tokens—making 100,000 classification decisions cost barely 40 cents.
4. Architectural Comparison: Decision-1 vs. Jev vs. OpenAI Decisions API
The emergence of dedicated decision engines creates a compelling three-way comparison across the leading architectures:
flowchart TD
M["<b>Microsoft-Decision-1</b><br/>Qwen3.5-9B Post-Trained Base • Microsoft Foundry<br/>$0.042 / 1M Input Tokens • $0 Output Fee<br/>Calibrated Multi-Class Probability Distributions"]
style M fill:#042f2e,stroke:#14b8a6,stroke-width:2px,color:#ffffff
flowchart TD
J["<b>TypeSafe AI Jev</b><br/>Proprietary Discriminative Core • BAML Native Schema<br/>$0.042 / 1M Input Tokens • $0 Output Fee<br/>Strictly Typed Boolean and Enum Primitives"]
style J fill:#1e1b4b,stroke:#818cf8,stroke-width:2px,color:#ffffff
flowchart TD
O["<b>OpenAI Decisions API</b><br/>Specialized GPT-6 Luna Classification Head • OpenAI Platform<br/>$0.10 / 1M Input Tokens • $0 Output Fee<br/>Typed Categorical Predicates and Bounded Scores"]
style O fill:#064e3b,stroke:#10b981,stroke-width:2px,color:#ffffff
1. Microsoft-Decision-1 (Probabilistic Reasoning on Open-Weights)
- Strengths: Backed by Microsoft’s enterprise infrastructure on Azure Foundry; provides mathematically rigorous, calibrated probability scores across arbitrary option lists; leverages the broad world knowledge of Qwen3.5-9B.
- Best For: High-stakes triage (medical, incident response, IT security alert scoring) where human supervisors require confidence metrics before executing automated remediations.
2. TypeSafe AI Jev (The Machine-Native System One Pioneer)
- Strengths: Deeply coupled with BAML (BoundaryML) schemas; custom non-autoregressive architecture delivering sub-70ms cold latencies; zero probability of schema violations or syntax corruption.
- Best For: Microservice-level routing, API gateway interceptors, edge deployment workers (Fly.io, Cloudflare Workers), and tight deterministic state machines.
3. OpenAI Decisions API (The Native Ecosystem Endpoint)
- Strengths: Natively integrated into the OpenAI API suite alongside GPT-6 and Dots; shared client SDKs and consolidated billing; zero pipeline bridging required for developers already built on OpenAI.
- Best For: High-speed guardrails and tool choice pruning inside existing OpenAI Agents API workflows.
5. The Dual-Cognition Agent Architecture in Production
How do engineering teams combine System One and System Two models into production systems?
The optimal pattern is a Hierarchical Reflex-Deliberation Pipeline:
flowchart TD
subgraph DualArchitecture["Production Dual-Cognition Pipeline"]
direction TB
In["Inbound User Request / API Event"] --> S1_Guard["System One: Microsoft Decision-1 (Guardrail Check)<br/>Latency: 50ms • Cost: $0.00004"]
S1_Guard -- "Unsafe / Prompt Injection" --> Reject["Instant Rejection & Alert Dispatch"]
S1_Guard -- "Safe Input" --> S1_Route["System One: Decision-1 / Jev (Intent Router)<br/>Latency: 60ms • Cost: $0.00004"]
S1_Route -- "Static Lookup / Simple Query" --> Cache["Direct Database / Cache Retrieval"]
S1_Route -- "Complex Planning / Code Generation" --> S2_Engine["System Two: Claude Sonnet 5.5 / GPT-6 Sol<br/>Extended Chain-of-Thought Reasoning<br/>Latency: 2,500ms • Cost: $0.005"]
S2_Engine --> S1_Verify["System One: Decision-1 (Output Safety & Schema Verification)<br/>Latency: 55ms • Cost: $0.00004"]
S1_Verify -- "Passed" --> Out["Deliver Verified Response to User"]
S1_Verify -- "Anomaly Detected" --> S2_Engine
end
style DualArchitecture fill:#0b0f19,stroke:#a855f7,stroke-width:2px,color:#ffffff
style S1_Guard fill:#042f2e,stroke:#14b8a6,stroke-width:1.5px,color:#ffffff
style S1_Route fill:#042f2e,stroke:#14b8a6,stroke-width:1.5px,color:#ffffff
style S1_Verify fill:#042f2e,stroke:#14b8a6,stroke-width:1.5px,color:#ffffff
style S2_Engine fill:#312e81,stroke:#a5b4fc,stroke-width:1.5px,color:#ffffff
style Out fill:#064e3b,stroke:#10b981,stroke-width:2px,color:#ffffff
style Reject fill:#450a0a,stroke:#f87171,stroke-width:2px,color:#ffffff
The Economic & Performance Impact
In a traditional architecture, every stage of this pipeline—guardrail checking, routing, execution, and verification—was routed to a single frontier model. A single request incurred 8,000ms of cumulative latency and cost $0.02.
In the Dual-Cognition architecture:
- 90% of requests are resolved or routed by System One models in under 150ms for a fraction of a cent.
- The heavy System Two model is invoked only when genuine deliberative reasoning is required.
- Output verification happens in a 55ms reflex pass rather than a second full LLM evaluation.
The result is a 10x to 20x latency reduction across the system, an 85% reduction in overall inference expenditure, and total elimination of JSON parsing failures.
6. The Verdict: Intelligence Gets Divided to Multiply
The simultaneous emergence of Microsoft-Decision-1, TypeSafe AI Jev, and OpenAI’s Decisions API signals the end of the “one model to rule them all” era in software engineering.
Just as computer hardware evolved from general-purpose CPUs to specialized GPUs, TPUs, and neural engines, AI model architectures are dividing along biological lines:
- System One delivers the instantaneous, deterministic reflexes that keep systems fast, secure, and affordable.
- System Two delivers the deep, contemplative reasoning that makes autonomous problem-solving possible.
By pairing Microsoft Decision-1’s 50ms calibrated probability scoring with frontier reasoning models, developers finally have the architectural tools to build agents that are as nimble as they are intelligent.
