Big Tech

Microsoft Launches Decision-1 Model, Bypassing OpenAI to Match TypeSafe's Pricing

Microsoft released Decision-1 on Alibaba's Qwen foundation just days after OpenAI's API debut, pricing it identically to TypeSafe's Jev while testing the model across its own product lines.

5 min read
Microsoft skipped OpenAI’s decision model and built its own on Alibaba’s Qwen

On Friday, Microsoft unveiled Microsoft-Decision-1 through its Foundry platform, arriving three days after OpenAI made its Decisions API available to all developers in public beta. Rather than relying on technology from its OpenAI partnership, Microsoft constructed the decision model by post-training Alibaba's Qwen3.5-9B, though the company indicated plans to rebase it later on both its own MAI models and OpenAI's offerings.

The pricing mirrors TypeSafe's Jev exactly: $0.042 per million input tokens with no output charges. This alignment came on the same day TypeSafe announced a Series A funding round of $870 million, valuing the company at $7.5 billion and led by a16z.

Microsoft trails the startup that pioneered this category by approximately three and a half weeks. In that interval, OpenAI, Upstage, Perplexity, Cloudflare, and AWS have each introduced their own decision models.

TypeSafe disclosed that nearly 30% of Fortune 500 companies have experimented with Jev, though the company has not publicly identified them. TypeSafe CEO Diogo Almeida stated on X that "29.4% of the Fortune 500 showed up" within the three weeks following the product's launch.

https://x.com/CompleteSkeptic/status/2108594987177021737?ref_src=twsrc%5Etfw

These same enterprises represent Microsoft's Azure customer base, making TypeSafe's rapid adoption difficult for Microsoft to overlook.

Microsoft's first customer is itself

Microsoft Chairman and CEO Satya Nadella unveiled the model on Friday via X, stating, "We're already testing it across Microsoft," with four internal teams providing concrete examples.

https://x.com/satyanadella/status/2108627923888754862?ref_src=twsrc%5Etfw

Xbox Research deployed it to organize more than 10,000 player feedback submissions. The Copilot team used it to evaluate chat and agent outputs. On-call engineers leveraged it to gather context during active incidents. Microsoft Discovery employed it to rate experiments before an agent reconsiders its approach.

Microsoft's internal results demonstrate the model operates more than 14 times faster than GPT-6 Sol while costing substantially less for Xbox operations, and achieves 46 times greater consistency within Discovery.

The model may serve a broader strategic purpose. On Wednesday, Microsoft announced that GitHub Copilot will soon determine whether to execute tasks locally or route them to cloud-based models, though the company has not detailed what Copilot transmits to the cloud.

Model routing ranks among the use cases Microsoft identifies for Decision-1, yet the company has not clarified whether this model will manage those routing decisions for Copilot. Given routing requirements spanning Copilot, GitHub, and Xbox, Microsoft holds incentive to manage these operations internally.

Decision models became cheap

Cognition's vice president of engineering, Jared Palmer, spent approximately $95 in Modal H100 compute time to adapt his open-source Kev models to Qwen3.5, and Cloudflare constructed Clef-flash using the identical Qwen3.5-9B foundation that Microsoft selected.

Jev's pricing has become the market standard: Palmer lists Kev-4B on OpenRouter at $0.042 per million input tokens, Perplexity charges $0.02, and OpenAI's rate exceeds Jev's by more than double. At these rates, earnings from individual decision transactions remain minimal, and Microsoft's principal opportunity lies in retaining agent traffic—including the generative operations surrounding each decision—within Foundry. The OpenRouter listing could additionally attract developers not currently on Azure into that ecosystem.

Achint Srivastava, vice president of software engineering in Microsoft's Office of the CTO, presented Decision-1 as a mechanism for incorporating decision-making into current applications, agents and workflows "in a secure, trusted environment."

Text only, no weights

This reveals Decision-1's limitations relative to competing offerings. Its Foundry documentation indicates it processes up to 32,768 tokens of text and outputs JSON, but lacks image support—a capability that OpenAI's Decisions API and Cloudflare's Clef, which incorporates a vision encoder, both provide.

Cloudflare released Clef under Apache 2.0 licensing, whereas Microsoft has not yet announced open weights. AWS, Upstage, and Ollama have adopted TypeSafe's System One API, establishing a shared interface across the category, but Microsoft has not confirmed whether Decision-1 achieves full compatibility, despite its Foundry sample code referencing a /systemone endpoint.

Calibration under adversarial pressure

Microsoft states that Decision-1's probabilities are calibrated, meaning a 90% prediction should prove accurate roughly nine times out of 10 across representative scenarios, but research on Jev demonstrates how substantially a confident score can deteriorate when inputs are crafted to deceive. Microsoft's own Foundry documentation recommends that customers verify calibration using their own datasets.

The JevOut preprint, authored by USC computer science researcher Zixiang Xu and colleagues, revealed that brief, plausible additions to context reversed 312 of Jev's 508 initially correct determinations. In 229 instances, Jev assigned at least 70% probability to the incorrect answer. Three additional scoring approaches exhibited flip rates between 64.9% and 73.2% under the same evaluation methodology.

Microsoft subjected Decision-1 to eight categories of perturbations, including reordered options and rephrased descriptions, which altered 1.3% of its answers on average. JevOut did not evaluate Decision-1, so the model's performance against comparable attacks remains uncertain. Microsoft has additionally not confirmed full System One API support that competitors have embraced, creating ambiguity regarding interoperability and the appropriate confidence level to assign to its probability estimates.

Source: The New Stack · Reporting supplemented by The Silicon Ledger staff.