Decision AI Models Explained: TypeSafe Jev vs Fastino GLiDE, GLiNER2.5-Decide and Open-Source Competitors

What are Decision AI Models?

Decision AI models are a new class of model that returns a decision, not a paragraph. You send text with typed questions. The model returns choices, scores or yes/no probabilities your code can branch on directly.

The category went mainstream when TypeSafe AI launched Jev after 2 years in stealth. TypeSafe calls it a ‘System One model,’ after Daniel Kahneman’s fast, intuitive System 1 thinking.

Within 3 weeks, Fastino Labs shipped 2 rival models, and open-source developers published several Jev-style reproductions. This article covers how the category works, where it fits, and how the options compare.

How Jev Works

Jev accepts a ‘state’ (a string, array or set of name-value pairs) and one or more questions. According to the TypeSafe docs, it supports 3 primitives:

  • Choice: pick one option from a list, with probabilities and confidence. TypeSafe says Jev supports up to 255 options.
  • Score: rate the state on a rubric of ordered levels, with probabilities and confidence.
  • Noul: a 0 to 1 probability that a statement is true. The name is short for Bernoulli.

Every question is evaluated in parallel and in isolation against the same state. Adding questions barely changes response time. Because Jev never generates strings, TypeSafe says it cannot return a type error.

Under the hood, TypeSafe describes a new architecture, a parallel sampler and a training method called Reinforcement Learning for Calibrated Decisions (RLCD). RLHF optimizes for human preference. RLCD optimizes for calibrated probabilities, where higher confidence should mean higher accuracy.

Pricing is the key thing to watch. Jev costs $0.042 per million input tokens, and output is free. OpenRouter lists a 32K context window. TypeSafe reports end-to-end responses in 70 to 500 milliseconds.

Interactive Explainer

Benchmarks: What TypeSafe’s Workflow Evals Show

TypeSafe built workflow evals across 4 tasks: security incidents, agent trace observability, invoice processing and customer service. Reference labels come from averaging GPT-6 Astra and Claude Fable 5.1 at high thinking.

  • Jev: 67.8% mean accuracy, $0.0004 per case, 0.4 seconds.
  • Claude Sonnet 5 (same workflow): 67.8%, $0.1174 per case, 78.1 seconds.
  • Best comparison model (OpenAI “sol”): 74.1%, $0.0836 per case, 23.3 seconds.

Jev matched Sonnet 5 on accuracy at a fraction of the cost and latency. It still trails the top frontier configuration by 6.3 points. Per task, Jev scored 76.0% on customer service but only 61.8% on invoice processing.

Use Cases: Where Decision Models Fit

The rule of thumb is simple. If your code needs a bounded answer it will branch on, a decision model is a candidate. If a human needs to read the output, use an LLM.

Agent control flow

  • Choosing the next tool or subagent: Vercel lists this as a primary use.
  • Continue, retry, ask the user, or stop: A single Choice question replaces a fragile JSON-parsing step.
  • Model routing: Send easy requests to a cheap model and hard ones to a frontier model. Fastino lists routing by destination, complexity or escalation level.

Classification and triage

  • Support ticket routing, email triage and intent detection: These make up much of Fastino’s 17-dataset benchmark.
  • Spam detection, label suggestions and prioritization: Simon Willison calls these natural classification fits.
  • Security alert triage: TypeSafe’s simplest published workflow decides whether to close an alert, pass it to an analyst, or contain it.

Verification and safety

  • Guardrails and jailbreak detection: TypeSafe pitches Jev for scoring prompts, reasoning traces and outputs.
  • Joint safety checks: GLiNER2.5-Decide can decode “safety” and “harm type” together so the answers never contradict.
  • Verified cascades: OpenRouter’s Jev guide describes drafting with a cheap model, checking with Jev, and escalating only on failure.

Evaluation and observability

  • LLM-as-a-judge replacement: Arize and Langfuse now run Jev evaluators on traces.
  • CI/CD gates: Buddy lets pipelines score, classify or gate runs with a Jev action.

Search and data processing

  • Reranking: Score 100 BM25 candidates for relevance in one parallel call.
  • Map-reduce over large datasets: TypeSafe pitches turning bulk data into features at low cost.
  • Context pruning: Fastino lists choosing which context to keep before an LLM call.

Real-time applications

  • Games and simulations: TypeSafe demoed Jev playing Doom and Wikiracing.
  • Latency-critical UX: Sub-second decisions make AI usable inside interactive flows.

When not to use a decision model

  • You need generated text, summaries or explanations.
  • The task needs exact arithmetic, counting or date math. TypeSafe’s Jev 1.13 jaggedness guide flags all 3.
  • The decision affects people’s livelihoods, such as hiring. Willison warns that hidden bias is hard to inspect.

Case Studies: Where Decision Models Are Already Working

1. Vercel AI Gateway adoption

Vercel reports that Jev became the fastest-adopted model in AI Gateway history. Within 24 hours, nearly 13% of paid teams were using it. That was 2x the GPT-5.6 family’s share and more than 6x Fable 5.1’s.

2. Search reranking

Simon Willison fetched 100 candidates with BM25, then had Jev score each for relevance. This pattern replaces an expensive LLM reranker with cheap, parallel Score questions.

3. LLM evaluation pipelines

Arize and Langfuse both shipped Jev-as-a-judge evaluators. Langfuse labels the feature “decision-model evaluators” so other models can be added later. Its Jev support is still marked experimental.

4. 6G edge network orchestration

A new arXiv paper used Jev to interpret service contracts at the network edge. At matched correctness, Jev cut median decision latency by 22.4% versus DeepSeek and 61.9% versus Gemini.

5. Real-time game agents

TypeSafe’s Doom demo ran Jev at 10 queries per second. The team estimated the cost at roughly $7 per hour.

Competitor Landscape: 7 Decision Models Compared

The table below covers Jev, its closest commercial rival, and the leading open-weight options. All specs come from each project’s own release page or model card.

ModelDeveloperLicense / accessSize and architectureQuestion typesReported latencyRuns locally
Jev 1.13TypeSafe AIClosed; hosted API at $0.042/M input, output freeNot disclosed; new architecture, parallel sampler, RLCDChoice, Score, Noul (up to 255 options)70 to 500 msNo
GLiDEFastino LabsClosed; Fastino API, 40K contextNot disclosed; adaptive thinking on uncertain casesSelected action, confidence and option probabilitiesNot disclosedNo
GLiNER2.5-DecideFastino LabsApache 2.0 open weights; also on Fastino API340M DeBERTa-v3-large encoderTyped decisions with joint constraints, plus spans and relations38.3 ms p50 (V100), 167.3 ms (48-vCPU CPU)Yes, CPU or GPU
LayaConvai InnovationsApache 2.0 open weights421M ModernBERT-large (English); 322M mmBERT-base (multilingual)Choice, Score, Noul; 100+ languages~33 ms per question on GPUYes
JevK5Independent developerApache 2.0 open weights4B, Qwen3.5-4B with merged LoRAProbability for every option in one passNot statedYes, GPU
OpenJevTheo LeeMITFrozen Qwen3.5-4B, reads option logitsTyped options in one pass5.21x faster than autoregressive JSONYes, consumer GPU
kev-0.5bJared PalmerApache 2.0 (base under Qwen license)LoRA plus readout head on Qwen2.5-0.5B; Jev-compatible APINoul, Choice (2 to 255), Score~160 ms for 6 questions (Apple M5)Yes, laptop
General LLM + structured outputManyMostly closed; input plus output tokensAutoregressive decoderAnything, as generated textTypically secondsDepends

Head-to-head numbers (each vendor’s own benchmark)

These scores come from different test suites. Do not compare numbers across rows.

Benchmark (reported by)Results
Workflow evals (TypeSafe)Jev 67.8%, Sonnet 5 67.8%, best LLM 74.1%
Decision Index 0.2.1 (Fastino)GLiDE 64.81, Jev 57.91
CLadder accuracy (Fastino)GLiDE 88.7%, Jev 72.6%
CRUXEval accuracy (Fastino)GLiDE 92.6%, Jev 73.0%
Fast Decisions, 17 datasets (Fastino)GLiNER2.5-Decide 60.1%, JevK5 57.5%, SemIf 56.4%, GLiFormer 49.0%, Laya 46.6%
102-row TypeSafe eval subset (OpenJev)Jev 0.883, OpenJev 0.845 balanced accuracy

Two caveats matter here. Fastino chose both the tests and the opponents for its GLiDE and GLiNER2.5-Decide claims. Its Fast Decisions comparison also uses JevK5, an open reproduction, not TypeSafe’s Jev.

How to choose

  • Pick Jev for a managed API with the broadest ecosystem support today.
  • Pick GLiDE to test harder, reasoning-heavy decisions, and verify Fastino’s claims on your own data.
  • Pick GLiNER2.5-Decide for air-gapped deployment, fine-tuning, or answers that must obey cross-question rules.
  • Pick Laya for multilingual decisions or very low per-question latency.
  • Pick kev, OpenJev or JevK5 to experiment locally or to avoid vendor lock-in.

Why Decision Models are the Latest Trend in AI

The idea is not new. Classifiers and rerankers have made decisions for years. Decision Transformer (2021) framed reinforcement learning as sequence modeling. DeepMind’s Gato (2022) showed one generalist model acting across many tasks. A 2023 survey even called ‘large decision models’ the next step.

What changed in 2026 is the packaging. 4 forces drove the shift:

  • Agents need cheap judgments: Routing, tool selection and guardrails happen thousands of times per workflow. Paying frontier prices for each one does not scale.
  • Calibration enables automation: A confidence score lets code act alone when sure and escalate when not.
  • The ecosystem moved fast: Within weeks, Jev landed on Vercel, OpenRouter, Arize, Langfuse and Buddy.
  • Competition arrived immediately: Fastino shipped GLiNER2.5-Decide on September 24 and GLiDE on September 30. Open reproductions such as kev, OpenJev and JevK5 appeared in parallel.

Key Takeaways

  • Decision AI models return typed answers with probabilities, not generated text.
  • Jev costs $0.042 per million input tokens, with free output.
  • In TypeSafe’s evals, Jev matched Sonnet 5 accuracy at 0.4 seconds per case.
  • Fastino’s GLiDE and GLiNER2.5-Decide lead a fast-growing field of rivals.
  • Use decision models for routing and scoring, and LLMs for reasoning and writing.

FAQ

  • Is a decision model the same as an LLM? The answer is No, A decision model reads text like an LLM but outputs probabilities over predefined answers. It cannot write, summarize or explain.
  • When should I use a decision model instead of an LLM? Use one for high-volume, bounded judgments: classification, routing, triage, guardrails, reranking and evals. Keep LLMs for open-ended reasoning and generation.
  • Can I run a decision model on my own hardware? Jev and GLiDE are API-only. GLiNER2.5-Decide, Laya, JevK5, OpenJev and kev all have open weights you can run locally.
  • What is the main competitor to Jev? Fastino Labs is the most direct commercial rival. It ships GLiDE as a hosted API and GLiNER2.5-Decide as open weights.

The post Decision AI Models Explained: TypeSafe Jev vs Fastino GLiDE, GLiNER2.5-Decide and Open-Source Competitors appeared first on MarkTechPost.



from MarkTechPost https://ift.tt/rtXEAUB
via IFTTT

Comments

Popular posts from this blog

Nanbeige4-3B-Thinking: How a 23T Token Pipeline Pushes 3B Models Past 30B Class Reasoning

Microsoft AI Proposes BitNet Distillation (BitDistill): A Lightweight Pipeline that Delivers up to 10x Memory Savings and about 2.65x CPU Speedup

Technical Deep Dive: Automating LLM Agent Mastery for Any MCP Server with MCP- RL and ART