Posts

Nace AI Open-Sources Drex 1.5: A 9B Decision Model That Scores Options, Not Text

Nace.AI has open-sourced Drex 1.5 , a 9B decision model for agents and backend workflows. The Drex 1.5 decision model does not write text. It reads a state and typed questions, then returns a probability for every option. Nace reports 58.08 on the public Decision Index 0.3.1 , the top score under 10B parameters. Weights are on Hugging Face, and a hosted version is live on OpenRouter . TL;DR Size: 8.95B parameters (dense), bf16 weights about 18 GB. Context is 16,384 tokens by default, up to 131,072. Runs on: 1 CUDA GPU in bf16 (tested on a 24 GB A10G). A Q8_0 GGUF (about 9.5 GB) runs on Apple silicon and CPU. Performance: 58.08 on Decision Index 0.3.1 (public), within the board’s tie band of Jev 1.13.0 (57.96). Best: 93.4% accuracy on 32K to 128K token documents. Worst: 7.4% per-review F1 on ACOS aspect sentiment, versus 29.5% for Jev. Bottom line: Best: open weights that match a closed model on 1 GPU. Worst: weak on broad knowledge and fin...

OpenAI Decisions API Hits Public Beta With 10x Faster Typed Answers

OpenAI has released the Decisions API in public beta. It turns text and images into typed answers your code can branch on. OpenAI team states the OpenAI Decisions API runs about 10x faster than the Responses API. It targets a common pattern: prompt an LLM, then parse its text into a label. TL;DR Size: GPT-6 Luna parameter count is not disclosed. Its model card lists a 1,050,000-token context window. Runs on: OpenAI-hosted API only, via POST /v1/decisions . No open weights, no self-hosting. Performance: About 10x faster than the Responses API, per OpenAI . Best: $0.10 per 1M input tokens, with no output, cache-read or cache-write charges. Bottom line: Best: fast, typed decisions with probabilities. Worst: 1 model, beta status, no independent evals yet. What is the OpenAI Decisions API? The Decisions API is an OpenAI endpoint that evaluates text, images or both and returns typed answers. It does not generate prose. A request has 3 fields:...

Google Cloud Launches Gemini Agent, One Universal Agent for Enterprise Work

Google Cloud has introduced the Google Cloud Gemini agent , a single agent for enterprise work. The Gemini agent is a cloud-hosted agent from Google Cloud that answers questions, does knowledge work, creates media, and writes and runs code. It does all of this from 1 prompt box and 1 API. For developers, the agent is the product and the model is a routing decision. TL;DR Runs on: Google Cloud (AI Hypercomputer). Reached from web, mobile, desktop, CLI, Workspace, Microsoft 365, Slack, or headless. Best: Bloomberg Media lifted SQL query accuracy by 63% by grounding data agents in Knowledge Catalog. Bottom line: Best: 1 governed agent with cloud memory, sub-agents and hard spend caps. Worst: buyers must evaluate it on customer anecdotes, not reproducible numbers. What is the Google Cloud Gemini agent? It is a delegation layer, not a chatbot. You give it objectives, not instructions. It plans the work, picks skills and tools, connects to company system...

Google Research RRSI Guide: Mastering Self-Improving AI Agents

In this tutorial, we implement RRSI (Regularized Recursive Self-Improvement) , a method that lets an LLM agent rewrite its own harness, prompts, tools, memory, control flow, and sub-agents around a frozen model, without the harness overfitting to the tasks it evolves on. The full RRSI loop drafts edits with Claude Opus on Vertex AI and scores them inside Docker benchmarks, which is not something a free notebook can run. The part of RRSI that actually carries the paper’s idea, the rules that decide which proposed edits to keep, is plain Python, and that is what we drive directly. We install the package from the official repository, walk through its estimator, its calibrated noise band, both branches of its selection algorithm, its annealed edit budget, its deterministic leakage screen, and its edit history, and then plug a simulated agent into RRSI’s own Domain interface. Because we built the simulated environment ourselves, we know the true effect of every edit, which lets ...