Posts

NVIDIA PivotOPD Teaches Multi-Turn AI Agents to Recover From Pivotal Mistakes

Image
NVIDIA researchers, with Princeton University and the University of Maryland, have introduced PivotOPD , an on-policy distillation method for multi-turn LLM agents. PivotOPD on-policy distillation trains an agent to avoid its most damaging early mistake, and to recover when it happens anyway. Against 13 baselines, it posts the best average on ALFWorld, WebShop and Search-based QA for Qwen3-1.7B and Qwen3-8B students. The takeaway: recovery is learnable, and standard OPD rarely teaches it. TL;DR Size: A training method, not a model. Tested on Qwen3-1.7B and Qwen3-8B students, plus a Nemotron-3.5-SFT student on SWE-Bench Verified. Runs on: Trained on NVIDIA H100 nodes. Adds 0 inference cost, so the trained agent runs wherever its base model runs. Performance: First on all 8 per-benchmark averages against 13 baselines, across 3 seeds. Best: Recovers from 72.7% of replayed pivotal mistakes, vs 20.3% for standard OPD. Worst: 55.9% on ALFWorld “Look”...

What Happens When a Trusted Model Repo Changes? Unsloth Studio Re-Checks Before It Runs

Image
Beta Lessons Learned After over 500 million downloads, years of requests from the open-source community, and being a top product on Hugging Face, Unsloth launched their beta desktop App, Unsloth Studio. Unsloth which makes it faster, easier, and more affordable to fine-tune and run AI models, including locally on your own hardware. The app centralizes features in one spot with Unsloth Studio so users can now use a dashboard install of manual installation.  Open-source projects rely on other code sources or platforms and in the case of Unsloth as early adopters to local modelling their product combined the freedom of the Hugging Face platform with the fine-tuning capabilities of Unsloth’s various packages.  After Unsloth Studio launched their product and have been updating on a fast scale for an OSS while adapting to quickly shifting safety environments in AI. For example a C ompromised LiteLLM versions 1.82.7 and 1.82.8 appeared on PyPI from a compromised Trivy s...

Anthropic Releases Claude Haiku 5.5: A Small Model With 1M Context Priced at $0.10 per Million Input Tokens

Anthropic has released Claude Haiku 5.5 , its cheapest and fastest small model to date. It targets high-volume work like summaries, compaction, classification and subagent tasks. It keeps a 1M token context window and up to 128K output tokens. Pricing starts at $0.10 per million input tokens and $0.50 per million output tokens. That is 90% below Claude Haiku 4.5 for prompts up to 100K tokens. Is it deployable? Yes, as a hosted API. Haiku 5.5 is generally available on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. What Anthropic Shipped Haiku 5.5 is the first Haiku-class model with an adjustable effort setting. Adaptive thinking is on by default, and the effort parameter defaults to medium . It takes text and images, outputs text, and has a June 2026 knowledge cutoff. Batch jobs support up to 300K output tokens in beta. It is important to note two key things. Non-default temperature , top_p or top_k values return ...

Liquid AI Releases Open-Weight d1-3B and d1-omni-600M: Multimodal Decision Models With Zero Output Tokens

Liquid AI has released Open d1 , two open-weight multimodal models in its d1 decision model family. d1-3B reads text and images. d1-omni-600M reads text with an image, or text with audio. Neither model writes text. Each returns calibrated, typed answers in one forward pass with zero output tokens. The target is real-time decisions on the NVIDIA stack: DGX servers, RTX workstations, and Jetson edge boards. Is it deployable? Yes. Both checkpoints are on Hugging Face, load through Transformers, and have day-one llama.cpp support. The LFM Open License v1.0 allows free commercial use below $10 million in annual revenue. d1-omni-600M is an early research release with no published latency figures. What is a decision model? A generative LLM writes its answer token by token, and your code parses it. A decision model takes a state and a set of named questions. It reads them once and returns a probability for every allowed answer. The Liquid AI define three question types: ...