NVIDIA PivotOPD Teaches Multi-Turn AI Agents to Recover From Pivotal Mistakes
NVIDIA researchers, with Princeton University and the University of Maryland, have introduced PivotOPD , an on-policy distillation method for multi-turn LLM agents. PivotOPD on-policy distillation trains an agent to avoid its most damaging early mistake, and to recover when it happens anyway. Against 13 baselines, it posts the best average on ALFWorld, WebShop and Search-based QA for Qwen3-1.7B and Qwen3-8B students. The takeaway: recovery is learnable, and standard OPD rarely teaches it. TL;DR Size: A training method, not a model. Tested on Qwen3-1.7B and Qwen3-8B students, plus a Nemotron-3.5-SFT student on SWE-Bench Verified. Runs on: Trained on NVIDIA H100 nodes. Adds 0 inference cost, so the trained agent runs wherever its base model runs. Performance: First on all 8 per-benchmark averages against 13 baselines, across 3 seeds. Best: Recovers from 72.7% of replayed pivotal mistakes, vs 20.3% for standard OPD. Worst: 55.9% on ALFWorld “Look”...
