Posts

NVIDIA Launches Open Agent Safety Platform: OpenShell Sandboxes Agents on Vera CPUs While Sentry on BlueField-4 Quarantines Them in Milliseconds

NVIDIA has launched the NVIDIA Open Agent Safety Platform , an open software platform and reference system design for AI agent security. It pairs the OpenShell secure runtime with NVIDIA Sentry, an out-of-band watchdog on BlueField-4 DPUs. The core idea is simple. Safety controls should not live inside the agent they are meant to control. Today, with over 100 industry partners, we introduced the NVIDIA Open Agent Safety Platform, bringing together OpenShell and Sentry. Artificial intelligence is extraordinary technology that will advance discovery, productivity, security, health, and prosperity for generations to… pic.twitter.com/dReAxwpRUn — Jensen Huang (@JensenHuang) September 28, 2026 Is it deployable today? Yes for OpenShell. It is Apache 2.0, installs on Linux, macOS (Apple Silicon) or Windows WSL 2, and its repo still labels it alpha. Why NVIDIA Moved Enforcement Below the Agent The NVIDIA technical report cites recent reports from several frontier lab...

Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens

Fireworks AI has released Ember-1 , a specialized model from Fireworks Research built by post-training Moonshot AI’s open-weight Kimi K3 . Ember-1 learns to produce shorter reasoning traces while keeping task accuracy. This is different from lowering the reasoning effort setting at inference time. According to the Fireworks release post , Ember-1 delivers Kimi K3’s quality with about 40% fewer tokens. Is it deployable? Yes, but only through the Fireworks serverless API as a Research Preview. Fireworks has not released Ember-1’s weights, training code, or exact training algorithms, so self-hosting is not an option today. The Problem: Reasoning Models Think Too Much Fireworks team reports that reasoning models like Kimi K3 sometimes spend more than 90% of generated tokens on internal reasoning. That cost compounds in multi-turn agentic workloads. Each turn replays prior reasoning back to the model. Context grows roughly quadratically with the number of turn...

Google Research Introduces an AI Video Co-Director: 4 Agentic Frameworks for Coherent, Minutes-Long Video Generation

Google Research has introduced an AI video co-director for long-form video generation. The suite of 4 agentic frameworks turns short clips into coherent, minutes-long stories. It targets identity drift and cascading errors, the 2 failures that break most multi-shot AI video pipelines today. Why Long AI Videos Fall Apart Diffusion models render high-fidelity clips in seconds. Stitching those clips into a story is harder. Most agentic pipelines chain modules with independent, handcrafted prompts. That causes semantic drift , where attire or scenery shifts between shots. It also causes cascading failures , where one bad upstream asset corrupts every later shot. Google team frames this as a credit assignment problem. A broken final video is hard to trace back to the prompt that caused it. How the AI Video Co-Director Works The system sits on top of Gemini and Veo. It is model-agnostic, so the same layer can drive other generators. Outputs inherit SynthID watermarki...

A Coding Guide to Google Research’s MSEB: Writing Sound Encoders to the Benchmark Contract and Scoring Them Across Classification, Clustering, Retrieval and Segmentation

In this tutorial, we work with MSEB , the Massive Sound Embedding Benchmark from Google Research, and approach it from the perspective of what a leaderboard number actually means: the evaluator surface. We install the package and map its three layers, then write two deliberately different encoders against the framework’s own abstract base class: one that measures loudness over time and one that measures timbre, and encode a small synthetic corpus we generate in the notebook so nothing has to be downloaded. We drive the classification, clustering, retrieval, and segmentation evaluators over those embeddings, call the metric functions directly to see what each one rewards, and finish by assembling the TaskMetadata a real submission carries. The result is a comparison in which the two encoders trade places depending on which evaluator is asked, which is the argument for a multi-task benchmark made in numbers rather than in prose. Copy Code Copied Use a different Browser import o...