Posts

IBM Brings Bob to Self-Hosted and Air-Gapped Environments: Agentic Software Development Without Moving Your Code

IBM has made a self-hosted deployment option for IBM Bob generally available. Bob is IBM’s agentic software development platform. It covers the full lifecycle: understanding code, planning work, executing changes and validating results. The new option lets enterprises run Bob on premises, in private or sovereign clouds, and in air-gapped networks. Is it deployable today? Yes. The self-hosted option is generally available to enterprise customers. You must source, license and host a supported model yourself. IBM has not published pricing and routes buyers to a demo request . What IBM Shipped at GA The self-hosted deployment includes: Bob running in customer-managed enterprise environments. Core Bob capabilities: the IDE experience, BobShell, parallel tool calling, the agent harness, skills and modes. Self-hosted, air-gapped and hybrid model configurations. Bring-your-own-license (BYOL) for eligible models on IBM’s supported list. Integratio...

Microsoft AI Releases MAI-Transcribe-2-Streaming: #1 Real-Time Speech-to-Text Model on Artificial Analysis

Microsoft AI has released MAI-Transcribe-2-Streaming , its first streaming speech-to-text (STT) model. It launched on October 1, 2026, alongside 2 text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Artificial Analysis ranks it #1 of 38 models for final and first partial transcript accuracy. The model targets voice agents, live captions and dictation, where latency decides the experience. What Microsoft Shipped MAI-Transcribe-2-Streaming is the real-time sibling of the batch MAI-Transcribe-2, released in September. It transcribes 60 languages with automatic, continuous language detection. Audio streams in continuously, and text streams back while the speaker is still talking. The model emits its first hypotheses, called partials, just over 100ms after receiving audio. It revises those partials as context arrives, then commits a stable final transcript. An agent can therefore start reasoning or calling tools mid-sentence. Microsoft team states its internal tests ...

Decision AI Models Explained: TypeSafe Jev vs Fastino GLiDE, GLiNER2.5-Decide and Open-Source Competitors

What are Decision AI Models? Decision AI models are a new class of model that returns a decision, not a paragraph. You send text with typed questions. The model returns choices, scores or yes/no probabilities your code can branch on directly. The category went mainstream when TypeSafe AI launched Jev after 2 years in stealth. TypeSafe calls it a ‘System One model,’ after Daniel Kahneman’s fast, intuitive System 1 thinking. Within 3 weeks, Fastino Labs shipped 2 rival models, and open-source developers published several Jev-style reproductions. This article covers how the category works, where it fits, and how the options compare. How Jev Works Jev accepts a ‘state’ (a string, array or set of name-value pairs) and one or more questions. According to the TypeSafe docs , it supports 3 primitives: Choice: pick one option from a list, with probabilities and confidence. TypeSafe says Jev supports up to 255 options. Score: rate the...

Datalab Introduces OmniExtractBench to Fix Bias and Opacity in Extraction Benchmarks

Image
Datalab has released OmniExtractBench , an open benchmark for structured document extraction. It tests how accurately a system fills a JSON schema from a PDF. The benchmark pools 620 documents from 4 existing benchmarks. One deterministic scorer grades all of them and explains each decision. The release lands while extraction vendors publish their own leaderboards. Datalab argues those leaderboards are hard to compare or audit. OmniExtractBench is its attempt at a shared yardstick. Is it deployable? Yes , the scorer installs from PyPI as omni-extract-bench (v0.1.7, Python 3.11+, SciPy only) under Apache 2.0. Rerunning vendors requires your own API keys and paid credits. What is OmniExtractBench? OmniExtractBench is a structured extraction benchmark built by Datalab. Each task gives a system a PDF and a JSON schema. The system returns JSON, which is scored value by value against a gold file. The code is on GitHub , and the data is on Hugging Face under CC BY 4.0. ...