Posts

Yandex Introduces Sona: A Single Generative Recommender That Replaces Entire Recommendation Cascade

Most production recommenders are cascades. Candidate generators feed a pre-ranker, which feeds a heavy ranker built on hundreds of engineered features. Yandex’s Sona Technical Report describes a different design. Sona is a generative AI model that brings candidate generation and ranking into a single system, replacing the multiple stages typically used in recommendation pipelines. Yandex tested the model in a seven-day live production experiment on its smart speakers. In an online A/B test, it replaced more than 15 candidate generators, the pre-ranking stage, and the ranking stage with one served transformer. What Problem Does Sona Solve? Cascades split one decision across separately trained models. Each stage optimizes its own objective, and the ranker only sees what upstream stages let through. Yandex’s previous stack on the Yandex Music surface consumed hundreds of features, including signals from Argus , Yandex’s earlier recommender transformer. Sona puts...

The Story of Qwen: Alibaba’s AI Models From 7B to 2.4T

In April 2023, Alibaba Cloud demoed a chatbot whose name roughly means ‘truth from a thousand questions.’ Three and a half years later, its descendant ships open weights with 2.4 trillion parameters. This is the story of how Qwen got there, release by release. Chapter 1 — 2023: a thousand questions Alibaba moved after ChatGPT, but not by much. On April 7, 2023, Alibaba Cloud began handing invitation codes to corporate customers for a model called Tongyi Qianwen. The name draws partly on the philosopher Mencius. Four days later, at the Alibaba Cloud Summit in Beijing, then-CEO Daniel Zhang unveiled it publicly. Alibaba said it would roll the model into every business , starting with DingTalk and the Tmall Genie voice assistant ( China Daily ). The real turn came in August. On August 3, 2023, Alibaba open-sourced Qwen-7B and Qwen-7B-Chat , a direct answer to Meta’s Llama 2. Qwen-7B was pretrained on over 2.2 trillion tokens with a 2,048-tok...

GPT-6 Astra vs GPT-6.1 Sol vs Gemini 4 Argon vs Claude Fable 5.1: Which Frontier Model Fits Which Job

Anthropic, OpenAI and Google DeepMind shipped 4 frontier-class models within 30 days. Claude Fable 5.1 arrived on September 1. GPT-6 Astra followed on September 3. GPT-6.1 Sol and Gemini 4 Argon landed in the last days of September. We covered each launch on its own. This piece puts them side by side. The benchmark scores overlap more than the launch posts suggest. The prices, access rules and cost per task do not. One change frames the lineup. OpenAI cancelled GPT-6.1 Astra on September 28 after it failed internal scope and authorization tests. GPT-6 Astra stays OpenAI’s top model for now. Specs, Pricing and Access Astra and Fable 5.1 share the same $10 input and $50 output list price. Sol and Argon list at one-fifth of that. Argon’s price is introductory and doubles later. Feature GPT-6 Astra GPT-6.1 Sol Gemini 4 Argon Claude Fable 5.1 Developer OpenAI OpenAI Google DeepMind Anthropic Released Sep 3, 2026 Sep 29, 2026 Announced Sep 30, 2026 Sep 1,...

Aleph Alpha Releases Kolibri: A 78.1B Open-Weight English-German MoE Model With Only 3.46B Active Parameters

Aleph Alpha has released Kolibri , an open-weight Mixture-of-Experts (MoE) language model built for German and English. Kolibri has 78.1B total parameters but activates only 3.46B, or 4.4%, per token. It accepts up to 1,048,576 tokens of context, lets users set reasoning effort per request, and ships under the Apache 2.0 license on Hugging Face . The target is sovereign deployment in regulated sectors such as public administration, industry and aerospace. Is it deployable? Yes. The FP8 checkpoint is about 78GB and runs on a single B200, B300 or H200, or on 2 H100 SXM5 GPUs, served through vLLM with dedicated Kolibri reasoning and tool-call parsers. What is Kolibri? Kolibri (Kolibri-1) is a bilingual English-German MoE transformer developed end to end by teams in Germany. According to the technical research report , Aleph Alpha team controlled the full pipeline: data, architecture, training infrastructure, post-training and evaluation. Training ran on infrastructure in Ge...