Yandex Introduces Sona: A Single Generative Recommender That Replaces Entire Recommendation Cascade

Yandex has replaced its multi-stage recommendation stack with a single generative model called Sona, running a live A/B test on smart speakers…

By Vane October 5, 2026 4 min read

Yandex has replaced its multi-stage recommendation stack with a single generative model called Sona, running a live A/B test on smart speakers that swapped out more than 15 candidate generators and ranking layers.

The test covered one week of traffic. The new system served a single transformer model instead of the previous pipeline. It handled candidate generation and ranking in one go. The experiment measured performance against the production control.

The old way and the new way

Traditional systems split the decision process across separately trained models. A candidate generator feeds a pre-ranker, which then feeds a heavy ranker. Each stage optimises its own objective. The final ranker only sees what upstream stages allow through.

Yandex’s previous stack on the Yandex Music surface consumed hundreds of engineered features. These included signals from Argus, the company’s earlier recommender transformer. Sona puts candidate generation and ranking around one shared user representation.

The encoder reads the listener’s history once per request. A decoder generates candidates. A Ranking Module scores them against the same encoder states. No component uses hand-engineered features. Inputs are logged event fields and learned Semantic IDs.

On Yandex smart speakers, playback can begin without the user first selecting an artist, genre, or mood. The research team describes this as a pure-recommendation setting.

How the architecture functions

Semantic tokenizer

Every track becomes a tuple of three discrete codes. A frozen multimodal LLM reads the mel-spectrogram of the first 90 seconds along with title, artists, and tags. It runs in prefill-only mode. A four-layer refinement transformer then aligns those features with listening behaviour, using InfoNCE on collaborative track pairs. Residual K-means quantises the result into three codebooks of 32,000 entries each. This beat a CLMR audio baseline: Recall@1000 rose from 0.8111 to 0.8524.

Encoder with history compression

Sona attends to 8,192 past events. Full attention over that length is expensive, so the encoder spends depth unevenly. The recent 2,048 events get a seven-layer self-attention stack. Older events pass through cross-attention and one full-history layer only. The paper reports this keeps most of the quality of full attention at about half the inference cost.

Decoder and Ranking Module

A two-layer decoder emits Semantic ID tuples through constrained beam search with width 1,024. A catalog trie blocks invalid prefixes. Each tuple expands to every track sharing it. The Ranking Module, which consists of four cross-attention layers, then scores those tracks against the shared encoder memory.

Training: A teacher that never ships

The Ranking Module learns from a frozen Teacher Ranker. The teacher is a 0.6B-parameter transformer, also without hand-engineered features. It is trained on a year of engagement events in two stages: next-item-prediction pre-training, then multi-head ranking fine-tuning. Removing pre-training dropped weighted pair accuracy from 0.6215 to 0.6153.

The team calls its distillation method Rollout Distillation. During training, the current decoder generates beam candidates. The teacher scores them, together with logged impressions. The Ranking Module regresses onto those scores with mean absolute error. The joint loss is L = L_NTP + L_rollout + L_impression. Both losses update the shared encoder. At serving time, the teacher is removed.

Training stays online. Events aggregate into sessions over a 15-minute window, feed a GPU trainer, and new weights reach serving every 10 minutes. End-to-end latency is 45 minutes at the median and 60 minutes at p99. Serving runs on NVIDIA Triton Inference Server with CUDA graphs and reaches 41% model FLOPs utilisation.

Results: Online A/B test on live traffic

The final experiment ran for seven days on 15% of randomly selected users per split. Every change below is statistically significant and relative to the production control:

  • Active Users (primary metric): +4.53%
  • Total Listening Time: +6.30%
  • Likes: +11.42%
  • “Repeat” Commands: +17.99%
  • Deeply Engaged Users: +7.37%

These gains stack on top of improvements retained from earlier deployments. On Active Users, Sona’s uplift is 2.35x the +1.93% increment Argus previously delivered on this surface.

Sona vs OneRec vs HSTU: Feature comparison

Sona is not the first end-to-end generative recommender in production. Kuaishou’s OneRec already serves a single encoder-decoder model. Meta’s HSTU Generative Recommenders reframed recommendation as sequential transduction over user actions in 2024. What Sona combines is a full cascade replacement, no hand-engineered features, and a distilled ranker, validated online.

FeatureSona (Yandex)OneRec (Kuaishou)HSTU GR (Meta)
DomainMusic streamingShort videoLarge internet platform, multiple surfaces
One served model replaces the cascadeYes, in A/B testYes, about 25% of total QPSNo, reported as a new architecture for recommendation models
User inputsLogged event fields only, no hand-engineered features“Multi-scale feature engineering” pathways, including uid, age, genderUser action sequences (sequential transduction)
Item output3-level Semantic IDs, 3 x 32,0003-level Semantic IDs via RQ-KmeansItem IDs
Ranking signalProduced from frozen 0.6B Teacher RankerRL with P-Score reward model (ECPO)HSTU ranking model
Reinforcement learningNo, fully supervisedYes (ECPO)Not reported
Scale reported8,192-event history, 0.6B teacher10x FLOPs of prior ranking model1.5 trillion parameters
Reported online gain+4.53% Active Users, +6.30% listening time, +11.42% likes+0.54% and +1.24% App Stay Time+12.4% in online A/B tests
Public code or weightsNoNot in the reportYes, GitHub

Sources: Sona, OneRec, HSTU. Online gains come from different platforms and metrics, so they are not directly comparable.

What it means

For users of smart speakers, the change is that the system now decides what to play without requiring prior selection. For the engineering teams, the shift removes the need to maintain separate models for generation and ranking. It reduces the complexity of feature engineering and allows the model to learn directly from raw event logs.

Scroll to Top