Yandex has replaced its multi-stage recommendation stack with a single generative model called Sona, running a live A/B test on smart speakers that swapped out more than 15 candidate generators and ranking layers.
In this article
The test covered one week of traffic. The new system served a single transformer model instead of the previous pipeline. It handled candidate generation and ranking in one go. The experiment measured performance against the production control.
The old way and the new way
Traditional systems split the decision process across separately trained models. A candidate generator feeds a pre-ranker, which then feeds a heavy ranker. Each stage optimises its own objective. The final ranker only sees what upstream stages allow through.
Yandex’s previous stack on the Yandex Music surface consumed hundreds of engineered features. These included signals from Argus, the company’s earlier recommender transformer. Sona puts candidate generation and ranking around one shared user representation.
The encoder reads the listener’s history once per request. A decoder generates candidates. A Ranking Module scores them against the same encoder states. No component uses hand-engineered features. Inputs are logged event fields and learned Semantic IDs.
On Yandex smart speakers, playback can begin without the user first selecting an artist, genre, or mood. The research team describes this as a pure-recommendation setting.
How the architecture functions
Semantic tokenizer
Every track becomes a tuple of three discrete codes. A frozen multimodal LLM reads the mel-spectrogram of the first 90 seconds along with title, artists, and tags. It runs in prefill-only mode. A four-layer refinement transformer then aligns those features with listening behaviour, using InfoNCE on collaborative track pairs. Residual K-means quantises the result into three codebooks of 32,000 entries each. This beat a CLMR audio baseline: Recall@1000 rose from 0.8111 to 0.8524.
Encoder with history compression
Sona attends to 8,192 past events. Full attention over that length is expensive, so the encoder spends depth unevenly. The recent 2,048 events get a seven-layer self-attention stack. Older events pass through cross-attention and one full-history layer only. The paper reports this keeps most of the quality of full attention at about half the inference cost.
Decoder and Ranking Module
A two-layer decoder emits Semantic ID tuples through constrained beam search with width 1,024. A catalog trie blocks invalid prefixes. Each tuple expands to every track sharing it. The Ranking Module, which consists of four cross-attention layers, then scores those tracks against the shared encoder memory.
Training: A teacher that never ships
The Ranking Module learns from a frozen Teacher Ranker. The teacher is a 0.6B-parameter transformer, also without hand-engineered features. It is trained on a year of engagement events in two stages: next-item-prediction pre-training, then multi-head ranking fine-tuning. Removing pre-training dropped weighted pair accuracy from 0.6215 to 0.6153.
The team calls its distillation method Rollout Distillation. During training, the current decoder generates beam candidates. The teacher scores them, together with logged impressions. The Ranking Module regresses onto those scores with mean absolute error. The joint loss is L = L_NTP + L_rollout + L_impression. Both losses update the shared encoder. At serving time, the teacher is removed.
Training stays online. Events aggregate into sessions over a 15-minute window, feed a GPU trainer, and new weights reach serving every 10 minutes. End-to-end latency is 45 minutes at the median and 60 minutes at p99. Serving runs on NVIDIA Triton Inference Server with CUDA graphs and reaches 41% model FLOPs utilisation.
Results: Online A/B test on live traffic
The final experiment ran for seven days on 15% of randomly selected users per split. Every change below is statistically significant and relative to the production control:
- Active Users (primary metric): +4.53%
- Total Listening Time: +6.30%
- Likes: +11.42%
- “Repeat” Commands: +17.99%
- Deeply Engaged Users: +7.37%
These gains stack on top of improvements retained from earlier deployments. On Active Users, Sona’s uplift is 2.35x the +1.93% increment Argus previously delivered on this surface.
Sona vs OneRec vs HSTU: Feature comparison
Sona is not the first end-to-end generative recommender in production. Kuaishou’s OneRec already serves a single encoder-decoder model. Meta’s HSTU Generative Recommenders reframed recommendation as sequential transduction over user actions in 2024. What Sona combines is a full cascade replacement, no hand-engineered features, and a distilled ranker, validated online.
| Feature | Sona (Yandex) | OneRec (Kuaishou) | HSTU GR (Meta) |
| Domain | Music streaming | Short video | Large internet platform, multiple surfaces |
| One served model replaces the cascade | Yes, in A/B test | Yes, about 25% of total QPS | No, reported as a new architecture for recommendation models |
| User inputs | Logged event fields only, no hand-engineered features | “Multi-scale feature engineering” pathways, including uid, age, gender | User action sequences (sequential transduction) |
| Item output | 3-level Semantic IDs, 3 x 32,000 | 3-level Semantic IDs via RQ-Kmeans | Item IDs |
| Ranking signal | Produced from frozen 0.6B Teacher Ranker | RL with P-Score reward model (ECPO) | HSTU ranking model |
| Reinforcement learning | No, fully supervised | Yes (ECPO) | Not reported |
| Scale reported | 8,192-event history, 0.6B teacher | 10x FLOPs of prior ranking model | 1.5 trillion parameters |
| Reported online gain | +4.53% Active Users, +6.30% listening time, +11.42% likes | +0.54% and +1.24% App Stay Time | +12.4% in online A/B tests |
| Public code or weights | No | Not in the report | Yes, GitHub |
Sources: Sona, OneRec, HSTU. Online gains come from different platforms and metrics, so they are not directly comparable.
What it means
For users of smart speakers, the change is that the system now decides what to play without requiring prior selection. For the engineering teams, the shift removes the need to maintain separate models for generation and ranking. It reduces the complexity of feature engineering and allows the model to learn directly from raw event logs.



