Alibaba shipped a 2.4 trillion parameter AI model in October 2026. This was the culmination of a seven-year strategy that began in April 2023 with a 7 billion parameter chatbot.
In this article
Chapter 1 — 2023: a thousand questions
Alibaba Cloud began distributing invitation codes for its model, Tongyi Qianwen, on April 7, 2023. The name references the philosopher Mencius. Four days later, then-CEO Daniel Zhang presented the system at the Alibaba Cloud Summit in Beijing. The company intended to integrate it into every business, specifically within DingTalk and the Tmall Genie voice assistant.
The major shift occurred in August. On August 3, Alibaba open-sourced Qwen-7B and Qwen-7B-Chat. This release targeted Meta’s Llama 2. The 7 billion parameter version was pretrained on over 2.2 trillion tokens with a 2,048-token context window. The license permitted commercial use for up to 100 million monthly users.
Qwen-VL, the first vision-language branch, launched in late August 2023. On September 13, Tongyi Qianwen opened to the general public, indicating Chinese regulatory approval. The team also published the Qwen Technical Report on arXiv that same month.
Alibaba released its 72B and 1.8B models for download around December 1, 2023. The family now ranged from laptop-sized to frontier-sized open weights.
Chapter 2 — 2024: the open-weight machine
2024 saw Qwen become a default choice for developers. The team shipped three generations in eight months.
Qwen1.5 (February): On February 5, 2024, the team released Qwen1.5. Framed as a beta of Qwen2, it covered 0.5B to 72B dense models with a stable 32K context at every size. A 14B MoE model with 2.7B active parameters followed, alongside a 110B dense model, the family’s first above 100B.
Qwen2 (June): The team announced Qwen2 in early June 2024 across five sizes, from 0.5B to 72B. This lineup included Qwen2-57B-A14B, its first open mixture-of-experts flagship. Training data added 27 languages beyond English and Chinese. The 7B and 72B instruct models handled up to 128K tokens. The licensing shift mattered most: every size except 72B moved to Apache 2.0.
Specialists arrived over the summer. Qwen2-Math launched in August, along with Qwen2-Audio. Qwen2-VL followed at the end of August, able to analyze videos over 20 minutes long.
Qwen2.5 (September): At the Apsara Conference on September 19, 2024, Alibaba released over 100 open-source models at once. The Qwen2.5 Technical Report states pretraining data grew from 7 trillion to 18 trillion tokens. Sizes ran 0.5B to 72B, with 128K context and 8K-token generation.
Adoption was already real. Alibaba said Qwen models had passed 40 million downloads and inspired over 50,000 derivative models on Hugging Face.
Coding and reasoning (November–December): Qwen2.5-Coder shipped its full family on November 11, 2024. Then came the first reasoning model. QwQ-32B-Preview arrived in late November under Apache 2.0, as an open challenger to OpenAI’s o1. QVQ-72B-Preview, an experimental visual reasoning model, closed the year on December 24.
Chapter 3 — Early 2025: answering DeepSeek
January 2025 belonged to DeepSeek-R1. Qwen’s response came in weeks, not months.
On January 26, the team shipped Qwen2.5-VL in 3B, 7B and 72B sizes. TechCrunch noted it could control PCs and phones. Three days later, on the first day of Lunar New Year, Alibaba launched Qwen2.5-Max, a large-scale MoE model. Reuters reported Alibaba’s claim that it surpassed DeepSeek-V3.
The bigger statement came on March 6. QwQ-32B, built on Qwen2.5-32B and trained with reinforcement learning, shipped under Apache 2.0. Qwen claimed performance comparable to DeepSeek-R1, a 671B model. VentureBeat put the hardware gap at about 24 GB of VRAM versus over 1,500 GB. Alibaba’s Hong Kong shares rose more than 7% that day.
Multimodality kept pace. Qwen2.5-Omni-7B arrived on March 26 under Apache 2.0. It took text, images, audio and video as input, and answered in text or speech.
Chapter 4 — 2025: Qwen3 and the trillion-parameter line
Qwen3 (April): On April 29, 2025 (Beijing time), the team released Qwen3: 6 dense models from 0.6B to 32B and 2 MoE models. The flagship was Qwen3-235B-A22B, with 22B active parameters. Every model shipped under Apache 2.0.
The key feature was hybrid thinking. One model could reason step by step or answer instantly, toggled by the user. The Qwen3 Technical Report lists pretraining on about 36 trillion tokens. Language coverage jumped from 29 to 119 languages and dialects. TechCrunch called it a family of “hybrid” reasoning models.
The 2507 refresh (July): On July 21, 2025, Qwen split hybrid thinking back apart. Qwen3-235B-A22B-Instruct-2507 shipped as a non-thinking model with a 262K native context. Separate Thinking-2507 checkpoints followed.
Agents and images (July–August): Qwen3-Coder-480B-A35B launched on July 22, 2025 for agentic coding. Alibaba paired it with Qwen Code, an open-source terminal coding agent. On August 4, the team open-sourced Qwen-Image, a 20B MMDiT model built for text rendering inside images.
An editing variant, Qwen-Image-Edit, was open-sourced on August 18, 2025.
The September sprint: Qwen3-Max-Preview, the first Qwen model over 1 trillion parameters, appeared on September 5. Unlike past flagships, it shipped API-only.
Days later came Qwen3-Next-80B-A3B, an architecture bet. It mixed Gated DeltaNet linear attention with gated attention in a 3:1 layout. The model card claims 10% of Qwen3-32B’s training cost and 10x inference throughput past 32K tokens.
On September 22, Qwen3-Omni and Qwen3-VL followed. At Apsara on September 24, Alibaba formally launched Qwen3-Max. CEO Eddie Wu said spending would exceed the RMB 380 billion (US$53 billion) three-year AI and cloud plan.
By year’s end, Qwen was national infrastructure for others. In November 2025, AI Singapore said its Sea-Lion model would switch from Llama to Qwen as its base.
What it means
For people building tools, the shift is from scarcity to abundance. The barrier is no longer access to weights, but the engineering to run them. The 2.4T parameter model sits at the top of this stack, designed for tasks requiring massive context and complex reasoning without human intervention.



