Alibaba has released Wan3.0, a beta video generation model that creates clips up to thirty seconds long from text, images, or documents. This duration doubles the output of the previous Wan2.5 version. The system accepts up to ten images, five video clips, and five audio files within a single prompt. It also processes web pages and PDFs to convert static data into moving visuals.
The model processes text, images, video, and audio simultaneously to maintain consistency in faces and spatial layouts. Alibaba claims this reduces the visual drift common in earlier generative tools. Pricing varies by resolution and speed, with a Standard tier costing six dollars for a thirty-second clip at 1080p and a faster Prime tier at eight dollars. The service is available via the wan.video website, Alibaba Cloud Model Studio, or an API on Qwen Cloud.
* Standard pricing for a 30-second 1080p clip is $6.00
* Prime pricing for the same clip is $8.40
* The model accepts inputs including PowerPoint presentations




