MiniMax has released the H3 video model weights, marking the first time an open-source system tops the Artificial Analysis video ranking. The 33-billion-parameter model processes text, images, video, and audio together to generate four to 15-second clips with stereo sound. It allows a single prompt to include up to nine reference images, three video clips, and three audio clips. The system ranks first in Video Editing, second in Text-to-Video, and third in Image-to-Video.
Two components remain closed, specifically the 2K resolution module and the H3-Context-IR translation tool. Running H3 locally in ComfyUI tops out at 768p, requiring users to handle context preparation using MiniMax’s published guides. The open weights do allow fine-tuning on custom footage, characters, or a specific visual style. One catch on the license side is that commercial use is only permitted for companies making under $20 million in revenue. ByteDance released its closed Seedance 2.5 the same day, which generates 30-second clips with built-in audio.
* Commercial licensing restricts firms with annual revenue above $20 million
* Local inference is limited to 768p resolution without the 2K module
* Context preparation must be handled manually by the user




