Runway has outlined a plan to shift AI video generation from a batch process to a live stream controlled in real time.
In this article
Current models require users to input a prompt, wait for the system to render a clip, and then start over if the output is unsatisfactory. Runway notes that creators frequently lose the most time on generation and revision cycles. The company aims to reduce the time to the first frame and stream video as the user dictates it.
Runway first detailed this method in March alongside Runway Characters. That tool relies on GWM-1, the company’s first General World Model, introduced in December 2025. GWM-1 builds on Gen-4.5, generates video frame by frame, and accepts inputs like camera movements, robot commands, or audio. A few weeks ago, Runway demonstrated Solaris, a system using Gen-4.5 to generate user interfaces frame by frame. It responds to clicks or voice input.
Real-time generation could cut waiting and GPU costs
Runway argues that real-time generation closes the gap between an idea and its execution. Instant feedback would mean users spend most of their time actively steering the video rather than waiting.
The company also points to lower costs. Faster models require less GPU time, making them more cost-efficient. According to Runway, the cost per output at a given quality level determines which applications make economic sense. Instant generation would lower that threshold, making previously unprofitable applications viable.
Small visual errors can grow into major distortions
A text model can correct itself mid-sentence, but a video model builds each frame on the previous one. This allows small errors to compound into major distortions over time. Runway describes this as the central problem with LLM-based approaches. It addresses the issue by training the model on its own outputs rather than only error-free inputs. This teaches the system to correct its own deviations instead of amplifying them.
Startup Decart used a similar approach for its real-time model MirageLSD, deliberately exposing it to flawed or distorted images during training. Google Deepmind says its world model Genie 3 keeps interactive worlds consistent for several minutes at 24 frames per second in 720p.
According to Runway, real-time generation shifts the compute load from training to use. The model must produce each frame fast enough to keep up with playback while running on hardware shared by several sessions at once.
World models would train robots and robotaxis
Runway sees interactive applications as the biggest long-term use case for AI-generated media. Education, gaming, and robotics need video that responds as quickly as the person watching it, the company says. Evaluating how robots or autonomous vehicles act in the world also calls for environments that generate in real time and respond instantly to edge cases. Runway previously introduced GWM Robotics, a variant of GWM-1 that generates synthetic training data for robots.
Waymo is taking a similar approach with the Waymo World Model, which is based on Genie 3 and adapted for road traffic. It lets Waymo simulate situations its fleet has never observed, such as an encounter with an elephant, a tornado, or a flooded residential neighborhood. According to Waymo, the Waymo Driver travels billions of miles in virtual worlds before encountering scenarios on public roads.
In March, Runway also showed a research preview of a real-time model developed with Nvidia at the chipmaker’s GTC conference. It runs on the Vera Rubin platform and is designed to deliver the first frame in under 100 milliseconds. Runway has not announced a timeline for availability.




