Our Research

Building general-purpose multimodal simulators of the world.

We believe models that use video as their main input/output modality, when supplemented by other modalities like text and audio, will form the next paradigm of computing.
Research from Runway
September 3, 2026
GWM Worlds 2
by Runway
GWM Worlds 2 turns high-fidelity video and audio generation into real-time interactive simulation. You define a world — its environment, subjects, visual style, physical rules and ambience — then steer it with free-form text actions and continuous camera motion, generating continuous 720p video at 24 fps and audio at 48,000 Hz that responds to your inputs as you explore. A research preview extending GWM Worlds with generated audio and rich subject and scene control....
September 1, 2026
Solaris: Towards Interfaces That Are Generated, Not Coded
by Yuval Alaluf, Omri Avrahami, Anastasis Germanidis, Michal Geyer, Kfir Goldberg, Guy Bukchin Leshem, Alejandro Matamala Ortiz, Elad Richardson + 13 authors
Digital interfaces are traditionally implemented through intermediate representations such as code, requiring their appearance and behavior to be specified in advance. We introduce Solaris, an interface world model that instead generates an interactive UI directly, frame by frame, in response to user actions. Solaris treats mouse interactions as conditioning signals and autoregressively synthesizes the resulting visual state at interactive speeds. To enable real-time generation while maintaining visual coherence over extended interactions, we combine autoregressive frame generation with few-step distillation and training on the model's own outputs....
July 19, 2026
Runway Characters: Real-Time Expressive AI Characters from a Single Image
by Michail Doukas, Taras Khakhulin, Kathleen Lewis, Yining Shi, Michael Tarasiou, Jimei Yang + 28 authors
Runway Characters transforms a single reference image of any style, from photorealistic human to cartoon mascot, to a real-time conversational video agent. The system produces audio-synchronized facial animation — including lip-sync, gaze dynamics, head movement and secondary motion — from the conditioned audio and the reference image at 24fps in HD resolution. Runway Characters generates the frames autoregressively and has an effective 37 milliseconds of model time per frame. The server-side turn-around from when the user stops speaking to when the character starts responding is 1.75 seconds. The real-time conversational characters represent a novel step towards real-time simulation of human conversational presence....
RNA Sessions AI and art research talks
RNA Sessions
An ongoing series of talks about frontier research in AI and art, hosted by Runway.
Learn more