Our Research

Building general-purpose multimodal simulators of the world.

We believe models that use video as their main input/output modality, when supplemented by other modalities like text and audio, will form the next paradigm of computing.
Research from Runway
July 19, 2026
Runway Characters: Real-Time Expressive AI Characters from a Single Image
by Michail Doukas, Taras Khakhulin, Kathleen Lewis, Yining Shi, Michael Tarasiou, Jimei Yang + 28 authors
Runway Characters transforms a single reference image of any style, from photorealistic human to cartoon mascot, to a real-time conversational video agent. The system produces audio-synchronized facial animation — including lip-sync, gaze dynamics, head movement and secondary motion — from the conditioned audio and the reference image at 24fps in HD resolution. Runway Characters generates the frames autoregressively and has an effective 37 milliseconds of model time per frame. The server-side turn-around from when the user stops speaking to when the character starts responding is 1.75 seconds. The real-time conversational characters represent a novel step towards real-time simulation of human conversational presence....
September 24, 2025
Autoregressive-to-Diffusion Vision Language Models
by Marianne Arriola, Naveen Venkat, Jonathan Granskog, Anastasis Germanidis
We develop a state-of-the-art diffusion vision language model, Autoregressive-to-Diffusion (A2D), by adapting an existing autoregressive vision language model for parallel diffusion decoding. Our approach makes it easy to unlock the speed-quality trade-off of diffusion language models without training from scratch, by leveraging existing pretrained autoregressive models....
June 2, 2025
Dual-Process Image Generation
by Grace Luo, Jonathan Granskog, Aleksander Hołyński, Trevor Darrell
Prior methods for controlling image generation are limited in their ability to be taught new tasks. In contrast, vision-language models, or VLMs, can learn tasks in-context and produce the correct outputs for a given input. We propose a dual-process distillation scheme that allows feed-forward image generators to learn new tasks from deliberative VLMs. Our scheme uses a VLM to rate the generated images and backpropagates this gradient to update the weights of the image generator. Our general framework enables a wide variety of new control tasks through the same text-and-image based interface. We showcase a handful of applications of this technique for different types of control signals, such as commonsense inferences and visual prompts. With our method, users can implement multimodal controls for properties such as color palette, line weight, horizon position, and relative depth within a...
RNA Sessions AI and art research talks
RNA Sessions
An ongoing series of talks about frontier research in AI and art, hosted by Runway.
Learn more