Runway Engineering
RW Mach: Runway's Engine for Fast Inference and Training
The optimization stack that trains and serves Runway's models, from bytes to frames.
Entries logged
mach/serve
Inference. Real-time generation, weight distribution, capacity control, queueing.
6- EntryDescriptionResult
MACH-010·2026-09-10·Production·Latest
Towards Instant Video Generation
Real-time engine · Inference · Distillation · DMD · Real time
The recipe behind real time: post-train the base model into a causal frame-by-frame generator, then distill each frame to a few denoising steps, off-policy first and on-policy on the student's own rollouts, on a sequence-length curriculum.
ResultFew steps per frameMACH-009·2026-08-31·Production
Introducing Solaris
Real-time engine · Inference · Distillation · Autoregressive · Interfaces
Gen-4.5 adapted into a real-time interface world model: autoregressive frames, denoising distilled to a few steps, then trained on its own outputs so quality holds across long sessions. Orders of magnitude cheaper to run than a standard video diffusion model.
ResultOrders of magnitude cheaperMACH-007·2026-07-02·Production
Borrowing the Night: Reclaiming Idle Inference GPUs for Research
deckard · Compute · Capacity · Queueing theory · Scheduling
by Brannon Dorsey and Matt Kafonek, Platform Engineering
A capacity controller sizes the production fleet against a week of observed demand using queueing theory, lends the overnight trough to research and takes it back before the morning peak. Fewer GPUs in production and shorter queues, at once.
ResultFewer GPUs, shorter queuesMACH-004·2026-05-04·Production
Building Real-Time Video Agent from a Single Image with Runway Characters
Real-time engine · Inference · Real time · KV cache · CUDA graphs
by Runway Research and Engineering Team
A conversational video agent from a single image at 24 fps in HD. Diffusion and VAE decode pipelined across devices, KV-cache eviction, CUDA graphs and fused Triton kernels bring model time to 37 ms per frame.
Result37 ms per frameMACH-003·2026-05-04·Production
60x Faster Cold Starts: Treating Peer GPUs as Weight Servers
NCCLBack · Weights · NCCL · Inference · Rollouts
by Jeevan Farias and Daniel Sammons, Platform Engineering
One worker downloads the weights; every other worker receives them from a loaded peer over NCCL. Cold starts fall from minutes to seconds, saving 347 TB of transfer and 6,500 minutes of inference time per day.
Result60× faster cold startsMACH-001·2021-11-01·Production
Distributing Work: Adventures in queuing
Queue broker · Inference · Queueing · Autoscaling · Serving
by Brannon Dorsey
Replaced push-based HTTP load balancing with a pull-based queue broker and eager workers for multi-second GPU requests. Two weeks of production traffic: p95 latency down 32%, p50 down 41%, failure rate unchanged.
Result−32% p95 latencymach/train
Training. Data decoding, distributed correctness, configuration, scheduling.
4- EntryDescriptionResult
MACH-008·2026-07-10·Open source
AVTensor: A High-Performance Rust Media Decoder for Training Pipelines
AVTensor · Data · Rust · FFmpeg · Dataloading
by Rik Heijdens, Research Engineering
pip install avtensorGitHubSingle-pass Rust demuxer with sample-accurate audio and video alignment, streaming straight from object storage. Dataloading about 30% faster; 1.8 MFU points on production audio-visual training runs.
Result+1.8 MFU ptsMACH-006·2026-05-18·Production
Why Distributed Training Is Hard: DTensor, Correctness and the Costs of Abstraction
DTensor hybrid · Training · DTensor · torch.compile · MFU
by Wei Zhang, Research Engineering
DTensor makes sharded gradients correct by construction, then quietly taxes throughput through redistributions and recompilation storms. The stack that wins: DTensor for correctness, hand-written collectives on the hot paths.
ResultHybrid stack winsMACH-005·2026-05-11·Open source
Stop Writing YAML: Configuring ML Systems with confingy
confingy · Tooling · Configuration · Python · Open source
by Ethan Rosenthal, Research Engineering
pip install confingyGitHubRetired thousands of lines of inherited YAML for typed Python configuration: one decorator gives serialization, lazy instantiation and early validation. Open sourced as confingy.
ResultYAML → 0MACH-002·2026-04-27·Production
No Idle GPUs: Managing Research Compute at Runway
Kueue scheduling · Compute · Kubernetes · Kueue · Scheduling
by Matt Kafonek and Brannon Dorsey, Platform Engineering
Kueue as a Kubernetes admission controller: reserved quotas for critical runs, a shared queue that borrows idle capacity and preemption that returns it when owners need it. Utilization at about 2× industry norms.
Result+20 pts GPU utilization
Start of log. Nothing earlier
- CareersWork on this team
- BlogAll engineering posts
Published from the inference and training engine
Runway AI, Inc.
