Anastasis from Runway discusses the evolution from generative art tools to world models that simulate physics, robotics, and real-time interfaces. He reveals how scaling video models using diffusion transformers leads to predictable improvements in understanding physics, and how world models can replace traditional front-end software by generating pixels directly from prompts.
Summarized by Podsumo
Runway pivoted to *world models* in late 2023, seeing video generation as a *general simulator* capable of physics, human dynamics, and robotics.
Their *interface world model* replaces front-end code (HTML/CSS) entirely — it renders pixels directly, accepts clicks and drags, and is envisioned as a *neural operating system*.
They achieved *real-time* video generation via *step distillation* and *auto-regressive models*, deploying the largest real-time video model for *talking avatars*.
A key insight: *third-person video data* is the most abundant and scalable source for training *robotics models*, requiring only hundreds of hours of fine-tune data versus millions hours of teleoperation data.
The *PhysicsIQ* benchmark shows that scaling video models predictably improves *intuitive physics* understanding, countering the idea that video models are merely "cute video" generators.
"The end game of something like interface world models is you have a *fully neural operating system*. — Anastasis"
"Scaling laws apply to video just like they apply to language models... we have no indications that this is not trending. — Anastasis"
"If we look at how humans learn tasks, a lot of it is by observing *others* — not from first person. That’s how video pre-training works. — Anastasis"