Runway Co-CEO Germanidis: The World Simulator is Key
Runway co-founder and co-CEO Anastasis Germanidis stated in his keynote speech at the Runway AI Summit held at The Masonic in San Francisco that the universal world simulator will be the most important technological advancement of this era. The company simultaneously published a record of the summit, which took place at 1111 California Street, from 9 AM to 6 PM.
The summit focused on world models, featuring speakers including Ming-Yu Liu, Vice President of NVIDIA Cosmos Labs, Quan Vuong, co-founder of Physical Intelligence, Jon Barron from Google DeepMind, Paril Jain, co-founder and CTO of The Bot Company, and Vivek Viswanathan, senior advisor to the California Governor's office. Co-founder Cristóbal Valenzuela previously described the event as a concentrated showcase of robotics, world models, and physical AI.
The product line has shifted from the Gen series of video generation to GWM world models. GWM-1 is designed for frame-by-frame generation, real-time interaction, accepting camera poses, robotic commands, and audio control; GWM Worlds 2 further provides playable worlds: continuous 720p, 24 frames, 48,000 Hz audio, breaking down "persistent scenes, subjects, physical rules" and "variable actions," with no preset duration for conversations. Solaris treats clicks and drags as action spaces, generating the interface frame by frame without writing code first. The company's direction is to integrate game worlds, robotic arms, and software interfaces into a single foundational model, only changing the action space.
Funding and valuation have aligned with pre-training. In February, General Atlantic led a $315 million Series E round, with NVIDIA and AMD Ventures participating, raising the valuation to approximately $5.3 billion, explicitly for next-generation world model pre-training and product expansion into new industries. In an interview before the summit, Germanidis described the long-tail problem of autonomous driving: the answer may lie in synthetic images never captured from a vehicle, comparing them with "200 million miles of real vehicle data" routes.
The public roadmap previously indicated that video models at sufficient scale would evolve into world models, as predicting the next frame requires learning how objects move, how forces are transmitted, and how actions lead to results; language models distill existing text, which cannot surpass physical reality. The company has provided a time anchor: about five years to achieve interactive simulations indistinguishable from the real world at a human scale, and about ten years for physics and biology to be sufficiently accurate to solve a significant proportion of scientific problems. In his Berkeley speech, he summarized the method in one sentence: using the prediction of the next frame to predict the world.
In market mechanisms, buyers include robotics, autonomous driving, and gaming companies seeking synthetic training grounds, while sellers are labs packaging video models into interactive simulators. Funding is shifting from film editing subscriptions to "wind tunnels"—running millions of scenarios in synthetic worlds without building real laboratories. Beneficiaries are model vendors mastering real-time video generation and action control, as well as chip manufacturers selling computing power for pre-training; those under pressure are teams relying solely on real vehicle mileage or manually written simulators, as well as old clients still viewing Runway as a Hollywood effects tool. The summit is driven by product narratives, not a one-time model weight open-source release.
On a supplementary note: Runway's safety head Conner McDowell is also on the agenda. Gen-4.5 is still referred to by the company as the latest foundational video model, with the GWM series trained on it by domain. Germanidis acknowledged that the current Worlds, Avatars, Robotics, and Solaris are still separate post-training applications, and the unified foundation is still an ongoing project, not a delivered single weight.
Source: Public Information
ABAB AI Insight
Runway has transitioned from Gen-2 text-to-video to GWM, treating "next frame" as a physics lesson rather than prompts as directorial instructions. Germanidis, along with Cristóbal Valenzuela and Alejandro, positions the company as a dual capability of art and science: Hollywood has paid to train motion priors, which are then sold to robots. The Series E round bringing in NVIDIA and AMD acknowledges that the bottleneck for world models lies in pre-training computing power, not in the palette.
Capital is shifting from subscription-based effects to simulation as data. The cost of real vehicle data over 200 million miles lies in collection and annotation; the synthetic world turns long-tail scenarios into infinitely replayable pixels. Solaris applies the same logic to the software interface, changing the action space from steering wheels to mice, indicating their bet on "one world model for all interactions," rather than creating another vertical driving model.
In comparison to Google's Genie, NVIDIA Cosmos, and Wayve's driving world models, the industry is in an expansion phase from generating clips to generating playable environments, with unified weights yet to be completed, initially using post-training to capture the market.
Structural judgments belong to technological substitution. Manual simulators and real vehicle logs are scarce capital; video prediction turns scarcity into computing bills. The mechanism is: whoever can consistently make the next frame adhere to forces and causality can lower the laboratory threshold from physical space to GPU; language models are limited to already written knowledge, while world models are competing for the physical that has yet to be written into text.
ABAB News · Cognitive Laws
- Predicting the next frame is paying the physics tuition.
- When laboratories are scarce, synthetic worlds first become wind tunnels.
- Action spaces can change; only then can a world model be considered universal.