WAIC 2026 | Shengshu's Luo Yihang: From Understanding Language to Understanding the World — General World Models Open a New Chapter for AI

Over the past few years, large language models have driven a major leap in AI. Machines have moved from understanding and generating text toward deep reasoning, tool use, and autonomous execution. But as AI enters dynamic scenarios like video, games, robotics, and autonomous driving, understanding language alone is no longer enough.

The real world is not a static knowledge base of text; it is a continuously changing complex system. Agents need to identify “what is happening,” predict “what will happen next,” judge “what consequences an action will produce,” and adjust decisions based on environmental feedback.

From understanding language to understanding the world — this is becoming the next critical capability leap for AI.

On July 19, the WAIC 2026 Qiming Venture Partners · Entrepreneurship and Investment Forum took place. ShengShu co-founder and CEO Luo Yihang delivered a keynote titled “General World Models: Bridging the Virtual-Physical Divide, Building the Foundation for Physical Intelligence.”

Luo Yihang keynote

World Models: From Parsing Information to Participating in the Real World

World models are not a new concept, but with rapid advances in video generation, multimodal models, and embodied AI, they are moving from academia to the center of industry. In Luo Yihang’s view, language models handle the relationship between humans and language, knowledge, and information, while world models address the interaction between agents and the entire world.

Human intelligence is not limited to language processing. Large areas of the human brain are dedicated to visual perception, action planning, and decision-making. The core of this intelligence is perceiving the physical world and making action decisions based on environmental changes.

“Relying solely on large language models to understand the world is like using only a small fraction of human intelligence. In the era of large models, we also need a world model at the LLM paradigm level.”

Video models need to understand how people and objects move; robots need to predict the outcome of actions; autonomous driving systems need continuous perception and real-time decision-making in complex environments. Behind these different scenarios lie similar foundational capabilities: perceiving the environment, predicting change, and taking action. A complete world model must form a continuous dynamic loop — perception, prediction, and action are not separate modules but should be built on a unified world representation.

Toward Generalization: The Path to Scalable World Model Deployment

Many world models today are built for specific scenarios — home robots, industrial operations, autonomous driving. But if each robot and each scenario requires its own model, data collection, training, and deployment costs become unscalable.

The key insight from large language models is to first build general capabilities through large-scale pretraining, then transfer to different tasks and scenarios. World models need a similar evolution.

“World models must not only form a closed loop of perception, prediction, and action, but also be as general and generalizable as language models.”

Generalized world models

Different agents can have different forms; the digital and physical worlds have different output modalities. But the underlying intelligence driving them should be unified. The real competition in world models is not who achieves higher success rates on individual tasks, but who can build a general foundation model with scaling effects, transferability, and the potential for intelligence emergence.

Video as a Key Entry Point for Understanding the World

Building a world model first requires AI to see and understand the world. Text is an abstract expression of reality; images capture a single moment; video continuously records how the world changes — how people act, how objects respond to forces, how spaces transform, how one event leads to another.

When a video model predicts the next frame, it learns not just pixel generation but the transition relationships between world states. Video models’ value extends beyond content production.

Video as an entry to understanding the world

Building on this, world models also need to absorb simulation data, first-person data, human operation data, and real robot data to translate understanding into real action. ShengShu thus builds a unified general world model foundation: first establishing understanding and prediction of the world, then decoding differently for the digital and physical spaces.

One Foundation, Two Wings: Driving Digital Generation and Physical Action

Around the unified world model foundation, ShengShu has built a product system covering world generation, real-time interaction, and world action.

Product system

The Vidu Q series targets digital content generation; the Vidu S series pushes video from offline generation to real-time, continuous interaction; Motubrain faces the physical world, converting the model’s understanding of environment and tasks into robot actions.

Three product lines for different scenarios, all sharing unified modeling of space-time, causality, and action. “Agents in the digital and physical worlds can vary enormously, but the intelligence behind them should be general, generalizable, and unified.”

From Generating the World to Acting in the World

Large language models enabled machines to understand and use language. Video generation models gave machines visual creation capabilities. General world models aim to push machines further — to understand, predict, and act in the world.

In Luo Yihang’s view, robots of every form will become agents in the physical world, achieving autonomous planning and execution just as digital agents do. The real breakthrough remains the model.

Luo Yihang closing remarks

“General world models are the most promising technical path for physical intelligence. In the future, agents in the digital and physical worlds will be diverse, but the intelligence behind them will be universal — communicating like humans, creating like humans, and acting like humans.”

From processing information to understanding change; from generating content to taking action — AI is entering a new stage of development. And general world models are the key infrastructure connecting digital and physical intelligence, pushing AI into the real world.