WAIC 2026 | Shengshu's Zhu Jun: General World Models as New Infrastructure Bridging Digital and Physical Worlds

On July 20, during WAIC 2026, ShengShu Technology and Wondershare co-hosted a forum on “A New Paradigm for AI Film Production Driven by World Models.” At the forum, ShengShu founder and Deputy Director of Tsinghua University’s Institute for AI, Zhu Jun, delivered a keynote titled “General World Models: New Infrastructure Bridging the Digital and Physical Worlds.”

Zhu Jun stated that AI is moving from understanding and generating digital content toward modeling, reasoning about, predicting, and acting upon the laws of the world. General world models will become critical infrastructure connecting the digital and physical worlds.

Zhu Jun keynote

What Is a World Model

Zhu Jun on stage

What is a world model? Consider the human world model: before riding a bike or driving, we form a mental model that helps us understand the environment, imagine outcomes, plan actions, and guide our movements.

Today, the world is focused on building machine world models. At their core, machine world models need three capabilities: first, precise environmental understanding; second, prediction and imagination based on the current state and intended actions; third, learning and acting during prediction. A machine world model must achieve an integration of understanding, imagination, and action.

The biggest shift in this field is from specialized models to general-purpose models with Scaling Law.

A Six-Layer Data Pyramid and a Unified Architecture

To build a general model, we first need to answer: what data should we use? We organized our data into a six-layer pyramid, from general internet video at the bottom to first-person human action video and robot operation data at the top.

Six-layer data pyramid

Video is the dominant format in this data system. Our principles are: large-scale data, comprehensive data types, and real data wherever possible.

We proposed the world’s first MoT (Mixture-of-Transformer) unified architecture, integrating understanding, generation, and action experts into a single model trained with a unified objective.

MoT unified architecture

As an integrated model, it can freely switch output forms: direct action output for robots, pixel-level video for creators. ShengShu is building a “general world model.”

World generation and world action models

From Digital Content to Real-Time Interaction

On the digital content side, models like Vidu Q3 achieve precise understanding and high-dynamic, high-consistency generation through large-scale training and Scaling Law.

Vidu capabilities

We recently released a real-time generation model. Video models can now serve as real-time interaction models. In digital human scenarios, we use fully streaming generation for more natural and human-like character behavior.

Real-time generation model

The model supports unlimited-length continuous dialogue, 540P video output, and up to 42 FPS.

Streaming generation

From the Digital World to the Physical World

Our ultimate goal goes beyond digital content generation. Both Vidu video generation and Vidu S1 streaming generation prepare for real-time physical interaction. These core technologies also support model action in the physical world.

Motubrain, released in April, can drive multiple heterogeneous robot bodies from a single model base to complete complex long-horizon tasks.

Embodied AI robots

A key advance is that with a strong foundation model, different robot bodies can rapidly learn complex skills. Users can give a text instruction, and the model decomposes the task into steps and generates precise action commands.

This is the fundamental difference from traditional VLA approaches. Traditional VLA fits trajectories, while our generative pretraining enables robots to plan actions based on imagination of the future, producing more reliable, stable, safe, and efficient actions.

Looking Ahead

Zhu Jun closing remarks

The general world model we discuss today is not merely an upgrade of video models or an extension of language models. It is a new paradigm driving AI from generating content to understanding, reasoning about, and interacting with the world in real time.

ShengShu has been positioned around foundation models since its founding, starting from video modeling. In just two years, video foundation models have gone from demonstrating technical possibility to genuinely entering industry and supporting production scenarios.

Future outlook

Video data is a rich, large-scale information carrier beyond language. With large-scale pretraining on video, models achieve precise understanding, high-quality imagination and prediction, and generalizable, cross-task, cross-embodiment action capabilities. Combined with efficient post-training, this drives rapid industry iteration and deployment.

We look forward to collaborating with more industry partners to explore the intelligence ceiling and industrial applications of general world models.