ShengShu Technology today announces Motubrain, a World Action Model that replaces multiple task-specific systems with a single, unified model that functions as a robotic brain for the physical world. Ranking highly on both WorldArena and RoboTwin 2.0, two of the field’s most rigorous benchmarks in embodied world models, Motubrain marks a decisive shift in an industry where robotic systems are typically built from task-specific or specialized systems.

Figure 3 Task scaling results. For each point on the curve, we train the model using data from only the specified number of tasks and keep optimization running until convergence, i.e., until the average success rate becomes approximately stable. As the number of training tasks increases, Motubrain shows a clear upward trend in average success rate and consistently stronger scaling behavior than conventional VLA baselines.
Figure 4 Data scaling results. For each data budget, we uniformly subsample demonstrations from every task to keep the per-task data distribution balanced. The number of training steps is set proportional to the total amount of training data; in particular, training on the full 27,500-trajectory dataset uses 50,000 optimization steps. Motubrain continues to benefit from additional training data, indicating that larger-scale supervision improves policy performance and robustness.
Best known for its leading video model Vidu, ShengShu Technology and its advancements in generative video for robotics earmarks an industry first. Generative video has laid the foundation for simulating robots in real-world environments at scale. Motubrain builds on this by turning those simulations into action, by enabling robots to learn from diverse, large-scale pre-training data while reducing reliance on traditional physical data collection.
“A true world model must be able to build a unified representation of the real world and predict how it evolves,” said Jun Zhu, Founder of ShengShu Technology. “Video is a critical foundation of that intelligence because it naturally captures time, space, motion, causality, and physical dynamics at scale. We believe general world models should not be built as stitched-together modules, but as a unified architecture that brings together perception, reasoning, prediction, generation, and action in a single system. That is what can ultimately bridge the digital world and the physical world.”