On March 12, Anyverse Dynamics, a general-purpose embodied intelligent robotics company, and Shengshu Technology, a multimodal foundation model and world model company, announced a comprehensive strategic partnership.
The two parties will leverage their respective core technology strengths, deeply integrating Shengshu Technology’s advantages in multimodal foundation models and world models with Anyverse Dynamics’ capabilities in embodied intelligence “universal brain” and “manipulation intelligence” R&D, integrated hardware-software system engineering, and full-scenario scalable deployment, to jointly overcome key technical bottlenecks in embodied intelligence and drive scalable applications globally.

Currently, the embodied intelligence industry is at a critical stage of transitioning from technical exploration to industrial deployment. While large language models have made breakthroughs in semantic understanding and knowledge reasoning, the real challenge for Physical AI lies in: how to enable models to not only understand “semantics” but also comprehend and internalize the physical laws of the real world.
The industry still widely faces the architectural problem of fragmented “perception — reasoning — action”, making it difficult for robot systems to operate stably in complex real-world environments. Therefore, paradigm innovation in foundational architectures and the construction of native world models are becoming the key path to general-purpose embodied intelligence.

At the signing ceremony, Anyverse Dynamics CTO Xu Wenda and Shengshu Technology’s World Model Head Tan Hengkai signed the agreement. Zhu Jun, Founder of Shengshu Technology, Luo Yihang, CEO of Shengshu Technology, Zhang Yufeng, Founder and CEO of Anyverse Dynamics, and Wang Yingyun, Co-founder of Anyverse Dynamics, witnessed the signing.
Video naturally records physical space-time, causal relationships, and dynamic evolution in the real world, serving as an important information carrier connecting perception and action. By deeply combining the “physical understanding and imagination capabilities” of video generation models with the “execution capabilities” of robot systems, both parties will jointly build generative world models with physical intuition, enabling robots to perform “mental rehearsal” before taking action.