On April 10, Shengshu Technology announced the completion of nearly RMB 2 billion in Series B financing. This round was led by Alibaba Cloud, with strategic investments from CNIC, Jiu’an Haitang, TAL Education, and Guanghe Ventures, alongside continued follow-on investments from existing shareholders including Xinglink Capital, Daitai Capital, CDFE Emerging Investment, BV (Baidu Ventures), and Zhuoyuan Asia.
Investment News
Shengshu Technology Closes Nearly RMB 2 Billion Series B Round, Powering Next-Generation Productivity for the Digital and Physical Worlds with a General World Model
Building AGI with a General World Model
We believe that compared to large language models, general world models inherently carry multimodal information from the physical world — vision, hearing, and touch — and can replicate how humans perceive, simulate, and interact with the physical world. This enables AI to achieve natural interaction comparable to real humans, realizing true physical AGI. Globally, Shengshu Technology is the first company to achieve a unified general world model spanning both the digital and physical worlds, committed to building a general intelligence system for precise modeling, reasoning, prediction, and action across both domains. At its core lies the Foundation World Model, powering the World Generation Model (WGM) for the digital world and the World Action Model (WAM) for the physical world.
Vidu Model: Powering Productivity and Creativity for Digital Content and Interaction
The Vidu model series achieves synchronized audio-video generation, long-duration output, high spatiotemporal consistency, and cinematic visual quality. It pioneered the “Reference-to-Video” technology, effectively solving the pain point of multi-subject consistency in commercial scenarios, while leveraging proprietary efficient training and inference architectures along with outstanding engineering optimization to achieve efficient generation and ultimate cost-effectiveness.
After the launch of Vidu Q3, it ranked first globally in the benchmark published by Artificial Analysis, an authoritative international AI testing institution. Vidu Q3’s “Built for Drama” feature supports up to 16 seconds of synchronized audio-video output with multi-shot switching, camera control, BGM and sound effect generation, and multilingual dialogue.
The Vidu model series serves global developers, creators, and enterprises through MaaS (Vidu AI Open Platform) and SaaS (Vidu Agent, Vidu Claw). Vidu is now also available on Alibaba Cloud’s Bailian platform, providing leading audio-video generation capabilities for industries including internet, advertising, animation, education, and tourism.
Motus Model: Building a Unified General-Purpose “Brain” for Embodied Intelligence
In December 2025, Shengshu Technology open-sourced Motus, the first world action model with a unified architecture based on a large video generation model. Built on the UniDiffuser unified modeling framework, it integrates multimodal knowledge to achieve unified expression and generation of language, video, and actions. The Motus model was the first globally to validate the Scaling Law for embodied foundation models, pioneering the “GPT-2 moment” in the embodied intelligence field.
As the “brain” for real-world embodied intelligence, Motus aims to solve the core pain points of traditional embodied intelligence — fragmented pipelines, data scarcity, and lack of generalization — driving robots from “modular execution” toward “unified intelligent agents.” Motus achieves approximately 40% higher success rates in multi-task scenarios compared to the leading international VLA model Pi0.5, demonstrating the strong generalization capability and scalable potential of general world models in the physical world.
From “understanding the world” to “generating the world” and then to “acting in the world,” Shengshu Technology is exploring and advancing a more comprehensive path to general intelligence. As the company’s proprietary core technologies continue to innovate and evolve, general world models will continue to unlock productivity value in the digital world while accelerating the adoption of general intelligence in physical industrial scenarios.
Investor Quotes
CNIC representative: CNIC has always anchored on national strategic priorities and fully supported independent innovation in critical core AI technologies. Shengshu Technology has deep expertise in the video generation track, building globally leading audio-video generation capabilities with its Vidu model series.
Cai Wei, Partner at Guanghe Ventures: General world models are becoming the next core path to AGI after large language models. Shengshu Technology, built on the U-ViT foundation, has connected multimodal perception with unified modeling capabilities, establishing a complete closed loop from understanding to generation to action.
Gao Xue, Managing Partner & CEO of BV (Baidu Ventures): “BV has firmly supported Shengshu Technology since the first round. Our continued investment in Series B reflects our strong recognition of the team’s core technical capabilities and our unwavering confidence in the general world model track.”
Zhu Jun, Founder of Shengshu Technology: “The core of world models is to equip AI with the ability to form unified representations and predictions of the real world. We aim to bridge the complete chain from perception to action through a unified model architecture, making general world models truly the bridge between the digital and physical worlds.”
Meanwhile, Shengshu Technology is accelerating its global talent strategy. The company is dedicated to assembling top-tier talent with deep technical expertise, outstanding product innovation capabilities, and global strategic vision, building a hardcore team to support frontier model development and global business expansion.