ShengShu Technology was one of the earliest teams globally to conduct research on multimodal generative algorithms. In September 2022, the team introduced the U-ViT architecture. In July 2024, Vidu launched globally, introducing the industry’s first Reference-to-Video capability. Moving beyond conventional text-to-video and image-to-video models, this innovation addresses one of the core challenges in commercial video generation: maintaining multi-entity consistency. Since its launch, ShengShu has continuously released Vidu Q1, Vidu Q2, and Vidu Q3, with each iteration further improving performance across key benchmarks, including consistency, semantic understanding, motion dynamics, stability, and inference speed.
The recently released Vidu Q3 is the world’s first video model built for storytelling. It supports 16-second synchronized audio-video generation, native 1080p output, advanced cinematic language, precise shot transitions, multilingual text rendering, and multi-language output.
According to the latest rankings from AI benchmarking authority Artificial Analysis, Vidu Q3 ranked No.1 in China and No.2 globally. The model placed ahead of several leading international video generation platforms, positioning Vidu among the world’s top-tier solutions. Artificial Analysis data also indicates that Vidu Q2 maintains the fastest generation speed globally among commercial-grade content generation models.

In December 2025, ShengShu Technology open-sourced its TurboDiffusion framework, enabling a 5-second video to be generated in just 1.9 seconds on a single RTX 5090 GPU, improving video generation efficiency by 100 to 200 times.