Introducing Vidu S1: Real-Time Interactive Video Generation

Over the past few years, video foundation models have made remarkable progress in generating high-quality video from text, images, and other inputs. But despite these advances, most models still follow the same workflow: users submit a prompt, wait for the model to finish generating a clip, and then watch the final result. Once generation begins, the content is largely fixed, leaving little opportunity for the model to respond to new user input.

Vidu S1 introduces a different approach.

Rather than treating video generation as a one-time creation task, Vidu S1 is designed for continuous, real-time interaction. The model can listen, understand, and respond while video is being generated, enabling AI avatars to engage in natural conversations that evolve dynamically with user input.

Real-time interaction requires far more than reducing latency. An interactive video model must continuously process new user input, preserve conversational context, maintain character consistency, and generate visually coherent responses frame by frame — all while the conversation is still taking place. Meeting these requirements demands tight coordination across data, model architecture, training algorithms, inference optimization, and deployment infrastructure. Vidu S1 was designed from the ground up with this goal in mind. By combining advances in generative modeling with highly optimized inference and serving systems, the model can continuously generate video, preserve long-term character identity, and respond immediately to new voice input throughout an interaction. This integrated architecture enables four core capabilities that distinguish Vidu S1 from conventional video generation models:

Voice-Driven Character Control

Voice-Driven Character Control

Go beyond lip sync. Drive character behavior with natural voice commands.

Unlimited Real-Time Interaction

Unlimited Real-Time Interaction

The world's leading generative video model for unlimited real-time interaction.

540P · 25 FPS Streaming Generation

540P · 25 FPS Streaming Generation

High-definition, low-latency interactive video generation at up to 42 FPS.

Custom Image & Voice

Custom Image & Voice

Create personalized AI characters instantly from your own image and voice.

These four capabilities set Vidu S1 apart from traditional video generation models, enabling AI avatars to remain continuously online — listening, responding, and interacting in real time.

Voice-Driven Character Control

Traditional AI avatars typically rely on speech-driven lip synchronization. While these systems can match mouth movements to speech, facial expressions, gestures, and body movements are often selected from predefined animation libraries. As a result, interactions remain predictable and limited, making it difficult for avatars to respond naturally to changing conversations.

Vidu S1 takes a fundamentally different approach.

Instead of treating speech merely as an input for lip synchronization, Vidu S1 interprets spoken language as behavioral intent. During a conversation, the model simultaneously understands what the user says, how it is said, and the current visual context, generating synchronized facial expressions, eye movements, gestures, posture, and full-body actions in real time.

The interaction pipeline evolves from:

Voice Input → Lip Synchronization

to:

Voice Input → Intent Understanding → Real-Time Behavioral Generation

This allows AI avatars to move beyond simply speaking on screen. They become generative characters capable of understanding users, responding naturally, and expressing behaviors that continuously adapt throughout an interaction.

Unlimited Real-Time Interaction

Most video generation models today produce clips with a fixed duration. Users provide an input, the model generates a complete sequence, and the video's future frames are largely determined from the start. If a new voice instruction arrives during generation, the model typically cannot modify the remaining content in real time; it must generate an entirely new clip. This is why most existing systems function primarily as offline video creation tools.

Vidu S1 takes a different approach.

Instead of generating an entire video upfront, Vidu S1 uses an autoregressive diffusion (AR + Diffusion) architecture that continuously predicts the next frame based on previously generated frames. This allows the model to incorporate new voice input while generation is still in progress, enabling users to influence how the video evolves in real time.

Previous Video Context + Live Voice / Instructions / Conversation ↓ Generate the Next Frame ↓ Update Character Memory ↓ Continue Generating

Rather than being a pre-generated sequence, the video becomes a continuously evolving interaction.

Maintaining stability over long conversations introduces another challenge: character consistency. Vidu S1 addresses this by maintaining two complementary forms of context during generation:

Long-Term Identity Anchor Maintains consistent character identity, appearance, clothing, and visual style throughout the interaction. Short-Term Sliding Window Preserves smooth continuity in recent facial expressions, body movements, and poses across consecutive frames.

Together, these mechanisms allow AI avatars to remain online for extended periods while continuing to respond dynamically to new voice input, enabling truly persistent interactive video generation.

Traditional Video Generation: Input Prompt → Generate a Fixed-Length Video → Playback Ends Our Autoregressive Video Generation: Previous Frames → Live Voice Input → Generate the Next Frame → Continuously Adapt Future Video Content

540P · 25 FPS Streaming Generation

Traditional video generation models were designed primarily for offline content creation. After receiving a prompt, the model spends time computing a complete video before returning the result. This workflow works well for short-form video production, but it is not sufficient for applications such as video conversations, interactive livestreaming, AI companionship, or real-time gaming, where the system must generate continuously, respond quickly, and maintain a stable frame rate.

Vidu S1 was redesigned specifically for real-time interaction. Its target is 540P HD resolution (960 × 540) at 25 FPS, with support for up to 42 FPS, allowing the model to behave more like a video call than a traditional offline video generator.

Achieving this required optimization at both the model level and the system level.

Model Layer: Reduce per-frame generation cost to accelerate every generated frame. System Layer: Optimize streaming scheduling to ensure stable, continuous video output.

On the model side, Vidu S1 integrates TurboDiffusion, SageAttention, SLA, SpargeAttention, model quantization, and inference kernel optimizations to significantly reduce the computational cost of generating each frame.

On the deployment side, Vidu S1 adopts a streaming serving architecture inspired by TurboServe. Interactive video generation is treated as a persistent online session rather than a one-time offline task. The system continuously maintains user input, AI avatar state, and historical visual context while dynamically allocating compute resources to keep generation stable and responsive.

Instead of following the traditional workflow: Submit Request → Wait for Video → Play Result Vidu S1 enables: Live Voice Input → Frame-by-Frame Generation → Continuous Playback → Continuous Response

To further reduce perceived latency, the model generates future frames while the current frames are being displayed, allowing new voice instructions to be injected immediately into subsequent generation.

As a result, 540P at 25 FPS represents far more than image quality or frame rate — it provides the technical foundation required for real-time AI video across video calls, live streaming, AI companions, gaming, and XR experiences.

Custom Image & Voice

Traditional AI avatar creation often requires multiple images or video assets, followed by modeling, rigging, lip-sync adaptation, and character-specific training. Depending on the workflow, creation can take anywhere from several minutes to an entire day.

Vidu S1 removes this process entirely.

Using a fully generative pipeline, users only need to upload a single reference image. The model immediately captures the AI avatar's identity, appearance, and visual style, then generates synchronized facial expressions, lip movements, gestures, and body motion in real time — without preprocessing, offline modeling, or character-specific training.

Vidu S1 also supports customizable voices, allowing visual identity and vocal identity to remain consistent throughout the interaction.

Character creation is simplified from:

Upload Assets → Train → Interact

to:

Upload One Image → Start Interacting Immediately

Making personalized real-time interactive characters dramatically easier to create.

The Technology Behind Vidu S1

Real-time interactive video generation is not the result of a single breakthrough. It requires coordination across model architecture, training algorithms, inference optimization, and deployment infrastructure.

At the inference layer, Vidu S1 integrates TurboDiffusion, SageAttention, SLA, SpargeAttention, model quantization, and optimized inference kernels to substantially reduce the computational cost of generating every frame.

At the serving layer, Vidu S1 adopts the streaming architecture introduced by TurboServe, transforming video generation from a one-time offline task into a persistent online generation session. The system continuously maintains user input, character state, and video history while dynamically allocating compute resources, enabling stable long-duration generation with low latency and rapid responses to new voice input.

The combination of model-side and system-side optimization forms the technical foundation that enables Vidu S1's real-time interactive capabilities.

Vidu S1 is now fully open for beta testing. Users can customize the initial image and experience real-time interaction, and the API platform is also available:

Try it now: https://www.vidu.com/vidu-stream
API Platform: https://platform.vidu.com/

References

  1. [1] TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times.
  2. [2] SageAttention: Accurate 8-bit attention for plug-and-play inference acceleration.
  3. [3] SLA: Beyond sparsity in diffusion transformers via fine-tunable sparse-linear attention.
  4. [4] SpargeAttention: Accurate and training-free sparse attention accelerating any model inference.
  5. [5] TurboServe: Serving Streaming Video Generation Efficiently and Economically.