ShengShu Technology has introduced major updates to the Vidu platform, including the One-click MV tool for automatic music video creation and the new Vidu S1 model, capable of generating interactive video characters in real time.

What Happened
As part of the Vidu platform update, the One-click MV tool was launched, allowing users to create music videos in up to 1080p resolution within minutes using only an audio file and a reference character image. Simultaneously, the Vidu S1 model was released, utilizing an autoregressive diffusion architecture to generate interactive video avatars at 540P resolution with frame rates up to 42 FPS, enabling interaction with characters via voice commands.
Context
The technological shift lies in the transition from traditional batch video generation on demand to streaming content generation. The use of autoregressive diffusion in the Vidu S1 model allows for the preservation of visual and emotional context throughout an entire dialogue, which is critical for creating lifelike interaction.
Why It Matters for the Industry
For the AI industry, this signifies the transformation of production tools into platforms for live interaction. The technology opens new markets for creating intelligent NPCs in video games, virtual streamers, and interactive AI assistants capable of maintaining long-term memory and emotional connection with the user.
Why It Matters for Users
Regular users gain the ability to instantly turn photos into "talking heads" or musical characters with a single click, as well as interact with AI characters in real time, where their facial expressions and movements adapt to the interlocutor's voice.
What Is Not Yet Known / Limitations
At this time, detailed data regarding model inference costs and the availability of specialized APIs for industrial production use are missing, leaving questions regarding the scalability of the solution.
Sources
Author
Look at AI, Editorial Team
