Kuaishou has released the new generation of its flagship video generation model, Kling 4.0: in a single pass, it creates coherent clips up to 30 seconds long with quality up to 4K 10-bit HDR and native stereo audio. The lightweight Kling 4.0 Flash version is already available in a limited mode, the full launch is scheduled for October, and the first public comparison with its direct competitor, ByteDance's Seedance 2.5, has already been published. The stakes in this race are advertising video and short dramas, where AI generation is for the first time claiming production-level quality.

image

What Happened

The official Kling 4 release notes were published on September 27, 2026, on the kling.ai website. According to them, a single generation produces clips from 3 to 30 seconds long. It claims resolution up to 4K 10-bit HDR, an ultrawide 21:9 format, support for up to 10 keyframes and up to 15 references — these can be images, video clips, characters, and voice. Audio is generated as two-channel stereo with lip-sync and dialogue in nine languages, and the allowable prompt length has increased to 8,000 tokens. The lightweight Kling 4.0 Flash version, which works with 3–20 second clips in 720p, is already available in a limited rollout via Kling Creator Studio; the full launch of the flagship version is scheduled for October. On October 2, 2026, the blogger "Blah Blah Pro Ai" published a video review with model tests and a comparison with Seedance 2.5.

Context

Kling is Kuaishou's flagship video generation line, and the fourth version is designed as a product, not a demo: the key innovations are aimed at scene control, not just text-to-image. The previous generation, Kling 3.0, limited a single generation to 15 seconds, which meant that long, coherent plots had to be manually assembled from short clips. Kling 4.0 shifts the approach from simple prompt-to-video to reference-conditioned generation: the model builds a scene around the provided faces, products, movements, and voice, maintaining them throughout the clip. The market is already competitive: ByteDance's Seedance 2.5 also produces 30-second clips with native audio and claims up to 50 reference inputs, and Google Veo is also competing in the fight for advertising video production and short dramas. It is precisely the duration, quality, and volume of available control elements that are becoming the main parameters for comparing models with each other.

Why This Matters for the Industry

For the industry, the main shift is the doubling of the duration of a single generative pass to 30 seconds: a short advertising clip or a short drama scene can theoretically be obtained in one pass, rather than stitched together from fragments. Multi-reference control with 10 keyframes and 15 references moves video generation into a staged production pipeline: characters, products, and voices can be reused between scenes, which is critical for serial shooting and brand content. Kling 4.0 Flash is already available, so studios and developers can prototype on the lightweight tier (3–20 seconds, 720p) right now: test the combination of keyframes and references, measure latency and the percentage of generations accepted on the first pass, and observe the stability of characters and voices on references. The competition is open: Seedance 2.5 claims up to 50 reference inputs, and it is precisely this parameter that will determine which model is more convenient for product workflows. Flagship features (30 seconds, 4K 10-bit HDR, the full set of references) are unavailable until the October launch, and there is currently no API description or pricing in public sources — meaning it is too early to integrate the model into a live production chain.

Why This Matters for Users

For readers and practitioners, this means that AI video is for the first time meeting the requirements of real tasks: long coherent scenes, reuse of characters, voices, and products, as well as on-screen text that remains readable during camera movement. This is particularly interesting for those who make advertising, titles, promo clips, or short series: part of what previously required a film crew can now be tried to be generated independently. Native stereo audio with lip-sync and dialogue in nine languages reduces the number of separate tools — voiceover and lip-sync do not need to be done after the fact. The lightweight Kling 4.0 Flash in Kling Creator Studio can be tried right now, and the "Blah Blah Pro Ai" video review with a comparison to Seedance 2.5 reduces the cost of choosing a tool until the full October launch of the flagship version.

What Is Still Unknown / Limitations

Confidence in production applicability currently rests on selected vendor examples and one amateur comparison — this is marketing and a crowdsourced demo, not a reproducible evaluation. The defect rate on routine prompts is unknown, as are real latency metrics; independent benchmarks and a technical report from Kuaishou do not yet exist. The quality of following a long prompt of up to 8,000 tokens has not been measured anywhere: the available video review has no disclosed prompts or stated comparison criteria. The architecture of the native audio is not disclosed — it is unclear whether it is a single multimodal model or a pipeline of separate modules, and this affects the stability of dialogue scenes against desynchronization. The practical limit of the useful number of references is also unknown: more inputs do not automatically mean a better result, and no source measures this dependency. There is no pricing for the full launch in public sources.

Sources

Author

Look at AI, editorial team