Google has released Gemini Omni 1.1 Flash — a production update to the Gemini Omni video model, available via the Gemini API in Google AI Studio, Gemini Enterprise Agent Platform, Flow, and the Gemini app. The model extends scenes up to 40 seconds in 10-second steps, relying on 10 seconds of previous context, generates video between specified first and last frames, and accepts up to 3 seconds of video reference. Drafts render in 360p 60% faster and three times cheaper than 720p, while the final output is upscaled to 1080p and 4K. Adobe Firefly, Runway, and Figma Weave have already integrated the model into their tools.

What happened
The update was released directly to production rather than shown as a research demo: the model is accessible via the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform, and is also available in Flow and the Gemini app. The key technical change concerns scene extension: previously, the model only considered the last second of the previous segment, but now it analyzes up to 10 seconds of context, allowing scenes to be built up in 10-second steps up to 40 seconds. Video generation between specified first and last frames has been added, and up to 3 seconds of video reference can be provided as input. Rendering follows a two-stage scheme: first, a fast 360p draft, which is 60% faster and three times cheaper than 720p, then upscaling of the final video to 1080p and 4K. In the announcement, Google also reports that the model took first place on Video Arena and that Adobe Firefly, Runway, and Figma Weave have already integrated it into their products.
Context
The background for reading this release is a paradigm shift in video generation. Until recently, the typical scenario was one-time clip generation from a prompt, where the result largely depended on luck, and the coherence of long scenes was a weak point. Expanding the context window from one to ten seconds is an architectural requirement for continuous storytelling: without memory of the previous segment, a long scene breaks down into disconnected pieces. Interpolation between the first and last frame and video reference are features of scene control, not a declarative quality increase: the industry is moving from a "prompt lottery" to a managed pipeline, where a cheap draft run is separated from an expensive final. This two-stage approach methodologically mirrors how iterative image generation has long been structured, where variants are explored at low resolution and upscaling is done at the end. The model's maturity is indicated not only by the vendor's word: third-party companies have integrated it into working tools for editing and design, which is a stronger argument for assessing production-readiness than a position in a vendor rating.
Why this matters for the industry
For the industry, the release means moving video generation from the category of demos to the category of working tools. Extending scenes based on 10-second context and interpolating between frames for the first time make multi-shot scenes with a continuous camera without cuts feasible — this is a direct challenge to Veo and Sora-class models in professional video production. The economics of iteration are changing: the cheap draft mode reduces the cost of exploring variants, so the "cheap draft → expensive final" pattern is expected to become established in product stacks, and idea selection will shift to the early stage of the pipeline. For startups whose value is reduced to "one clip from a prompt," there is price pressure: Google makes the same primitive generally available through the API, Flow, and the app. If the stated characteristics are reproducible, scene control — extension, frame interpolation, references — risks becoming a hygienic minimum, and competitors will have to respond either with comparable features or with public comparison methodologies. The expected wave is keyframe-first interfaces, storyboard tools, and agentic pipelines "script → draft → upscale," where differentiation moves to orchestration and vertical templates.
Why this matters for users
The practical benefit for the reader is that experiments cost nothing at the start: in Google AI Studio, 360p drafts are available for free, so scene variants can be explored without a rendering budget, and finalization to 1080p and 4K is available to Google AI Plus, Pro, and Ultra subscribers in the Gemini app and in Flow. Editing skills are not required: a scene can be assembled from key frames by specifying the first and last frame, extended in 10-second steps up to 40 seconds, and up to 3 seconds of video reference can be passed to the model to set the style or movement. Realistic first scenarios are concept previews, short scenes up to 40 seconds, storyboard selection, and quick iterations on a reference before ordering a final render.
What is still unknown / limitations
All key facts in the material come from the vendor's announcement: there is no technical report, methodology, or independent quality assessment in the source, so the internal mechanism of the 10-second context — conditioning, video tokenization, memory — remains undisclosed. The announcement does not include absolute prices or latency data. The first place on Video Arena is known only from Google's post; the integrations of Adobe Firefly, Runway, and Figma Weave are a more substantial independent signal of maturity, but they do not replace measurements. Quantitative claims about the speed and price of drafts are falsifiable — access via the Gemini API allows reproducing measurements in days — but for now, these are the vendor's words. Conclusions about the commoditization of raw generation and the emergence of a new industry standard are interpretations on top of a single announcement, and they should be verified by independent tests.
Sources
Author
Look at AI, editorial team
