Google DeepMind released Gemini Omni 1.1 Flash on August 27, the latest version of its text-to-video model aimed at developers building creative and editing tools rather than end consumers directly. The headline change is length: scene extension, which continues an existing clip by analyzing up to 10 seconds of prior context, now chains up to four times to reach a cumulative 40 seconds -- up from 10 seconds in the prior release.
Gemini Omni 1.1 Flash, in short
- Released
- Aug. 27, 2026
- Max clip length
- 40 seconds
- Standard price
- ~$0.10/sec
- Draft price
- ~$0.03/sec
- Where
- Google AI Studio, Gemini API, Gemini Enterprise, Google Flow
Two other controls ship alongside the length increase. Developers can now specify a first and last frame and let the model fill the continuous shot between them -- useful for camera moves that are hard to prompt into existence from text alone -- and a new 360p draft mode renders a preview at roughly a third of the standard cost and up to 60% faster, so a team can iterate on several cheap drafts before paying full price to upscale the one that works to 1080p or 4K. Google's own reference pricing puts standard 720p output at about $0.10 per second; the draft tier runs roughly $0.03 per second -- about $0.30 for a 10-second draft.
Google says the model now tops its own Text-to-Video Arena leaderboard at 1,515 points -- a figure from Google's own August 27 announcement that, as of this writing, no independent benchmark has replicated. The claim lands in a video-generation field that is simultaneously losing one competitor and facing another: OpenAI's Sora 2 API is scheduled to shut down September 24, while ByteDance's Seedance 2.5 already generates clips up to 30 seconds with native 4K export and up to 50 simultaneous reference inputs -- a wider ceiling than Omni 1.1 Flash's four-chain, 40-second cap. Two early integration partners quoted in Google's own release describe the model's edge as accuracy rather than raw duration: Louisa Guo of GMI Cloud singled out “its accuracy,” and Itay Schiff of Figma Weave called it “one of the strongest video models available” inside that product.
- Google DeepMind released Gemini Omni 1.1 Flash on Aug. 27, its latest text-to-video generation model.
- Scene extension now stretches continuous clips from 10 seconds up to a cumulative 40 seconds.
- A new 360p draft mode previews clips 60% faster at about a third of standard cost.
- Standard 720p generation runs about $0.10 per second; 4K upscaling is a separate paid step.
- Caveat: Google's claimed #1 Text-to-Video Arena ranking of 1,515 points is a self-reported score.