Gemini Omni Flash API: conversational video editing at a per-second cost

Google released the Gemini Omni Flash API, the first model in its new Omni family, at $0.10 per second of generated 720p video. The company frames the family's ambition as creating anything "from any input," and the API rollout is the part that puts that framing in front of enterprise teams. The rollout also adds a stateful interactions API that lets edits accumulate across turns rather than restarting from a blank prompt.

The pricing comparison the source provides is where the deployment decision starts to sharpen. At $0.10 per second of 720p, Omni Flash matches Veo 3.1 Fast at the same resolution, runs double Veo 3.1 Lite, and undercuts standard Veo 3.1 by three-quarters. The same table exposes the ceiling: Omni Flash only generates at 720p, with no 1080p or 4K option, while the Veo 3.1 tiers scale to 4K. Internal training and most social formats can live at 720p. Premium brand work meant for a large screen cannot, and Veo 3.1 still has a job in that space.

The unification story is the part decision-makers should weigh first. Many teams have been assembling AI video the hard way, bolting together an LLM for a script, a text-to-image model, an image-to-video model, a separate lip-sync tool, and a voice generator, each with its own contract, billing, and data path. Collapsing those into a single model means fewer vendors and a single place to monitor output and enforce data-handling rules. Organizations that have avoided generative video because stitching the tools together was not worth the overhead will see the equation shift.

The cost structure complicates that shift. Every conversational edit is a fresh generation billed at the same per-second rate. An edit-heavy session that produces ten iterations of a ten-second clip is ten dollars of inference, not one. The source's framing of the change as "book a reshoot versus send a note" undersells this: a sent note in the Omni model is itself a paid generation pass. The stateful interface does not change the cost of an edit; it reduces the number of wasted ones, since context carries forward and the generations go toward refining a take that mostly works.

The reference-input handling is where the model earns its commercial weight. Omni accepts up to seven images and up to three video clips of three seconds or less as inputs, and it carries those specifics into the result. Hand the model a photograph of a specific product and ask it to place that object in a scene, and the output reproduces the real thing's coloring and rough shape rather than inventing a generic stand-in. Two of Google's four highlighted strengths speak directly to enterprise work: a world model that handles physical consistency (rain and puddles render reflections of people and objects in wet pavement) and text and logo insertion that rewrites signage in another language or drops a brand logo into a scene. The source's own testing found that sign tracking in complex scenes was not always perfect and that some text slipped back to the original language between frames. Output that ships as a training video with on-screen labels or an ad with a logo placed in-scene still needs a human review pass.

On quality, the early signal is strong but narrow. In LMArena's Text-to-Video Arena, where voters rank head-to-head outputs from competing models, Omni Flash sat at number one with a score of 1527. Arena rankings are preference signals from a specific population of voters on specific prompts, not controlled capability evaluations, and the source does not characterize the prompt distribution or the breakdown of which kinds of outputs won. The score suggests Omni Flash is competitive against the models in the comparison set, but it does not establish how the model behaves on the specific shots a given marketing or training team needs to produce.

Provenance work is the layer regulators and security teams are most likely to scrutinize, and it ships alongside the model itself. Every Omni clip carries Google's SynthID watermark, Google is extending C2PA Content Credentials across its generative tools, and it has launched an AI Content Detection API that flags AI-generated media from both its own and other vendors' outputs. The model also draws a deliberate line: it will not take a still photo of a person plus an audio clip and lip-sync them into speech, an explicit move to limit deepfakes. It will take a recording of someone talking and translate it into another language, a path that is useful for localizing global training content. Those constraints and the baked-in provenance function as features in regulated environments, where output policy can be enforced at the model level rather than bolted on downstream.

The constraint that most directly shapes the deployment decision is the 10-second cap. Clips run 3 to 10 seconds at 720p native, in landscape (16:9) or portrait (9:16). To produce anything longer, the developer generates chunks and edits them together, and audio cannot be uploaded as an input yet though the model generates audio alongside the video it produces. The source does not report how cleanly Omni Flash handles continuity across those stitched chunks, and Google's own model card is candid that holding consistency across edits and rendering accurate text remain open problems.

Omni is not alone in chasing these budgets. Veo 3.1 remains Google's production-grade option when higher resolution is required, and rivals from Bytedance, Alibaba, and OpenAI are all pursuing the same enterprise video line items. The capability Omni adds is editing itself: a video becomes a living document rather than a one-shot render. The deployment case depends on whether teams can verify throughput, consistency, and cost under their own access patterns, none of which the source independently validates.

Subscribe to AI Enthusiast Log

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe