Google's next-generation conversational video generation and editing model
Gemini Omni Flash (model ID: gemini-omni-flash-preview) is Google's next-generation video generation and editing model, designed for real-time, omnimodal interactions. It accepts text, image, and video inputs and outputs short video clips (currently up to 10 seconds at 720p). Available through Google AI Studio, the Gemini API, Vertex AI (Agent Platform), the Gemini app, and Google Flow, it is purpose-built for speed, interactive video workflows, and conversational video editing. It is a paid-tier-only model with no free tier, priced on output token consumption at approximately $0.10 per second of 720p video output.
Who it's for
developersvideo creatorsAI researchersproduct teams building video-generation features
Pricing · usage-based
checked 1d ago
Plan
Price
Includes
Standard (Paid API)
Free
~$0.10 per second of 720p video output · $17.50 per 1M output tokens (billed at 5,792 tokens/second) · $1.50 per 1M input tokens · No Batch API discount available · No free tier; paid Gemini API access required
AI-researched pricing — verify on the official site before subscribing.
Use it for
— Generating short AI videos from text or image prompts
— Conversational video editing via natural language instructions
— Integrating video generation into apps via the Gemini API
— Building interactive, low-latency video workflows in Google Flow
— Prototyping video AI features in Google AI Studio
Get the most out of it
01Use the model ID 'gemini-omni-flash-preview' when calling the API — it is not the same as other Flash models
02Budget carefully: at ~$0.10/second of output video, costs scale quickly for batch or high-volume use cases
03Note that Google uses your content to improve its products with this model, even on the paid tier — review data usage policies before sending sensitive content
04There is no Batch API discount for Gemini Omni Flash, unlike most other Gemini models, so plan your usage accordingly
05Test and prototype in Google AI Studio before integrating into production pipelines to understand output quality and token usage patterns
A stack enabling indie developers to produce engaging video content using AI tools.
How the workflow runs
01gemini-omni — Create scripts and storyboards using multimodal content understanding.
02diffusiongemma — Generate high-quality video scripts tailored for specific game content.
03gemini-omni-flash — Produce and edit the video content automatically using the generated scripts.
Combining multimodal understanding with script generation and automated video production accelerates video content creation, allowing indie developers to focus on creative storytelling.