Skip to main content
Gemini Omni Flash is Google DeepMind’s high-quality, cost-efficient video generation and conversational editing model. First introduced at Google I/O 2026 as part of the Gemini Omni family, it combines Gemini’s multimodal reasoning with native video creation, enabling developers to generate, edit, and remix videos through natural conversation.

What Gemini Omni Flash is good at

  • Conversational video editing: Refine and edit videos using natural language: swap characters, relight scenes, alter angles, add or remove objects while maintaining original audio and video tracks
  • Multimodal input: Combine text, images, and video inputs to guide generation. Natively generates synchronized audio with every video output
  • World knowledge and simulation: Combines physics understanding with Gemini’s knowledge of history, science, and cultural context, enabling meaningful storytelling beyond photorealism
  • Text and action synchronization: Render legible text and graphics directly into video, syncing kinetic typography with on-screen movements
  • Pricing: $0.10 per second of video output, matching Veo 3.1 Fast pricing

Use it in ComfyUI

Gemini Omni Flash workflows

Run the text-to-video, image-to-video, and video edit workflows in ComfyUI, locally or on Comfy Cloud