All models

omni-flash

GoogleVideo
Get your API key
omni-flash

Create videos and edit scenes with text and reference materials

omni-flash is the video creation entry point for Google Gemini Omni Flash, suitable for creating new shots starting from text concepts, still images, or existing videos. It can generate visuals from prompts and use reference materials to adjust scenes, styles, and elements. Through this platform, you can choose landscape or portrait aspect ratios and output resolutions, and integrate into asset production workflows through asynchronous tasks.

GoogleModel brand
VideoModel type
VideoTask capability

Specifications and API features

Clarify capacity, inputs and outputs, and invocation methods before selecting a model.

Creation methods
Text-to-video, reference image guidance, video editing / video reference
Output aspect ratios
16:9, 9:16; default 16:9
Output resolutions
720p, 1080p; default 720p
Image references
Submit one or more image links via image_urls
Video references
Up to 1 video link; at least 1 reference image must also be provided
Task delivery
Supports asynchronous polling and completion callbacks; returns video_url upon success

The above specifications are within this platform's omni-flash invocation scope; extended features of different native Gemini Omni versions must be distinguished separately.

Core Capabilities

Learn what omni-flash can bring to your work.

Turn camera intent into actionable descriptions

When starting from text, you can specify the subject, scene, action, camera movement, and lighting mood in the prompt to generate video footage for review. Compared with writing only a theme, this is better suited to fully expressing shooting intent: for example, the subject slowly turns around, the camera pushes forward, and the environment maintains soft backlighting, bringing creative discussions down to specific shots.

Bring static assets into video creation

By guiding generation with reference images, you can continue creating dynamic footage from product photos, character images, or illustrations. Images provide the visual basis, while text describes the direction of movement, scene relationships, and desired style. This is suitable for production stages where design assets already exist but video shots have not yet been formed; clearly state whether the image serves as a subject reference or an overall visual reference.

Define edit goals around the original video

Existing videos can also serve as a starting point for creation, generating new videos with reference images and editing instructions. You can request changes to the season, adjustments to the scene style, or additions and removals of visual elements, while specifying which compositional relationships should be retained. This workflow is suitable for creating creative variations around the same asset, rather than redescribing the entire scene each time.

Use Cases

Start with specific tasks to find where the model can make an impact.

Turn product images into dynamic showcases

Input selected product images, describe the product action, camera direction, and background atmosphere, and create showcase videos for the marketing team to review. Landscape format is suitable for pages or presentation visuals, while portrait format is suitable for vertical content layouts. After delivery, focus on checking whether the product outline, composition, and motion match the creative intent before deciding whether to proceed to formal editing.

Visual variations of the same scene

Input an existing scene video and reference images, and use text to request changes to the weather, season, or visual style while listing the positional relationships that need to be retained. For example, turn a sunny beach into a winter scene while preserving the layout of the trees and boats. The output can be used to compare new versions across different creative directions, reducing the work of reimagining from a blank frame.

Integrate into automated asset production

The application backend submits prompts, aspect ratios, and asset links, creates tasks with async, saves the task_id and then queries the result, or receives completion notifications through callback_url. After obtaining a successful status and video link, download and archive it, then pass it to the review or editing workflow. This is suitable for production applications that do not want to continuously maintain a generation request connection.

How to Choose This Model

Choose based on task complexity, input materials, and expected results.

Choose a Creation Method Based on Existing Materials

Choose text-to-video when you only have creative text; add image references when the subject's appearance or visual direction has already been determined; use a combination of video and images when you want to modify existing shots. The value of choosing omni-flash is that it brings these material formats into the same video workflow. For editing tasks especially, clearly state separately “what to change” and “what to preserve,” rather than submitting only vague style terms.

Distinguish the Video Endpoint from Native Versions

omni-flash is for video generation and editing and should not be confused with Gemini Flash text chat models. Gemini Omni 1.1 Flash also has a separate native version name, so do not apply first-and-last-frame or extension controls directly to this endpoint just because the names are similar. If your goal is guidance from existing materials and scene rewriting, choose the workflow here; if you depend on specific version features, select models by version.

Get Started

From a small-scale task to formal integration.

01

Prepare the Task and Materials

Define the goal, required inputs, and output requirements, and use real business examples as a starting point.

02

Try It in the API Debugging Area

Open the trial page, confirm the parameters supported by this endpoint, then submit a small-scale task to review the results.

03

Integrate According to the API Documentation

Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.

Usage Limits

Before formal use, understand output quality and capability boundaries.

  • Video editing does not work by submitting only a video: video_urls accepts at most one link, and you must also provide at least one image_urls. Reference images should be relevant to the modification target, and the prompt should clearly specify the editing effect; missing images will cause a request parameter error and prevent entry into the normal editing workflow.
  • You can choose an output resolution of 720p or 1080p, but selecting a resolution does not guarantee that subject details and actions will meet requirements. Especially for scene replacement or adding and removing elements, review each item of the layout and appearance that needs to be preserved; do not treat textual instructions to “keep unchanged” as a promise of no frame-by-frame changes.
  • An asynchronous submission returning task_id does not mean the video has been completed. You need to wait for the succeeded status before obtaining the final video; the video link may be empty during the pending stage. Generated links have a retention period, so download them promptly to your own storage after completion to avoid using temporary result links as long-term asset addresses.

Frequently Asked Questions

Answers to common questions about using omni-flash.

Is omni-flash a standard Gemini conversational model?

No. The omni-flash here is designed for video creation: submit text and optional assets through POST /gemini/videos to receive video results. It is not a Gemini Flash model that returns regular text answers through a conversational interface; during development, organize inputs, status queries, and result storage around video tasks.

Can I generate a video with just one image?

Yes. You can submit an image link through image_urls and also provide the required prompt. We recommend clearly describing how the subject should move, how the camera should move, and which visual characteristics should be maintained. Images are used to guide generation; do not simply write “make it move,” or it will be difficult to convey the shot effect you actually need.

What assets are needed to modify an existing video?

You need to submit a video link in video_urls, provide at least one reference image in image_urls, and use prompt to describe the desired modifications. You can make requests around style, scene, or visual elements, and should also specify the layout that needs to be preserved. Sending only a video without an image does not meet this workflow's input requirements.

Can I choose portrait orientation and 1080p output?

Yes. aspect_ratio supports 16:9 and 9:16, and resolution supports 720p and 1080p, with defaults of 16:9 and 720p respectively. We recommend determining the delivery aspect ratio before creation, especially the subject position and camera movement direction, to avoid cropping after generation that causes the composition to deviate from the original intent.

How do I retrieve a video after asynchronous generation?

After setting async to true, save the returned task_id and submit a query request to /gemini/tasks using that value as the id; you can also set callback_url to receive completion notifications. After the task succeeds, read the video_url in the result and download and save it. Continue waiting if it is still pending, and check the error information if it fails.

Model information · Updated: 2026-10-01. For request parameters and billing rules, see the API and pricing sections.

Use omni-flash for your next task

Start with a clear goal and judge from actual results whether it suits your work.