Google Any-input video generation
Gemini Omni Video
Bring text, images, audio and video references into one video workflow.
Gemini Omni Video is a multimodal video route in TokenHot's catalog. It accepts text, images, audio and video, returns video and uses a resolution-and-duration expression with 720p, 1080p and 4K conditions. Use it when a creative brief needs several reference types coordinated in one shot, while keeping job parameters and review evidence explicit.
- Text + image + audio + video
- Video output
- Resolution and duration tiers
The catalog expression changes with input video, resolution and duration. The route does not remove the need for media preparation, rights review or frame and audio inspection; verify the endpoint contract before estimating a deliverable.
Try this model
Choose an aspect ratio, resolution, and duration to generate a video.
gemini-omni-videoInput
Generated result
Enter a prompt, choose your settings, and generate.
USD /s
Pricing
| Group | Tier | Price |
|---|---|---|
| default | 4S/1080p | $0.495 |
| default | 4S/720p | $0.495 |
| default | 6S/1080p | $0.66 |
| default | 6S/720p | $0.66 |
| default | 8S/1080p | $0.825 |
| default | 8S/720p | $0.825 |
| default | 基础价格 | $1 |
| default | 10S/1080p | $1 |
| default | 10S/720p | $1 |
| default | 4S/4K | $1.15 |
| default | 720p/input_video | $1.32 |
| default | 1080p/input_video | $1.32 |
| default | 6S/4K | $1.32 |
| default | 8S/4K | $1.5 |
| default | 10S/4K | $1.65 |
| default | 4K/input_video | $2 |
Public list prices are shown in USD. Final charges may vary by account group and usage tier.
Overview
What is Gemini Omni Video?
Omni Video is most useful when the references have distinct roles: an image for identity, an audio clip for timing, a video for movement and text for direction. Declare those roles and the one property the model may change.
The pricing expression includes several duration and resolution factors. Keep a job manifest with the chosen resolution, seconds and input-video flag so estimates can be explained after the fact.
Assign each reference a role
State what must be preserved from the image, audio or video and what the text brief controls.
Write time-coded beats
Describe action, camera, audio cue and final hold in a short, reviewable sequence.
Inspect visual and audio tracks together
Review frame continuity, timing, intelligibility, rights and export settings before delivery.
From brief to deliverable
Put Gemini Omni Video to work
Creative production teams
Compose a multimodal shot
Combine supplied image, audio and video references into a one-shot creative brief with explicit invariants.
- Tools you need
- Media preprocessing, generation queue, player and editorial review.
- Keep in mind
- Multiple references can conflict; the model cannot decide rights or hidden source intent.
View a task brief
Map each supplied reference to a role in one bounded shot. Return a prompt, source-role map and review rubric. Flag conflicts instead of silently choosing one.
Product and marketing teams
Prototype a product demonstration
Create a short product beat with a reference image, motion cue and sound direction while preserving brand invariants.
- Tools you need
- Reference upload, generation, player and brand review.
- Keep in mind
- A generated product video is not evidence that the product works or that copy is legally approved.
View a task brief
Plan a product demonstration from the supplied references. Return action, camera, audio and brand QA checks, and keep functional claims outside the generated shot.
Editors and content teams
Remix a reference clip
Use a video reference and audio or image cue to explore one controlled transformation.
- Tools you need
- Frame extraction, media storage, generation and editor review.
- Keep in mind
- Input references guide output but do not guarantee timing, identity or geometry across every frame.
View a task brief
Plan a controlled remix from the supplied clip and references. State the invariant list, frame checkpoints and export validation; do not claim the remix rendered.
Capabilities in context
Key features
| Feature | What you get | Why it matters |
|---|---|---|
| Input coverage | Text, image, audio and video input are listed | Keep a source manifest with role, encoding, rights and duration for every reference. |
| Video output | The route returns video | Treat generation as an asynchronous job and retain status and output metadata. |
| Resolution tiers | 720p, 1080p and 4K labels appear in the expression | Choose resolution explicitly and use the live table for estimates. |
| Duration tiers | 4, 6, 8 and 10 second labels appear | Keep duration alongside the prompt because it changes the catalog expression. |
| Video-input factor | The expression changes when input video is present | Record whether video references are included and validate the endpoint request before billing estimates. |
The route identity, modalities and pricing conditions come from TokenHot's current catalog expression. Verify supported durations, reference formats, audio behavior and effective rates on the endpoint.
Audience & fit
Where Gemini Omni Video fits
Creative teams with mixed references
Use a source-role map to coordinate text, images, audio and clips in one shot brief.
Product and marketing teams
Prototype product beats while separating generated visuals from functional, legal and rights claims.
Editors and content platforms
Keep job parameters and frame/audio review notes attached to every selected take.
Cost context
Plan the cost of a useful result
Keep duration and resolution visible
The expression has several conditions; a quote without those parameters is incomplete.
Price media preparation
Transcription, frame extraction, storage and rights checks add cost before and after generation.
Budget rejected takes
Mixed-reference workflows need review and reruns. Measure cost per accepted shot, not just per call.
Use the current pricing table for available rates and account groups. Model IDs and prices come from TokenHot’s catalog; provider pricing and subscriptions are separate.
API
Code examples
curl --location 'https://api.tokenhot.ai/v1/video/generations' \
--header 'Content-Type: application/json' \
--data '{
"model": "gemini-omni-video",
"input": {
"prompt": "Slow push-in as the character walks out from a neon-lit street",
"image_urls": [
"https://your-host.com/scene1.png"
],
"duration": "6",
"aspect_ratio": "16:9",
"resolution": "720p"
}
}'FAQ
Frequently asked questions
What is Gemini Omni Video best suited for?
Gemini Omni Video is suited to the input, output, and capability types listed on this page. Test production workloads before choosing it for a critical system.
How is Gemini Omni Video priced?
Pricing is shown in the pricing section above. Actual cost depends on usage volume, account group, and the request payload.
How can I call Gemini Omni Video?
Use the model ID shown above with one of the supported API protocols. Code examples are provided when usage data is available.
What inputs does Gemini Omni Video accept?
The catalog lists text, image, audio and video input with video output. Confirm exact formats and limits in the current endpoint.
How does the route price video?
The expression varies by input-video presence, resolution and duration. Use the dynamic pricing table for the selected combination.
Does it automatically synchronize every reference?
No. Assign each source a role, review timing and inspect conflicts. Reference inputs guide the model but do not guarantee a usable final mix.
What should I save with a job?
Store the model ID, source-role map, resolution, duration, input-video flag, status, output metadata and review notes.
Get started
Build with Gemini Omni Video
API