MiniMax Omni-modal video generation
MiniMax H3
Bring reference media and sound into one video brief.
MiniMax H3 is an omni-modal video route in TokenHot's catalog. It accepts text, image, audio and video references and produces video with synchronized audio described in the catalog. Use it for short-form video, advertising and reference-driven character or camera work, not as a general chat model.
- Text + image + audio + video
- Video output
- Synchronized audio route
The route is video-specific and uses the openai-video protocol entry. Verify reference shapes, duration, resolution and audio behavior in the API response; do not infer a finished clip from the prompt alone.
USD /s
Pricing
| Group | Tier | Price |
|---|---|---|
| default | 768P | $0.08 |
| default | 2K | $0.13 |
Public list prices are shown in USD. Final charges may vary by account group and usage tier.
Overview
What is MiniMax H3?
The best H3 tasks have a clear shot, reference role and review rubric. State whether an input controls identity, motion, timing or sound, then check the resulting audio-visual alignment.
Keep 768P and 2K settings explicit because the catalog billing expression distinguishes them. Compare accepted takes, not just the first generation, before selecting a workflow.
Assign each reference a job
Say what the image, video or audio reference should influence so the result can be diagnosed.
Review the two tracks together
Check synchronization, intelligibility, motion and continuity as one deliverable.
Design for async review
Store the job, parameters and output metadata, then pass accepted clips to an editor.
From brief to deliverable
Put MiniMax H3 to work
Creative and advertising teams
Build a reference-driven ad shot
Use a product image, a motion reference and a sound cue to explore one short advertising shot.
- Tools you need
- Video generation, asset storage, player and design review.
- Keep in mind
- Reference inputs guide a take but do not guarantee exact product geometry or brand copy.
View a task brief
Use the supplied reference image or clip, motion direction, audio cue, resolution and duration. to plan build a reference-driven ad shot. Return a reviewable ad take with an audio-visual checklist., state the evidence limits and list the checks required before accepting the result. Do not claim an output or interaction was verified without direct evidence.
Directors and content teams
Test a character or camera move
Describe one visible action and one camera move while preserving the reference identity.
- Tools you need
- Reference upload, generation queue and an editor for assembly.
- Keep in mind
- One clip cannot establish multi-shot character continuity.
View a task brief
Use the supplied character reference, shot brief, camera move and target ending. to plan test a character or camera move. Return a motion study with continuity notes., state the evidence limits and list the checks required before accepting the result. Do not claim an output or interaction was verified without direct evidence.
Social and short-form teams
Prototype a sound-aware beat
Coordinate a short visual action with a spoken line or sound cue, then review timing and intelligibility.
- Tools you need
- Audio review, captions and an editing workflow.
- Keep in mind
- Pronunciation, rights and final mix need a separate review.
View a task brief
Use the supplied visual brief, dialogue or sound cue, target duration and rejection rules. to plan prototype a sound-aware beat. Return a selected take and a timing report., state the evidence limits and list the checks required before accepting the result. Do not claim an output or interaction was verified without direct evidence.
Capabilities in context
Key features
| Feature | What you get | Why it matters |
|---|---|---|
| Input range | Text, image, audio and video inputs are listed | Attach each reference with an explicit role and verify media limits before sending. |
| Output | Video with synchronized audio is described in the catalog | Review picture and sound together and retain output metadata. |
| Resolution tiers | Billing exposes 768P and 2K labels | Choose the resolution deliberately and use the dynamic table for cost estimation. |
| Protocol | TokenHot lists an openai-video endpoint | Use the existing API docs and do not substitute a chat request shape. |
| Model boundary | Video-specific route, not general chat | Keep captions, editing, retrieval and other text tasks on appropriate text routes. |
The page uses TokenHot catalog metadata and its current video billing expression. Verify reference, resolution, duration and audio behavior on the selected endpoint.
Audience & fit
Where MiniMax H3 fits
Short-form and advertising teams
Use H3 to explore reference-driven clips where audio and visual timing should be reviewed together.
Directors and content creators
Prototype one shot or beat at a time, then assemble accepted takes in an editor.
Developers building media pipelines
Treat the route as an asynchronous job and preserve reference metadata, status and output details.
Cost context
Plan the cost of a useful result
Count resolution and reruns
Resolution tiers and rejected takes affect spend. Budget review, storage and editing alongside generation.
Separate audio finishing
Mixing, captions, rights checks and loudness normalization remain outside the model call.
Keep the job record
Store inputs, parameters, model ID and output status so a selected take can be reproduced or audited.
Use the current pricing table for available rates and account groups. Model IDs and prices come from TokenHot’s catalog; provider pricing and subscriptions are separate.
API
Code examples
curl --location --request POST 'https://api.tokenhot.ai/v1/video/generations' \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data-raw '{
"model": "MiniMax-H3",
"prompt": "A cat is playing on the grass."
}'FAQ
Frequently asked questions
What is MiniMax H3 best suited for?
MiniMax H3 is suited to the input, output, and capability types listed on this page. Test production workloads before choosing it for a critical system.
How is MiniMax H3 priced?
Pricing is shown in the pricing section above. Actual cost depends on usage volume, account group, and the request payload.
How can I call MiniMax H3?
Use the model ID shown above with one of the supported API protocols. Code examples are provided when usage data is available.
Is MiniMax H3 a chat model?
No. The catalog identifies it as an omni-modal video route. Use a text model for general conversation and H3 for video generation.
Can H3 use audio and video references?
The catalog lists text, image, audio and video inputs. Confirm exact media encoding and limits in the API contract.
Does synchronized audio guarantee a final mix?
No. Review timing and intelligibility, then complete mixing, captions and rights checks in your editing workflow.
Why does the price change with resolution?
The current billing expression exposes separate 768P and 2K labels. Use the dynamic table and keep resolution beside the job estimate.
Get started
Build with MiniMax H3
API