MinimaxMiniMax Omni-modal video generation

MiniMax H3

Bring reference media and sound into one video brief.

MiniMax H3 is an omni-modal video route in TokenHot's catalog. It accepts text, image, audio and video references and produces video with synchronized audio described in the catalog. Use it for short-form video, advertising and reference-driven character or camera work, not as a general chat model.

  • Text + image + audio + video
  • Video output
  • Synchronized audio route

The route is video-specific and uses the openai-video protocol entry. Verify reference shapes, duration, resolution and audio behavior in the API response; do not infer a finished clip from the prompt alone.

USD /s

Pricing

GroupTierPrice
default768P$0.08
default2K$0.13
default
768P$0.08
2K$0.13

Public list prices are shown in USD. Final charges may vary by account group and usage tier.

Overview

What is MiniMax H3?

The best H3 tasks have a clear shot, reference role and review rubric. State whether an input controls identity, motion, timing or sound, then check the resulting audio-visual alignment.

Keep 768P and 2K settings explicit because the catalog billing expression distinguishes them. Compare accepted takes, not just the first generation, before selecting a workflow.

Omni-modal direction

Assign each reference a job

Say what the image, video or audio reference should influence so the result can be diagnosed.

Audio-visual quality

Review the two tracks together

Check synchronization, intelligibility, motion and continuity as one deliverable.

Video workflow

Design for async review

Store the job, parameters and output metadata, then pass accepted clips to an editor.

From brief to deliverable

Put MiniMax H3 to work

Start with a concrete task and decide what a useful result looks like.

Creative and advertising teams

Build a reference-driven ad shot

Use a product image, a motion reference and a sound cue to explore one short advertising shot.

Tools you need
Video generation, asset storage, player and design review.
Keep in mind
Reference inputs guide a take but do not guarantee exact product geometry or brand copy.
View a task brief

Use the supplied reference image or clip, motion direction, audio cue, resolution and duration. to plan build a reference-driven ad shot. Return a reviewable ad take with an audio-visual checklist., state the evidence limits and list the checks required before accepting the result. Do not claim an output or interaction was verified without direct evidence.

Directors and content teams

Test a character or camera move

Describe one visible action and one camera move while preserving the reference identity.

Tools you need
Reference upload, generation queue and an editor for assembly.
Keep in mind
One clip cannot establish multi-shot character continuity.
View a task brief

Use the supplied character reference, shot brief, camera move and target ending. to plan test a character or camera move. Return a motion study with continuity notes., state the evidence limits and list the checks required before accepting the result. Do not claim an output or interaction was verified without direct evidence.

Social and short-form teams

Prototype a sound-aware beat

Coordinate a short visual action with a spoken line or sound cue, then review timing and intelligibility.

Tools you need
Audio review, captions and an editing workflow.
Keep in mind
Pronunciation, rights and final mix need a separate review.
View a task brief

Use the supplied visual brief, dialogue or sound cue, target duration and rejection rules. to plan prototype a sound-aware beat. Return a selected take and a timing report., state the evidence limits and list the checks required before accepting the result. Do not claim an output or interaction was verified without direct evidence.

Capabilities in context

Key features

FeatureWhat you getWhy it matters
Input rangeText, image, audio and video inputs are listedAttach each reference with an explicit role and verify media limits before sending.
OutputVideo with synchronized audio is described in the catalogReview picture and sound together and retain output metadata.
Resolution tiersBilling exposes 768P and 2K labelsChoose the resolution deliberately and use the dynamic table for cost estimation.
ProtocolTokenHot lists an openai-video endpointUse the existing API docs and do not substitute a chat request shape.
Model boundaryVideo-specific route, not general chatKeep captions, editing, retrieval and other text tasks on appropriate text routes.

The page uses TokenHot catalog metadata and its current video billing expression. Verify reference, resolution, duration and audio behavior on the selected endpoint.

Audience & fit

Where MiniMax H3 fits

Short-form and advertising teams

Use H3 to explore reference-driven clips where audio and visual timing should be reviewed together.

Directors and content creators

Prototype one shot or beat at a time, then assemble accepted takes in an editor.

Developers building media pipelines

Treat the route as an asynchronous job and preserve reference metadata, status and output details.

Cost context

Plan the cost of a useful result

Count resolution and reruns

Resolution tiers and rejected takes affect spend. Budget review, storage and editing alongside generation.

Separate audio finishing

Mixing, captions, rights checks and loudness normalization remain outside the model call.

Keep the job record

Store inputs, parameters, model ID and output status so a selected take can be reproduced or audited.

Use the current pricing table for available rates and account groups. Model IDs and prices come from TokenHot’s catalog; provider pricing and subscriptions are separate.

API

Code examples

curl --location --request POST 'https://api.tokenhot.ai/v1/video/generations' \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data-raw '{
    "model": "MiniMax-H3",
    "prompt": "A cat is playing on the grass."
}'

FAQ

Frequently asked questions

What is MiniMax H3 best suited for?

MiniMax H3 is suited to the input, output, and capability types listed on this page. Test production workloads before choosing it for a critical system.

How is MiniMax H3 priced?

Pricing is shown in the pricing section above. Actual cost depends on usage volume, account group, and the request payload.

How can I call MiniMax H3?

Use the model ID shown above with one of the supported API protocols. Code examples are provided when usage data is available.

Is MiniMax H3 a chat model?

No. The catalog identifies it as an omni-modal video route. Use a text model for general conversation and H3 for video generation.

Can H3 use audio and video references?

The catalog lists text, image, audio and video inputs. Confirm exact media encoding and limits in the API contract.

Does synchronized audio guarantee a final mix?

No. Review timing and intelligibility, then complete mixing, captions and rights checks in your editing workflow.

Why does the price change with resolution?

The current billing expression exposes separate 768P and 2K labels. Use the dynamic table and keep resolution beside the job estimate.

Get started

Build with MiniMax H3

Use in Console