MiniMax H3 · Multimodal Video Generation API

MiniMax H3 Video Generation

Generate 4–15 second videos from text, up to 9 reference images, or up to 3 reference audio clips with a single REST API. Unified asynchronous tasks, callbacks, billing, and CDN delivery integrate directly into your content product.

Text-to-video / Image-to-video / Audio-guided768P / 2KAsynchronous Tasks + Webhook
🎬
3
Video Generation Modes
⏱️
4-15s
Available Video Duration
🖼️
1-9
Reference Images
🎵
1-3
Reference Audio Clips

One API, Three Creative Starting Points

No action or multiple paths required; the API automatically determines the generation mode based on the content assets.

✍️

Text-to-Video

Generate short videos directly from descriptions of scenes, actions, camera shots, and styles.

🖼️

First-and-Last-Frame Image-to-Video

Control the opening and ending visuals with first_frame and last_frame.

🎵

Multimodal References

Combine reference images, videos, and audio to guide subjects, actions, sound, and rhythm.

📡

Asynchronous Delivery

Receive a task ID immediately, then get the final CDN video through polling or callback_url.

curl
curl -X POST 'https://api.acedata.cloud/minimax/videos' \
  -H 'Authorization: Bearer YOUR_API_KEY' \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "MiniMax-H3",
    "content": [{
      "type": "text",
      "text": "A red fox running through a snowy forest at dawn, cinematic tracking shot"
    }],
    "resolution": "2K",
    "ratio": "16:9",
    "duration": 5
  }'

Start Generating with Just a Few Lines of Code

Parameters directly express creative intent, while the platform handles long-running task execution, result storage, and precise billing.

1

Build content

All modes must provide text; images, videos, and audio indicate their purpose through their corresponding types and roles.

2

Submit an Asynchronous Task

The creation API is always asynchronous and returns a task_id immediately upon success.

3

Get Permanent Results

Retrieve the Ace Data Cloud CDN video from task.content.url via /minimax/tasks or a callback.

Typical scenarios from source material to finished video

Three real API samples cover product activation, character storytelling, and audio-visual reference workflows.

Commercial video shot of a silver fragrance device amid blue-purple lighting and water reflections
MiniMax-H3 real sample representative frame, 16:9 first-frame image-to-video

Brand and e-commerce shorts

Create commercial shorts from product visuals and shot descriptions, ideal for product showcases, launch assets, and social media ads.

MiniMax-H3 · 16:9 · First-frame image-to-video
A fictional courier wearing an amber jacket and carrying a teal shoulder bag stands at a train station on a rainy night
MiniMax-H3 real sample representative frame, 16:9 multi-image reference video

Consistent character storytelling

Use multiple proprietary reference images to improve consistency in characters, clothing, props, and scene styles, ideal for continuous shots and serialized content.

MiniMax-H3 · 16:9 · Multi-image reference
Vertical video shot of a fictional dancer performing in front of a glowing teal and amber doorway
MiniMax-H3 real sample representative frame, 9:16 video and audio reference

Audio-visual rhythm shorts

Combine proprietary motion video and audio references to guide dance movements, camera motion, and visual rhythm, ideal for vertical content creation.

MiniMax-H3 · 9:16 · Video and audio reference

Go live in three steps

Reuse Ace Data Cloud's unified authentication and task system without building a separate media pipeline.

01

Get a Token

Create an application in the console and obtain a unified Bearer Token.

02

Submit generation

Choose the duration and aspect ratio, then add text, image, or audio assets.

03

Deliver video

Poll the task or receive a callback, then integrate the CDN video into your product workflow.

Unified Video API for Production

Consolidate task execution, error semantics, usage records, and file delivery into a stable contract.

CapabilityAce Data Cloud MiniMax H3Some Others
Multimodal inputUnified entry point for text, images, video, and audioMultiple APIs to assemble yourself
Long-running tasksPolling and WebhookNeed to build your own task queue
Result filesAutomatically stored on the platform CDNInconsistent link lifecycles
Failure billingNo charge for failuresNeed to reconcile bills yourself

Select a workflow by content

The model is fixed as MiniMax-H3, and the mode is automatically inferred from the official V2 multimodal content array.

TEXT

text

Pass text content only to generate visuals directly from ideas, scripts, and shot descriptions.

FRAME

first_frame / last_frame

Use the first frame or the first and last frames to control the video's opening and ending visuals.

REFERENCE

reference media

Combine reference images, video, and audio to control the subject, motion, sound, and rhythm.

TASK

task_id / callback

Create asynchronous tasks and retrieve results through task queries or callbacks.

Transparent resolution pricing

Billed by video seconds, with charges only for successful tasks.

768P

$0.057143 / second
  • As low as $0.228572 for 4 seconds
  • Ideal for rapid creative validation
  • Three input modes
View pricing

Failed tasks

$0 / time
  • No charge for failures
  • Free task queries
  • CDN result delivery
View parameters

Frequently Asked Questions

Key notes on modes, assets, duration, and task execution.

Do I need to pass action?

No. When content only has text, it is text-to-video; first_frame / last_frame are image-to-video; reference_* are multimodal reference-to-video. All modes require non-empty text.

What aspect ratios and durations are supported?

Resolution is required, with support for 768P and 2K; supported ratios include adaptive, 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16; duration is an integer from 4–15 seconds.

How do I get the results?

The creation API is always asynchronous and immediately returns task_id; use /minimax/tasks to query it. Upon success, get the video from task.content.url, or provide callback_url.

Are failed tasks charged?

No. Only tasks that successfully complete and return video results are billed based on the final duration.

Integrate MiniMax H3 into Your Product

One Token, one generation endpoint, one task system—start building multimodal video experiences.