MiniMax/minimax-h3
MiniMax H3 is a general-purpose omni-modal generation model with unified understanding of text, image, video, and audio context, native stereo audio-video output, strong instruction following, and high-quality 2K video generation.
Output: $0.12 / use or 8 uses / $1
Input
text * string
Required text description of the video to generate.
first_frame image
Optional first-frame image used to guide the opening frame and visual content.
last_frame image
Optional last-frame image used to guide the ending frame and visual content.
resolution * enum
Generated video resolution.
duration * int
Generated video duration in seconds. Enter an integer from 4 through 5.
ratio enum
Generated video aspect ratio. Adaptive selects the most suitable ratio from the input; the actual ratio is available in the query response's ratio field.
Reset
Output
{
  "task_id": "j2n94b12anrmw0cwdf1bpjde6r",
  "user_id": 1,
  "version": "c98270c13de8ee8e86597b77ec3493956f1bd8b8ea4d0021bb0641541babbd3b",
  "error": null,
  "total_time": 240,
  "predict_time": 240,
  "logs": null,
  "output": [
    "https://vmodel.ai/data/dev/model/minimax/minimax-h3/121061af-e563-4bc7-9b58-b9bddc2ff9b3.mp4"
  ],
  "status": "succeeded",
  "create_at": 1771334410,
  "completed_at": 1771334960,
  "input": {
    "text": "A cinematic product shot of a silver sports car driving through a neon-lit city at night, reflections moving across the bodywork, dynamic tracking camera.",
    "resolution": "768P",
    "duration": 5,
    "ratio": "16:9"
  }
}
Generated in: 240 seconds
Download
Examples
Pricing
Model pricing for minimax/minimax-h3. Looking for volume pricing? Get in touch.
When
model variant is 768p
$0.120000
per generation at 768P
or 8 generations for $1
When
model variant is 2k
$0.180000
per use
or 5 uses for $1
Readme

Loading...