MiniMax Video Generation
POST /v2/video_generation — MiniMax-H3 native video protocol (text / image / reference to video)
MiniMax-H3 generates video from text, images (first frame / last frame / first+last), or multimodal references. This family uses its own native entrance — the request shape differs from the /api/v1 families.
Models
| Model | Notes |
|---|---|
MiniMax-H3 | MiniMax H3 (model name is case-insensitive; minimax-h3 works too) |
Create a job
/v2/video_generationcurl https://api.modelsite.ai/v2/video_generation \
-H "Authorization: Bearer $MODELSITE_API_KEY" -H "Content-Type: application/json" \
-d '{
"model": "MiniMax-H3",
"content": [{ "type": "text", "text": "Epic space-opera trailer: a captain watches her fleet jump away through the observation window" }],
"resolution": "768P",
"ratio": "16:9",
"duration": 5
}'Request body (OpenAI-multimodal-style content[], everything flat at the top level)
| Field | Type | Notes |
|---|---|---|
model | string (required) | MiniMax-H3 (case-insensitive) |
content | array (required) | Content items, below. A non-empty text item is required (the prompt) |
resolution | string (required) | 768P (standard) / 2K (high definition) |
duration | integer (required) | [4, 15] seconds |
ratio | string | Aspect ratio; the rule depends on the job — below |
content items
| type | Fields | Notes |
|---|---|---|
text | text | The prompt (required; the first text item wins) |
image_url | image_url.url + role | Image: role is first_frame / last_frame (≤1 each; defaults to first frame when omitted) or reference_image (≤9). Frames and references are mutually exclusive |
video_url | video_url.url | Reference video (≤3 clips, 2–15 s each, ≤15 s total, ≤50MB) |
audio_url | audio_url.url | Reference audio (≤3 clips, 2–15 s each, ≤15MB) |
Image limits: JPG/JPEG/PNG/WEBP/HEIC/HEIF, sides [256, 5760]px, aspect 0.4–2.5, ≤30MB.
ratio rules (per job)
| Job | Rule |
|---|---|
| Text to video | Required, and must not be adaptive: 21:9 / 16:9 / 4:3 / 1:1 / 3:4 / 9:16 |
| Image to video | Always adaptive (follows the input image); other values are ignored |
| Multimodal reference | Optional, default adaptive; an explicit ratio is honored |
Other vendor-native parameters (e.g. callback_url) go flat at the top level and pass through verbatim.
Create response
{ "task_id": "enc_v2_xxxxx" }Polling
/v2/query/video_generation/{task_id}curl "https://api.modelsite.ai/v2/query/video_generation/$TASK_ID" \
-H "Authorization: Bearer $MODELSITE_API_KEY"{
"task": {
"id": "enc_v2_xxxxx",
"status": "succeeded",
"content": { "url": "https://.../video.mp4" }
}
}| Field | Notes |
|---|---|
task.id | Job ID (the token returned at creation; pass it back verbatim) |
task.status | Lowercase: queued / running / succeeded / failed |
task.content.url | Video URL, only on success; download promptly |
task.error | Only on failure: {"code": "...", "message": "..."} |
This entrance has no delete endpoint — a job cannot be cancelled once created; wait for its terminal state.
Billing
Billed by actual generated seconds; 768P and 2K have different rates; failed jobs are not billed. Rates on the Models page.
Error handling
This entrance uses the MiniMax-native error shape, not the /api/v1 one:
{
"type": "error",
"error": { "type": "bad_request_error", "message": "...", "http_code": "400" }
}Creation-time parameter errors return 400 + bad_request_error. Platform-level errors (auth, rate limits, balance) are covered in Errors.