Wan2.7 Reference-to-Video
POST /api/v1/services/aigc/video-generation/video-synthesis — Wan2.7 reference-to-video (identity / voice consistency)
The Wan 2.7 reference-to-video model accepts multimodal references (images, video, audio) and produces a video that keeps subject identity and voice consistent — single-character performance or multi-character interaction.
Models
| Model | Notes |
|---|---|
wan2.7-r2v | Wan 2.7 reference-to-video |
Create a job
/api/v1/services/aigc/video-generation/video-synthesiscurl https://api.modelsite.ai/api/v1/services/aigc/video-generation/video-synthesis \
-H "Authorization: Bearer $MODELSITE_API_KEY" -H "Content-Type: application/json" \
-d '{
"model": "wan2.7-r2v",
"input": {
"prompt": "The character from Video 1 holds Figure 2 and plays a gentle folk tune; Figure 1 walks past and puts down what she carries",
"media": [
{ "type": "reference_image", "url": "https://example.com/girl.jpg",
"reference_voice": "https://example.com/girl-voice.mp3" },
{ "type": "reference_video", "url": "https://example.com/role2.mp4" },
{ "type": "reference_image", "url": "https://example.com/prop.png" }
]
},
"parameters": { "resolution": "720P", "ratio": "16:9", "duration": 10 }
}'Request body
| Field | Type | Notes |
|---|---|---|
model | string (required) | wan2.7-r2v |
input.prompt | string (required) | ≤5000 characters. Refer to media entries as "Figure 1 / Video 1" (images and videos count separately, in array order); with a single asset, "the reference image" also works |
input.negative_prompt | string (optional) | ≤500 characters (in input) |
input.media | array (required) | reference_image + reference_video together 1–5 entries; optionally plus at most 1 first_frame for joint control |
parameters | object (optional) | Below |
media entry fields
| Field | Notes |
|---|---|
type | reference_image (subject/scene) / reference_video (subject + voice; avoid empty shots) / first_frame (joint control, ≤1) |
url | Asset URL or Base64. Images ≤20MB, sides [240, 8000]px; video mp4/mov, 1–30 s, ≤100MB |
reference_voice | Only on reference entries: a voice-reference audio URL (wav/mp3, 1–10 s, ≤15MB). Voice only, independent of content; takes precedence over the source audio of a reference_video |
parameters
| Parameter | Type | Notes |
|---|---|---|
resolution | string | 720P / 1080P (default). Sellable tiers follow the price configuration |
ratio | string | 16:9 (default) / 9:16 / 1:1 / 4:3 / 3:4. Ignored when first_frame is present (follows the frame) |
duration | integer | Default 5. [2, 10] when a reference video is present, [2, 15] otherwise |
prompt_extend | boolean | Default true |
watermark | boolean | Default false |
seed | integer | [0, 2147483647] |
Keep a single character per reference asset. Billed duration = input video duration + output duration (the poll's usage.duration).
Polling and result
/api/v1/tasks/{task_id}curl "https://api.modelsite.ai/api/v1/tasks/$TASK_ID" \
-H "Authorization: Bearer $MODELSITE_API_KEY"Response fields
| Field | Meaning |
|---|---|
output.task_id | Job ID (same value returned at creation) |
output.task_status | PENDING queued / RUNNING processing / SUCCEEDED done / FAILED failed |
output.video_url | The generated video URL — only on SUCCEEDED; download promptly |
output.code / output.message | Present only on failure, with the reason |
usage | Usage stats (duration, resolution tier, …); counted only on success |
request_id | Unique request ID — include it when reporting issues |
Status flow: PENDING → RUNNING → SUCCEEDED / FAILED.
Billing
Billed by actual generated seconds — with a reference video, its duration counts too; failed jobs are not billed. Rates on the Models page.
Error handling
Creation-time errors return {"code": "...", "message": "..."}; mid-job failures surface through the poll's output.code / output.message. General error codes: Errors.