Can wan2.6-t2v generate videos directly from images?
This model is designed for text-to-video generation. If creation requires an existing image as the starting frame, choose wan2.6-i2v; if you need to continue a character's appearance from a video, choose wan2.6-r2v. Selecting the appropriate model based on the source material helps clarify how the visuals are controlled.
How long and at what resolution can it generate videos?
wan2.6-t2v has a native maximum duration of 15 seconds, with 720P and 1080P resolution options. It is suitable for creating short videos focused on a single subject. For longer stories, it is recommended to split them into multiple clips, then connect them through editing while maintaining the same subject and scene settings for each segment.
How should prompts for multiple shots be written?
Write the prompt as a brief shooting outline: first describe the subject and scene, then describe the opening, main action, and ending in order. Keep the character's appearance, environment, and atmosphere consistent, and avoid redefining the subject for each shot. When multiple shots are needed, you can select shot_type as multi and review the generated continuity.
Do generated videos include sound by default?
audio defaults to false in requests on this platform, so audio creation should explicitly set audio to true. This model natively supports audio input and output, but you should still check how the sound matches the visuals. If dialogue, music, or ambient sound has strict production requirements, further editing can be done after generation.
How do I submit a task and obtain the video result?
Submit a request to POST /wan/videos with model set to wan2.6-t2v, use prompt to describe the video content, and use text2video for text-to-video generation. After setting async to true, obtain task_id, then query the final status through /wan/tasks; successful results may provide information such as the video URL, dimensions, and thumbnail.