Generate cohesive multi-shot short films with sound from text scripts
wan2.6-t2v is the text-to-video model in the Alibaba Wan 2.6 series, designed for production from text ideas to narrative short films. It organizes visuals through intelligent shot scheduling, emphasizing consistency of subjects, scenes, and atmosphere across different shots. It natively supports up to 15 seconds and includes audio input and output capabilities. It is suitable for product stories, social content, and short-film concept previews, especially during creative stages when image or video assets have not yet been prepared.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Creation method
Text-to-video; natively supports text and audio input
Output modalities
Video, audio
Maximum native duration
15 seconds
Native resolution options
720P, 1080P
Shot organization
Intelligent multi-shot storytelling; requests can optionally use single or multi
Audio toggle
audio defaults to false; must be enabled manually for sound-enabled creation
Invocation and delivery
POST /wan/videos; supports asynchronous tasks, with results including video URL and dimension fields
Duration and resolution are described according to this model's native specifications; the audio toggle, shot options, and task delivery are explained according to this platform's invocation method.
Core capabilities
Turn short scripts into shot-based narratives
wan2.6-t2v is suitable not only for describing a dynamic moment, but also for organizing continuous plots into multi-shot short films. When creating, clearly write the sequence of events by opening, main action, and ending, then add shot scale and atmosphere. Intelligent shot scheduling provides organizational capability for short stories, making it suitable for expressing a clear visual theme within a limited duration.
Maintain continuity around the same subject
The focus of multi-shot creation is to connect characters, scenes, and atmosphere, rather than stitching together unrelated visuals. This model emphasizes consistency across these elements. Prompts can fix the subject's appearance, clothing, location, and lighting while varying only actions and viewing angles, making it easier to evaluate whether the story is coherent.
Plan visuals and sound together
Native capabilities include audio input and audio output, making it suitable to treat sound as part of short-film design rather than considering only silent visuals. When invoked on this platform, audio is disabled by default and must be actively enabled for sound-enabled tasks. When writing scripts, describe the intended sound at the same time, then check whether the audiovisual expression matches the creative intent after completion.
Applicable Scenarios
Product Story Concept Video
Enter a textual description of the product use scenario, subject actions, and ending shot to generate a short video that showcases the creative idea. For example, design a departure, use, and arrival storyline around travel products for internal discussion of communication direction. If existing product images must be reproduced exactly, use an image-starting creation model instead.
Social Content Script Test Shoot
Write a short-video idea as a brief shooting outline, describing the opening hook, main actions, and ending state, then generate a dynamic draft for review. It is suitable for comparing different story openings or shot arrangements. Change only one creative focus at a time to make it easier to determine which expressions are worth continuing to produce and edit.
Narrative Shot Previsualization
Enter character settings, scene atmosphere, and the sequence of actions to turn a text storyboard into dynamic previsualization material. Directors, designers, or content teams can use it to discuss visual rhythm and shot transitions before deciding on subsequent shooting plans. The deliverable is a short-video previsualization and should not be treated as a final cinematography plan locked frame by frame.
How to Choose This Model
Start from text, choose t2v
When you only have a script, scene concept, or creative direction, wan2.6-t2v is the direct choice. If you already have an image and want to use it as the opening frame, choose wan2.6-i2v; if you want to extract a character's appearance from an existing video and bring it into a new scene, consider wan2.6-r2v. Prioritize choosing the model based on the creative starting point rather than sending all materials to text-to-video processing.
Determine the short-video structure first, then choose the shot mode
For one subject action or a complete presentation process, use single first to organize a single-shot expression; for a short story with clear plot turns, try multi. wan2.6-t2v has a native duration limit of 15 seconds and is suitable for focusing on one theme. For longer narratives, split them into segments and arrange post-production transitions rather than stacking too many people, locations, and events into a short video.
Get Started
Prepare the input for this task first
In prompt, clearly describe the subject, environment, action sequence, shots, and sound goals; this model starts from text, so do not use a first-frame image as a substitute for description.
Choose the actual model and output
Specify model=wan2.6-t2v and action=text2video for /wan/videos, starting with 5 seconds and 720P; Wan 2.6 does not use Wan 3's 30 seconds, automatic duration, or all-type media parameters.
Retrieve the video and check audio and visuals
Use async=true to save task_id, query /wan/tasks or receive a callback; after completion, retrieve video_url, check the subject, actions, and sound, then proceed to editing.
Trial suggestion: an audio short narrative from a text storyboard
Input and goal
Early morning in a café, a barista grinds beans, pours water, and hands over a cup of coffee; the scene presents three distinct actions in sequence, keeping the person and in-store lighting consistent, with sounds of the coffee machine and soft clinking of cups and saucers.
Acceptance criteria and next steps
Use text rather than reference images; first check the action sequence and shot continuity, then preview the audio; complex processes are best produced as storyboards, rather than assigning all multiple reference assets in a series to T2V.
Usage boundaries
Multi-shot consistency is a key capability of this model, but that does not mean character appearance, prop positions, or action transitions can be locked frame by frame. When requirements for character details or product appearance are strict, review each completed shot individually; tasks with existing clear visual assets are better suited to the corresponding image or reference-video model.
The native maximum length is 15 seconds, so complex plots need to be compressed or split up. Prompts that simultaneously require frequent scene changes, interactions among multiple characters, and continuous actions can easily diffuse the main focus. It is recommended to keep a clear main storyline and handle subtitle layout, brand logos, and precise editing in post-production.
For audio creation, note that audio is disabled by default. Native audio capability does not mean you can specify any audio format, precisely control the timing of every line of dialogue, or directly complete professional mixing. Before publishing, check the audio content, volume, and coordination with the visuals; add voice-over or edit the sound further if necessary.
Frequently Asked Questions
Can wan2.6-t2v generate videos directly from images?
This model is designed for text-to-video generation. If creation requires an existing image as the starting frame, choose wan2.6-i2v; if you need to continue a character's appearance from a video, choose wan2.6-r2v. Selecting the appropriate model based on the source material helps clarify how the visuals are controlled.
How long and at what resolution can it generate videos?
wan2.6-t2v has a native maximum duration of 15 seconds, with 720P and 1080P resolution options. It is suitable for creating short videos focused on a single subject. For longer stories, it is recommended to split them into multiple clips, then connect them through editing while maintaining the same subject and scene settings for each segment.
How should prompts for multiple shots be written?
Write the prompt as a brief shooting outline: first describe the subject and scene, then describe the opening, main action, and ending in order. Keep the character's appearance, environment, and atmosphere consistent, and avoid redefining the subject for each shot. When multiple shots are needed, you can select shot_type as multi and review the generated continuity.
Do generated videos include sound by default?
audio defaults to false in requests on this platform, so audio creation should explicitly set audio to true. This model natively supports audio input and output, but you should still check how the sound matches the visuals. If dialogue, music, or ambient sound has strict production requirements, further editing can be done after generation.
How do I submit a task and obtain the video result?
Submit a request to POST /wan/videos with model set to wan2.6-t2v, use prompt to describe the video content, and use text2video for text-to-video generation. After setting async to true, obtain task_id, then query the final status through /wan/tasks; successful results may provide information such as the video URL, dimensions, and thumbnail.