Create short videos with sound using text and first/last frames
Seedance 1.5 Pro is ByteDance's video generation model, and doubao-seedance-1-5-pro-251215 is its specific invocation version. It supports text-to-video, image-to-video, and sound generation, making it suitable for turning shot descriptions or still images into short films. When you need to start from a specified image, control the final composition, or add sound to a scene, this version provides a clear creative workflow.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Creation modes
Text-to-video, image-to-video, first/last frame control
Video duration
Platform invocation: 4–12 seconds; duration=-1 for automatic duration
Output resolution
Platform options: 480p, 720p, 1080p
Aspect ratios
16:9, 4:3, 1:1, 3:4, 9:16, 21:9, adaptive
Sound generation
Enable with generate_audio=true; disabled by default
Generation controls
Random seed, fixed camera, watermark options, last-frame return
The values above are the invocation ranges for this version on this platform. Creation modes and controls are used according to this model and do not represent the capabilities of other Seedance versions.
Core capabilities
Plan visuals and sound together
1.5 Pro can enable sound while generating video, so you do not need to split every creation into separate visual and audio tasks. Prompts can describe character actions, environments, and the sounds you want to hear at the same time, such as wind blowing through hair and the sound of wind. When silent footage is needed, turn off sound generation and arrange music or editing afterward.
Use first and last images to define a shot's start and finish
Image-to-video can begin from a first-frame image, or include a last-frame image to specify the ending shot, then use text to describe the action and camera movement in between. It is suitable for shot designs with an established composition; unlike providing only character references, first and last frames emphasize the beginning and end of the image rather than identity references across scenes.
Adjust generation for the delivery format
Landscape, portrait, square, and ultra-wide formats can serve presentations, mobile short videos, and conceptual shots respectively. During invocation, you can explicitly set the aspect ratio and duration, and use a fixed camera or random seed depending on the task. When you need to continue planning the next segment, you can request the last frame to be returned as an image for preparing the following shot.
Applicable Scenarios
Turn Product Images into Showcase Shots
Use a product image as the first frame, add subject motion, background changes, and camera direction, and generate short shots for advertising edits. For example, make steam rise from a cup and have the camera slowly move closer, rather than merely asking for it to “look premium.” The deliverable is downloadable video footage, making it easy to add brand text and editing rhythm later.
Creative Test Clips for Scenes with Sound
Enter a clear short scene, specifying the subject, action, and sound—for example, rain falling on a street while a person walks by holding an umbrella—and enable generate_audio. This is suitable for first validating creative directions that combine sound and visuals before deciding on subsequent production; after delivery, the audio track should be reviewed, not just individual frames.
Dynamic Previsualization Between Storyboards
When you already have opening and ending storyboard frames, set the two images as the first and last frames respectively, describe the movement, turn, or camera changes in between, and generate a transition preview. It can help directors, animation teams, or design teams discuss whether the shot works; the delivered video can also be used for presentations rather than replacing full-length production.
How to Choose This Model
Choose 1.5 Pro for Short Videos with Sound
If the task is a text- or first/last-frame-driven short video and you want to generate sound directly, 1.5 Pro is the clear choice. Compared with the 1.0 series, which does not support generate_audio, it adds a path for sound-enabled creation. If you only need silent visuals, there is no need to enable audio just for the version number; consider the relevant 1.0 models separately for text-to-video and image-to-video tasks.
Choose Another Version for Complex References and Editing
If you need character reference images, reference audio, or reference video, consider the 2.0 series rather than adding these assets as roles in 1.5 Pro. If you need to edit existing video, extend clips, or generate up to 30 seconds, choose 2.5; if you need 4k output, consider 2.0 Standard. Your choice should follow the task, not just the version number.
Getting Started
Organize content and Assets
Provide text in content; images can use first_frame/last_frame to represent the starting and ending visuals; the last frame should be consistent with the creation method and asset composition.
Choose the Correct Version and Shot Settings
Specify model=doubao-seedance-1-5-pro-251215 for /seedance/videos; first test with duration=5, resolution=720p, and a clearly defined aspect ratio. When sound is needed, explicitly set generate_audio=true; it is disabled by default.
Save the Task and Final Frame
For asynchronous processing, first obtain the task_id, then query /seedance/tasks or receive a callback; after completion, check the subject, action, ending, and audio, then save the final video and returned last frame as needed.
Trial recommendation: First-frame-driven sound scene
Input and goal
Use an image by a window on a rainy day as the first frame. The camera slowly moves closer to the glass, raindrops slide down, the indoor lighting is soft, and the sound is rain outside the window, with no voices or music.
Acceptance and next steps
Request sound with generate_audio=true, and check the rain sound and raindrop motion segment by segment; do not use the 2.x multimodal reference_image method.
Usage boundaries
1.5 Pro is suitable for 4–12-second clips and can also use automatic duration, but it should not be planned as a 30-second long-video task. Longer narratives can be split into multiple shots and then edited together; returning the last frame helps prepare the next shot, but does not mean an existing video is automatically extended.
For image input, use first_frame or last_frame. Do not treat reference_image, reference_audio, or reference_video as asset modes for this model. Enabling sound generation also does not mean uploading an audio reference; the two correspond to different creation methods.
Each text content item may contain up to 1000 characters; focus on describing the subject, action, camera, and sound. Do not use the frames control supported only by the 1.0 series, and do not set the 2.5-specific edit, extend, or retrieval tools for this model; prioritize placing parameters in top-level fields.
Frequently Asked Questions
How does 1.5 Pro generate videos with sound?
Set generate_audio=true in the request, and describe the scene and the sounds you want to appear in the text. This option is disabled by default, so merely writing a sound description without enabling the parameter should not be considered a sound-enabled delivery solution. After generation is complete, check both the visuals and the actual audio track.
How should first-frame and last-frame images be submitted?
Add an image_url item to the content array, place the address in the image_url.url object, and set the role to first_frame or last_frame respectively. First- and last-frame tasks should also include a textual description of the changes in between; image_url cannot be written directly as a string.
Can it generate 4k or 30-second videos?
This version is used with 480p, 720p, 1080p, and durations of 4–12 seconds, and also supports automatic duration. For 4k tasks, consider 2.0 Standard; for tasks up to 30 seconds, consider 2.5. Do not directly use the resolution or duration settings of other versions with 1.5 Pro.
Can I use a person's photo as a character reference?
You can use a person's photo as the first frame, so the shot begins from that image; however, this differs from a reference_image character reference. If the goal is to preserve the person's identity while changing the scene, consider the 2.0 series that supports this reference method. The 1.5 Pro image workflow focuses on first and last frames.
How do I obtain the video file after making a call?
Submit model and content to /seedance/videos, with the model set to doubao-seedance-1-5-pro-251215. Asynchronous mode returns task_id; then query the task, or use callback_url to receive notifications. After the task is complete, download the video through data.video_url.