Lightweight Video Generation Model for Multimodal Short-Form Creation
Seedance 2.0 Mini is a lightweight video model in ByteDance's Seedance 2.0 series, suitable for short-form previews, character shots, and advertising concept validation. It can generate video from text or images, and can also combine image, audio, and video references to organize visuals, motion, and pacing. On this platform, you can choose 480p or 720p, landscape or portrait aspect ratios, and audio output to create short shots for clearly defined asset purposes.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API Features
Creation methods
Text-to-video, image-to-video, first and last frame control, multimodal reference
Up to 9 images, 3 audio files, and 3 videos; the total duration of audio and video must not exceed 15 seconds each
Audio control
generate_audio can enable videos with audio; disabled by default
Invocation and delivery
POST /seedance/videos; supports asynchronous tasks, callback notifications, and video link delivery
The above describes the creation and invocation scope for this model on this platform; native model specifications and available service parameters should be understood separately.
Core Capabilities
Organize shots by assigning roles to assets
Multimodal references do more than attach an image to a video: reference_image is used for character or subject reference, reference_video for motion and camera movement reference, and reference_audio for sound and rhythm reference. The prompt then specifies the role each asset plays, making it suitable for organizing existing visual assets into new short shots.
Distinguish composition locking from subject reference
When motion needs to begin from a specified image, use first_frame; when the ending of a shot needs to be defined, add last_frame. If you want a person or subject to enter a new scene, use reference_image and describe the new environment and action. The two methods serve different creative goals, so there is no need to mistake subject reference for locking the original image composition.
Include sound in short-form planning
After enabling generate_audio, you can generate video with sound; the prompt can describe both visual actions and the desired sounds. Reference audio, meanwhile, participates in creation as another type of input asset. The two have different meanings: providing an audio reference does not mean enabling audio output, so sound goals and reference usage should be planned separately during creation.
Use Cases
Ad Shot Concept Preview
Input product images and shot descriptions to create different short video concepts around display angles, backgrounds, lighting, and camera movement. You can first choose 480p to review composition and motion, then use 720p to produce the selected shot. The deliverable is video material for team review, making it easier to discuss creative direction rather than generating a complete ad project in one go.
Character Short Video Creation
Use authorized clear images of people or characters as reference_image, combined with prompts for scenes, clothing, and actions, to create short shots of the same subject in different environments. Choose 9:16 for vertical content, and clearly specify character behavior and camera movement; after delivery, focus on checking whether the appearance, actions, and scene match the specifications.
Storyboard and Dynamic Concept Validation
Use storyboard images as the first frame or first and last frames, describe the motion between the images, and generate short shots to demonstrate composition changes. If the focus is on referencing a particular action or camera movement, use reference video mode instead. Generated results can be used for director communication, animation previsualization, or concept presentation, helping determine the next production direction.
How to Choose This Model
How to Choose Between Mini, Fast, and Standard
Mini is suitable for lightweight multimodal short-shot creation, especially for tasks where 480p and 720p already meet preview and material presentation needs. 2.0 Fast is the faster version in the same series, and both support these two resolutions; if the project explicitly requires higher resolution, consider 2.0 Standard, which supports up to 4k. Do not infer specific time or quality differences solely from version names.
When to Switch to Seedance 2.5
For tasks focused on 4–15-second generation, first-and-last-frame control, and image, audio, or video references, choose Mini. If you need up to 30 seconds, creation based solely on audio references, more reference materials, or editing and extending existing videos, Seedance 2.5 is more suitable. These are differences in creative workflow, not capabilities that can be obtained by adding fields to Mini.
Get Started
Organize content and Materials
In content, set specific roles for images, videos, and audio, and explain the purpose of each asset in the text; organize first-and-last-frame constraints and subject references separately according to the task.
Choose the Correct Version and Shot Settings
Specify model=doubao-seedance-2-0-mini-260615 for /seedance/videos; first test with duration=5, resolution=720p, and a clear aspect ratio. When sound is needed, explicitly set generate_audio=true; it is disabled by default.
Save Tasks and Final Frames
For asynchronous requests, first obtain the task_id, then query /seedance/tasks or receive a callback; after completion, check the subject, action, ending, and audio, then save the finished video and returned final frame as needed.
Trial suggestion: character short-shot test
Input and goal
The cartoon character in the character reference image peeks out from beside the door, waves, then retreats indoors, maintaining the character design, with a relaxed scene and a fixed camera.
Acceptance and next steps
Use 480p or 720p to verify the motion and character reference; when no video footage is available, first create a clear single shot, and do not infer model scale from the lightweight version name.
Usage limits
Mini supports 480p and 720p and should not be selected for 1080p or 4k final-video requirements; generated duration is 4–15 seconds, so it is also not suitable for directly handling continuous 30-second clips. Automatic duration lets the model determine the clip length; it does not mean the duration range of this model is removed.
First frame, first-and-last frames, and full-modal reference are mutually exclusive modes; first_frame or last_frame cannot be used together with reference_image, reference_video, or reference_audio. Choose first-and-last frames when the start and end images must be specified precisely; choose reference mode when combining assets is needed.
Reference audio must be wav or mp3, 2–15 seconds each and no more than 15MB each; reference video must be mp4 or mov, 2–15 seconds each. The total duration of each type of asset must not exceed 15 seconds. Authorization should be obtained for real-person and character assets, and clear, unobstructed images are more suitable as character references.
Frequently Asked Questions
How do I accurately call this Mini version?
Submit a request to POST /seedance/videos, set model to doubao-seedance-2-0-mini-260615, and provide the content array. For text items, use type=text, with a maximum of 1000 characters per item; it is recommended to write the resolution, aspect ratio, and duration as top-level fields for clear control over generation settings.
Can Mini make a single image move?
Yes. Put the image in the image_url object, specify it as the starting frame with first_frame, then describe the motion and camera movement in text. The image address must be written in the url field within image_url; do not assign a string directly to image_url. If the subject changes scenes, reference_image is more suitable.
Can I upload only audio to generate a video?
Mini supports reference_audio for multimodal reference, but when using audio reference, you must also provide an image or video reference and explain the purpose of each asset in the prompt. Audio reference does not mean audio output is enabled; to generate a video with sound, separately set generate_audio=true. If you need to create a video based solely on audio reference, you can choose Seedance 2.5.
Can reference videos be used for direct editing or extension?
Mini's reference_video is used to reference motion, camera movement, and other content during generation; it is not equivalent to dedicated video editing or extension tasks. To replace content in an existing video or extend an original clip, use the corresponding modes in Seedance 2.5; Mini is better suited for creating new short videos based on reference assets.
How do I obtain the generated video after submission?
You can set async=true to first obtain a task_id, then query the task result; you can also provide callback_url to receive a notification when the task is complete. The video link in the completed result is used to download the finished video; do not treat the returned task ID as a video file. If you need a last-frame image, enable return_last_frame.