All models

kling-v2-6

KuaishouVideo
Get your API key
kling-v2-6

A short-video model with first and last frame control and synchronized audio

kling-v2-6 is Kuaishou Kling's V2.6 video model, suitable for turning text concepts or static images into short videos. Its key value is that pro mode provides both first and last frame control and synchronized audio, making it useful for product showcases, atmospheric clips, and narrative shots; in talking-photo workflows, it can also combine portrait photos with existing audio to create lip-synced videos.

KuaishouModel brand
VideoModel type
Text · Image guidanceCreation method
STANDARD APIs · QUICK SETUP

Bring this model into your workflow

Submit requests to the public API at api.acedata.cloud using the documented parameters, then use the results in your application.

API hostapi.acedata.cloud
modelkling-v2-6

Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.

Specifications and API features

Creation method
Text-to-video, image-to-video; also includes a lip-sync workflow using a photo and audio
Video duration
The platform supports 5 or 10 seconds
Generation modes
std, pro
Aspect ratios
16:9, 9:16, 1:1
First and last frame control
Pro image-to-video supports a first frame plus a last frame; the last frame cannot be used alone
Synchronized audio
Pro supports generate_audio; disabled by default
Result delivery
Video URL, video ID, task ID, and status; supports asynchronous processing and callbacks

The above are this platform's V2.6 invocation specifications. Photo lip-sync is a separate workflow and is not the same as synchronized audio during video generation.

Core capabilities

Generate sound and visuals together

V2.6's pro mode supports enabling audio while generating video, making it suitable for short-form creations that need both visuals and sound. Prompts can be organized around the subject, actions, and scene, then audio can be enabled with generate_audio; when only visual assets are needed, keep it disabled for easier later dubbing and editing.

Structure shots with a beginning and an end

Image-to-video can start from a specified first frame, and pro mode can also add a last frame to set clear opening and closing visuals for a clip. It is suitable for projects with existing product images, character images, or storyboards, allowing creation to develop around the selected composition; first and last frames constrain visual boundaries rather than providing frame-by-frame editing of every intermediate frame.

Turn portrait photos into talking clips

The talking-photo workflow accepts a portrait image and existing audio, first animates the photo, then completes audio lip-sync, delivering the final video in a single submission. You can use a prompt to describe movements and expressions during the animation stage, and explicitly select kling-v2-6; this workflow is suited to tasks with existing recordings rather than directly synthesizing speech from text.

Use Cases

Product Showcases and Social Shorts

Provide a main product image and a description of the display action to create short videos suitable for landscape, portrait, or square publishing. If you have already designed the closing shot, use pro to add an end frame so the product showcase lands on the intended composition; the delivered video link can enter the editing workflow, where you can further add brand text, subtitles, and a complete marketing arrangement.

Soundscapes and Story Shots

Starting from descriptions of scenes, character actions, and events, use text-to-video to create short shots, and enable synchronized audio in pro mode. This is suitable for first validating a story segment or soundscape concept before deciding whether to move into formal editing; organizing the task into clear individual clips makes checking easier than cramming an entire long-form script into a single generation.

Talking-Head Videos with Existing Recordings

Prepare a clear front-facing photo of one person, along with a recording that matches the target video length, and submit them through the talking photo entry point. The output includes the final lip-synced video, and you can also obtain the video link from the photo animation stage, making it easy to separately check character movement and lip-sync results. It is suitable for character introductions, brief explanations, and talking-head samples.

How to Choose This Model

Choosing Between V2.5 Turbo and V2.6

If you need first and last frames, the pro modes of both can be considered; if you want audio at the same time as generation, V2.6 pro better meets the task requirements, while V2.5 Turbo does not provide this feature. For clips that need neither sound nor an end frame, start with V2.6 std; choose pro when sound or a clearly defined closing shot is needed, rather than judging solely by the version name.

When to Switch to V3 or Omni

V2.6 is suitable for fixed short durations, text- or first-frame-driven generation, and audio creation in pro mode. If the task requires integer durations of 3–15 seconds, native 4K mode, or dedicated camera movement controls, consider the corresponding V3 capabilities; if you need to combine multiple reference images, reference videos, or directly edit existing videos, choose O1 or V3 Omni rather than applying these capabilities to V2.6.

Getting Started

Choose a text or image starting point

text2video starts from a prompt; image2video provides start_image_url. If you need to specify the ending composition, you can also provide end_image_url under pro.

Choose a mode based on whether the task has audio

Specify model=kling-v2-6 for /kling/videos, and choose 5 or 10 seconds. When synchronized audio is needed, use mode=pro and generate_audio=true; for silent tasks, audio can remain disabled.

Review the visuals and selected audio

Save the async task ID or use callback_url to retrieve results; check the subject, action, and ending. For tasks with audio enabled, also listen to the dialogue, sound effects, and timing; when not enabled, proceed to editing as silent footage.

Trial suggestion: a product shot with synchronized audio

Input and goal

Keep the appearance of the coffee machine in the reference image, generate a shot of coffee being poured into a cup, with the sound of the machine operating and coffee flowing, and keep the image stable.

Review and next steps

When synchronized audio is needed, choose pro and generate_audio=true, and listen to the mechanical sound; std does not inherit pro's audio or end-frame capabilities.

Usage boundaries

  • std does not support synchronized audio or end frames; choose pro when using these two features. Image-to-video requires a first-frame image, and an end frame can only be used together with a first frame; before submitting, clearly describe the desired action sequence to avoid mistaking the constraints of the starting and ending images for precise control over the entire motion path.
  • V2.6 does not support mode=4k or camera_control-specific camera movement parameters. Camera intent can be expressed in the prompt, but this differs from parameterized camera movement control; multi-image references, reference videos, and video editing are workflows for other models, not general input capabilities of this model.
  • Photo lip-sync requires accessible image and audio links, and a clear front-facing image of a single person is recommended. Supported audio formats are mp3, wav, m4a, and aac, with files no larger than 5MB and a recommended duration no longer than the target video; this workflow relies on existing audio and does not replace voice creation or long-form spoken-content production.

Frequently Asked Questions

How should I choose between std and pro in V2.6?

If you only create clips without synchronized audio and do not need to specify an end frame, you can choose std. If you need synchronized audio or start/end frame control, use pro. Both modes support 5-second or 10-second videos; these pro features must be enabled explicitly and are not all automatically turned on simply by selecting this mode.

How do I enable synchronized audio in V2.6?

In the /kling/videos request, explicitly set model to kling-v2-6, mode to pro, and generate_audio to true. This switch is off by default. Synchronized audio and submitting existing audio for photo lip-sync are two different approaches; choose based on whether you already have a recording.

Can I generate a video using only an end-frame image?

You cannot use an end frame alone as input for this image-to-video workflow. Use action=image2video, provide start_image_url, and add end_image_url as needed in pro mode. If you only have one image, use it as the start frame and describe the desired action in the prompt.

How do I make a photo speak according to a recording with V2.6?

Submit image_url and audio_url to /kling/talking-photo, explicitly specify model=kling-v2-6, and choose 5 seconds or 10 seconds. The prompt can describe the movements and expressions during the photo animation stage. The final video_url is the lip-synced video, while source_video_url is the intermediate photo animation video.

How do I receive generation results without keeping the connection waiting?

You can set async=true, save the returned task_id first, then retrieve the status and result through task queries; you can also configure callback_url to receive completion notifications. After obtaining video_url, download it or proceed to the editing workflow. Both generation endpoints should explicitly specify kling-v2-6 to avoid using other default models.