Dreamina is an **AI digital human (talking photo)** service based on ByteDance **OmniHuman 1.5**: upload a portrait photo and an audio clip to generate a video where the person in the photo **speaks with synchronized lip movements**. No modeling, no green screen, no real person on camera required—a front-facing photo + a voice clip can produce a natural digital human talking-head video, suitable for digital human presentations, memorial videos, virtual hosts, livestream product promotions, educational explanations, brand endorsements, and other scenarios. ## Core Capabilities - **Audio-driven lip sync**: The person naturally speaks according to the audio content, with lip movements, expressions, head, and shoulder movements all synchronized with the speech, saying goodbye to fake "face-pasted PowerPoint" movements. - **Single image is enough**: Generate with one clear front-facing portrait photo + one driving audio clip, with no need for multi-angle materials or pre-training. - **Controllable emotions and style**: Optionally use `prompt` to control expressions, emotions, stability, and style, making the presentation better match the tone of the content. - **Specify a subject in multi-person images**: Supports `mask_url` to specify and drive a particular person in a group photo. - **HD output + CDN hosting**: Outputs approximately 1664×1248, and generated results are automatically transferred to our CDN, ready to use and download directly. - **Both synchronous / asynchronous modes**: Get results directly for short tasks synchronously; for long tasks, use the `callback_url` callback or `async: true` + polling, stable and reliable. ## Use Cases - **Digital human presentations / virtual hosts**: Use one profile photo to produce talking-head videos in batches, replacing real-person recording. - **Memorial / remembrance videos**: Let people in old photos "speak" for warm scenarios such as family memorials. - **Product promotion presentations / marketing short videos**: Quickly generate digital human on-camera content for product explanations and brand endorsements. - **Education and training**: Pair script audio with an instructor image to generate course explanation videos in batches. - **Multilingual broadcasts**: Combine with audio in different languages to produce multilingual presentations using the same image. ## Input Specifications | Input | Requirements | | -------------- | ---------------------------------------------------- | | Image `image_url` | Publicly accessible; a clear, well-lit **front-facing portrait** works best, with the face unobstructed and appropriately sized | | Audio `audio_url` | Publicly accessible mp3/wav; recommended duration ≤60 seconds (for 1080p, recommended ≤30 seconds; for 720p, ≤60 seconds) | ## API Overview | API | Description | Billing | | ----------------------- | --------------------------------------------------- | ------- | | `POST /dreamina/videos` | Generate a digital human talking-head video (returns synchronously; becomes asynchronous when `callback_url` or `async:true` is provided) | Charged by video duration | | `POST /dreamina/tasks` | Query task results by `task_id` / `trace_id` | Free | ## Quick Start ```bash curl -X POST https://api.acedata.cloud/dreamina/videos \ -H "Authorization: Bearer $ACEDATACLOUD_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "omnihuman-1.5", "image_url": "https://cdn.acedata.cloud/4hfydw.jpg", "audio_url": "https://cdn.acedata.cloud/assets/examples/fish/5ade0339-5f11-487e-aacc-06a908271706-8e3fcb0e5547.mp3", "prompt": "Gentle, calm, natural, with slight head movements" }' ``` The `data.video_url` in the response is the generated video URL. If using asynchronous mode (`async: true`), first obtain the `task_id`, then poll: ```bash curl -X POST https://api.acedata.cloud/dreamina/tasks \ -H "Authorization: Bearer $ACEDATACLOUD_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "action": "retrieve", "id": "" }' ``` ## Billing Billing is based on the **duration of the generated video**, with the largest package costing approximately **¥1/second** (for example, a 10-second video costs approximately ¥10). The task query API `/dreamina/tasks` is free. A single API Token can be used to call all platform services, and free credits are provided upon the first application. ## Frequently Asked Questions - **What are the image requirements?** Clear front-facing portraits, even lighting, and unobstructed faces work best; side profiles, blur, and overly dark images will affect lip-sync and expression quality. - **How long can the audio be?** It is recommended to keep it within 60 seconds; overly long audio may be truncated or affect stability. - **Can I specify a person in a multi-person photo?** Yes, specify the driving subject by passing a subject mask through `mask_url`. - **How long does generation take?** It depends on the audio duration; for long tasks, it is recommended to use `callback_url` or `async:true` + `/dreamina/tasks` polling to avoid request timeouts. - **Will the results be saved?** Generated videos are automatically transferred to the CDN and returned through `video_url`; please download and save them to your own storage in time.