Create dynamic short videos with consistent characters from multiple reference images
happyhorse-1.1-r2v is the reference-image-to-video model of HappyHorse 1.1, suited for tasks with existing character, clothing, or prop assets where you want to continue creating short videos in new scenes. It combines ordered reference images with textual shot descriptions, using visual assets to constrain subjects and prompts to arrange actions, environments, and camera work, with a focus on short-video creation that requires high character consistency.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Clarify capacity, inputs and outputs, and invocation methods before selecting a model.
Creation mode
Reference-image-to-video; input prompt and 1–9 image_urls
Asset references
Use names such as character1 and character2 in image order
Native output
480P, 720P, 1080P; 24 fps; MP4
Invocation resolution
720P, 1080P; default 1080P
Video duration
3–15 seconds; default 5 seconds
Available aspect ratios
16:9, 9:16, 1:1, 4:3, 3:4
Task delivery
Supports asynchronous queries and completion callbacks, returning video_url
Native specifications include audio capabilities and 480P output; this platform entry provides reference-image generation calls at 720P and 1080P.
Core capabilities
Learn what happyhorse-1.1-r2v can bring to your work.
Maintain character identity with reference images
Unlike using text alone to describe appearance, reference images provide clear visual guidance for characters, clothing, and props. You can change environments and actions in prompts while continuing to use the same character assets, making it suitable for creating short shots in different scenes around one character; consistency is a creative goal, not a guarantee of being completely unchanged frame by frame.
Assign asset roles in order
Images are not unordered attachments: use character1 and character2 to point to the corresponding assets, and specify which image provides the subject and which provides the clothing or visual style. For example, have character1 walk through grass while using the leather and gold-trim style of character2, establishing a clear relationship between assets and shot descriptions.
From generation task to video delivery
Submit tasks through POST /happyhorse/videos, explicitly select reference_to_video and this model, then set duration, resolution, and aspect ratio. When persistent connections are inconvenient, use async or callback_url, save the task_id to query results or receive completion notifications, and obtain the video link for download and subsequent editing.
Applicable Scenarios
Start with specific tasks to find where the model can be effective.
Character Series Shorts
Provide reference images of the same character, describe the scene, action, and filming approach for each shot, and produce editable short clips. For example, create shots of a character walking on grass or pausing on a street, first verify appearance consistency, then select results that meet narrative requirements to form a series of content.
Creative Previsualization for Clothing and Props
Prepare character, clothing, or prop assets, clearly specify the purpose of each reference image in the prompt, and generate dynamic previews of characters carrying props or presenting looks. The deliverables are suitable for discussing visual direction, action arrangements, and scene atmosphere; when product details are involved, check whether textures, logos, and structures match the original design.
Content Assets for Multiple Aspect Ratios
Using the same set of character reference assets, create landscape, portrait, or square short clips for creative test videos in different display placements. When providing input, adjust both the aspect ratio and composition description rather than only changing the ratio; after output, check subject placement and action space, then pass it to the editing workflow to add subtitles and brand information.
How to Choose This Model
Choose based on task complexity, input materials, and expected results.
Reference Images or a Fixed First Frame
If you want a character to continue appearing in a new environment and do not require the reference image itself to become the opening frame, choose happyhorse-1.1-r2v. If the task focuses on animating a finalized image as the first frame, choose happyhorse-1.1-i2v; if there are no image assets and you mainly rely on script and shot text, use happyhorse-1.1-t2v.
Distinguish Versions from Editing Tasks
Both 1.1-r2v and 1.0-r2v generate from reference images. The stated difference in public specifications is that 1.1 adds a native 480P option, while this entry still provides 720P and 1080P. Do not interpret version numbers as a fixed quality increase. If you already have a video that needs wardrobe changes, style transfer, or local replacement, choose happyhorse-1.0-video-edit.
Get Started
From a small-scale task to formal integration.
01
Prepare the Task and Materials
Define the goal, required inputs, and output requirements, using real business examples as a starting point.
02
Try It in the API Debugging Area
Open the trial page, confirm the parameters supported by this entry, then submit a small-scale task to review the results.
03
Integrate According to the API Documentation
Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.
Usage Limits
Before official use, understand the output quality and capability scope.
Reference-image generation is not video editing, nor is it motion replication. This model's creative inputs are images and prompts; do not treat video_url in the shared API as its video reference capability. When you need to preserve existing camera movement or modify the original footage, use the corresponding editing workflow.
A single generation lasts 3–15 seconds, making it unsuitable for directly delivering long, continuous narratives. You can first split the content into individual shots, then edit them into a finished video; character details, spatial relationships, and action continuity across segments still require manual review, and reference images cannot replace a complete continuity check.
Images must be submitted via publicly accessible URLs, with their order consistent with prompt references. If multiple assets impose conflicting requirements on the same character's clothing, appearance, or style, the creative intent becomes unclear; it is recommended to select assets carefully and clearly define the role of each image.
Frequently Asked Questions
Answers to common questions about using happyhorse-1.1-r2v.
Will the reference image directly become the first frame of the video?
This model uses reference images to guide characters and visual elements, and should not be used as a fixed first-frame workflow. If you need an image with a specific composition as the opening frame, choose happyhorse-1.1-i2v. r2v is better suited for preserving the reference subject while rearranging the environment, actions, and camera through text.
How do I make prompts correspond to multiple reference images?
Put 1–9 image URLs in image_urls, and reference them in order in the prompt using character1, character2, and so on. In addition to identifying the corresponding assets, clearly describe the subject's actions, scene, and the purpose of each asset, rather than merely listing names without explaining their relationships.
Which key items need to be explicitly specified when making a call?
Set action to reference_to_video, model to happyhorse-1.1-r2v, and submit prompt and image_urls. Do not rely on the default text-to-video action. Then choose resolution, ratio, and duration according to your needs to form a complete reference-image generation request.
Can it generate audio or preserve the original video's sound?
This model has native audio capabilities, but reference-image generation does not mean preserving the original video's sound, nor does it mean specifying voice-over or lip synchronization. The origin usage for preserving original audio belongs to video editing tasks; if audio has specific delivery requirements, review the generated result and arrange any necessary audio post-production.
How do I retrieve asynchronously generated videos?
After submitting with async, save task_id and query the task through /happyhorse/tasks; you can also provide callback_url to receive completion notifications. The result includes the task status and video_url; statuses may be pending, succeeded, or error. Download the video for review only after it has completed successfully.
Model information · Updated: 2026-10-01. For call parameters and billing rules, see the API and pricing sections.