All models

nano-banana-pro ★

GoogleImage
Get your API key
nano-banana-pro

A refined image creation model for complex compositions and multilingual design

Nano Banana Pro is an image generation and editing model built by Google DeepMind on Gemini 3 Pro, suited for tasks with high requirements for text readability, brand consistency, and compositional detail. It can transform text descriptions, product photos, and character references into posters, infographics, or scene composites, and adjust lighting, focus, and image elements through natural language.

GoogleModel brand
ImageModel type
Generation · EditingCreation method
STANDARD APIs · QUICK SETUP

Bring this model into your workflow

Submit requests to the public API at api.acedata.cloud using the documented parameters, then use the results in your application.

API host
api.acedata.cloud
model
nano-banana-pro
Get your API key

Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.

Specifications and API features

Creation method
Text-to-image generation; editing and compositing with text and reference images
Resolution
1K, 2K, 4K; select via resolution through the dedicated endpoint
Dedicated endpoint aspect ratios
1:1, 3:2, 2:3, 16:9, 9:16, 4:3, 3:4
Native reference capabilities
Blend up to 14 reference images; maintain consistency for up to 5 characters
Creation controls
Local edits, camera angles, focus, color grading, and scene lighting
Dedicated endpoint generation count
count is 1–4, defaulting to 1
Result delivery
The dedicated endpoint returns image_url; the compatible endpoint can return url or b64_json

The number of reference images and characters is within native capability limits; aspect ratio, generation count, and result format are set separately according to the selected endpoint.

Core Capabilities

Make text part of the image

The model emphasizes the readability and visual presentation of text within images, making it suitable for integrating headlines, product selling points, and descriptions into posters or packaging mockups. It can also convert multilingual content, allowing text to be adjusted together with the original design. Providing accurate copy, language, and typographic hierarchy is more helpful for controlling results than simply asking to “make an advertisement.”

Organize multiple assets into a complete composition

It can combine references for people, products, and environments to generate scenes with unified lighting and visual relationships. Its strength is not just stitching assets together, but also maintaining consistency in character appearance and brand elements within complex compositions. Prompts should explain the purpose of each image and list the appearance, clothing, or product details that must be retained.

Refine photographic effects with natural language

Changing daytime into night, shifting focus to a product, altering the shooting angle, or adjusting the color mood can all serve as clear editing goals. When making local changes, describing separately what “needs to change” and what “must remain unchanged” helps revise an existing visual direction rather than redesigning the entire image each time.

Use Cases

Brand campaigns and multilingual posters

Input brand assets, campaign copy, and design direction to create key visuals with headlines and selling points, then adapt the text for different languages. Suitable for promotional tasks that need to balance image style and information delivery. Before delivery, focus on checking spelling, brand names, and text hierarchy, then choose a landscape or portrait format according to the placement.

Product scenes and outfit concepts

Provide product photos or references for people and clothing, describe the target environment, camera, and lighting, and generate product scene changes, styling displays, or lifestyle assets. For product tasks, clearly specify that shape, color, and logos must be preserved; outfit results are suitable as visual concepts, while actual fit, materials, and wearing effects still need to be verified against the physical items.

Turn reference materials into infographics

Transform organized knowledge points, processes, or handwritten notes into infographics, diagrams, and storyboard drafts. Specify the audience, information order, and visual focus in the input so the model can organize the image around the content. Deliverables are suitable for instructional explanations or proposal communication; values and technical relationships should be reviewed item by item.

How to Choose This Model

Prioritize Pro for Complex Visual Tasks

When a task involves substantial text, multiple reference assets, character consistency, or detailed photography controls, Nano Banana Pro is better suited to professional asset production. The original Nano Banana corresponds to Gemini 2.5 Flash Image and is geared more toward fast, easy editing; simple creative exploration does not need to pursue complex compositions from the start.

Choose Between It and Nano Banana 2 Based on Workload

Nano Banana 2 corresponds to Gemini 3.1 Flash Image and is positioned as a balanced model that combines speed with general-purpose creation; Pro focuses on complex visual tasks, brand consistency, and fine control. Consider the former for high-frequency routine content, while key visuals with text-heavy layouts, complex asset relationships, or a need for careful adjustment are better suited to Pro.

Getting Started

First Determine Whether to Generate or Edit

Choose generate for text-based creation; choose edit to modify existing assets, provide reference images with image_urls, and specify separately what to retain and what to change.

Select the Full ID and Aspect Ratio

Specify model=nano-banana-pro, action, and prompt for /nano-banana/images; start with aspect_ratio=1:1, resolution=2K, and count=1, setting the aspect ratio and resolution separately.

Save the Result Before the Next Editing Round

Retrieve the image from data[].image_url; for asynchronous requests, query or receive callbacks using task_id. When continuing edits, resubmit the selected image and narrow the scope of changes in each round.

Trial Recommendation: Brand Visual Localization

Input and Goal

Preserve the subject, brand colors, and composition of the reference poster, replace the title with the provided English copy, rearrange the text width, and maintain the lighting and product appearance of the original image.

Review and Next Steps

Verify the localized copy, brand identity, and portrait details word by word; if the text placement is unsuitable, describe only the text area and layout in the next round.

Usage Limits

  • Stronger text generation capabilities do not mean character-by-character accuracy. Longer passages, small fonts, mixed languages, and brand marks may still have omissions or distortions. Important copy should be provided explicitly in its original form and checked item by item after image generation; text in images should not be treated directly as proofread final copy.
  • Consistency for people and products is a generative capability, not pixel-level locking. When changing pose, camera angle, or lighting, faces, clothing textures, and small marks may change accordingly. Areas that must be strictly preserved should be clearly listed in the instructions, and key commercial assets should still be compared with the original reference.
  • Native reference image specifications do not mean every endpoint accepts the same quantity; when editing, select assets directly relevant to the task. Knowledge reasoning also does not mean every generation automatically accesses the internet; time-sensitive information such as weather and prices should be provided in the input, and structural diagrams cannot replace engineering validation.

Frequently Asked Questions

What is the relationship between Nano Banana Pro and the original Nano Banana?

They are different models: Pro corresponds to Gemini 3 Pro Image, while the original corresponds to Gemini 2.5 Flash Image. Pro focuses on complex compositions, text in images, and fine-grained control; when calling it, use nano-banana-pro, and the service name nano-banana does not mean the original model is being used.

How do I generate 2K or 4K images?

When using /nano-banana/images, specify model as nano-banana-pro, action as generate, and select 2K or 4K through resolution. Set the image aspect ratio separately with aspect_ratio; resolution tiers and landscape or portrait composition are two different choices.

How do I edit using an existing image?

The dedicated endpoint uses action=edit, provides image addresses through image_urls, and describes the changes in prompt. You can also use /openai/images/edits, explicitly specifying the model and image. It is recommended to also state what to preserve, for example, preserve the product shape and change only the background and lighting.

Can I keep modifying the same design?

Yes, you can continue editing an existing result, for example, first finalize the composition, then adjust the text or lighting. In practice, you can use the previous image as the reference for the next round and clearly specify what to change this time only. Saving satisfactory intermediate results helps prevent subsequent changes from affecting details that have already been confirmed.

How can generated results be used by an application?

The dedicated endpoint provides image_url in data and returns task_id and trace_id for associating tasks. Compatible image endpoints can return either an image URL or a Base64 result; when asynchronous processing is needed, use async and callback_url to integrate completed results into asset management or content workflows.