Fine-Grained Image Creation Model for Complex Compositions and Multilingual Design
Nano Banana Pro is an image generation and editing model built by Google DeepMind on Gemini 3 Pro, suited for tasks with high requirements for text readability, brand consistency, and compositional detail. It can transform text descriptions, product photos, and character references into posters, infographics, or scene composites, and adjust lighting, focus, and visual elements through natural language.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API Features
Clarify capacity, inputs and outputs, and invocation methods before choosing a model.
Creation Methods
Text-to-image generation; editing and compositing with text and reference images
Resolution
1K, 2K, 4K; select via resolution through the dedicated endpoint
Dedicated Endpoint Aspect Ratios
1:1, 3:2, 2:3, 16:9, 9:16, 4:3, 3:4
Native Reference Capability
Fusion of up to 14 reference images; consistency preservation for up to 5 characters
Creation Controls
Local edits, camera angles, focus, color grading, and scene lighting
Dedicated Endpoint Generation Count
count is 1–4, with a default of 1
Result Delivery
The dedicated endpoint returns image_url; the compatible endpoint can return either url or b64_json
The number of reference images and characters falls within native capability limits; aspect ratio, generation count, and result format are configured separately according to the selected endpoint.
Core Capabilities
Discover what nano-banana-pro can bring to your work.
Make Text Part of the Image
The model emphasizes the readability and visual presentation of text in images, making it suitable for integrating headlines, product selling points, and descriptions into posters or packaging mockups. It can also convert multilingual content, adjusting the text together with the original design. Providing accurate copy, language, and typographic hierarchy offers more control over the result than simply asking it to “make an ad.”
Organize Multiple Assets into a Complete Composition
It can combine references for people, products, and environments to generate scenes with consistent lighting and visual relationships. Its strength goes beyond simply stitching assets together; it also includes maintaining continuity in facial appearance and brand elements within complex compositions. Prompts should explain the purpose of each image and list the appearance, clothing, or product details that must be retained.
Refine Photographic Effects with Natural Language
Changing daytime into a night scene, shifting focus to the product, altering the camera angle, or adjusting the color mood can all be specified as clear editing goals. When making local edits, describing what “needs to change” separately from what should “remain unchanged” helps revise an existing visual direction rather than redesigning the entire image each time.
Use Cases
Start with specific tasks to find where the model can make an impact.
Brand Campaigns and Multilingual Posters
Input brand assets, campaign copy, and design direction to create key visuals with headlines and selling points, then adapt the text for different languages. Suitable for promotional tasks that need to balance image style with information delivery. Before delivery, focus on checking spelling, brand names, and text hierarchy, then choose a landscape or portrait format based on the publishing placement.
Product Scenes and Outfit Concepts
Provide product photos or references for people and clothing, describe the target environment, camera, and lighting, and generate product scene changes, styling displays, or lifestyle assets. For product-related tasks, clearly specify that shape, color, and logos must be retained; outfit results are suitable as visual concepts, while actual fit, materials, and wearing effects still need to be verified against physical items.
Turn Reference Materials into Infographics
Convert organized key knowledge points, processes, or handwritten notes into infographics, diagrams, and storyboard drafts. Specify the audience, information order, and visual focus in the input so the model can organize the image around the content. The deliverables are suitable for instructional explanations or proposal communication; numerical values and technical relationships should be reviewed item by item.
How to Choose This Model
Choose based on task complexity, input materials, and expected results.
Consider Pro First for Complex Visual Tasks
When a task includes substantial text, multiple reference assets, character consistency, or detailed photography control, Nano Banana Pro is better aligned with professional asset creation. The original Nano Banana corresponds to Gemini 2.5 Flash Image and is more suited to fast, easy editing; simple creative exploration does not need to pursue complex compositions from the start.
Choose Between It and Nano Banana 2 Based on Workload
Nano Banana 2 corresponds to Gemini 3.1 Flash Image and is positioned as a balanced model that combines speed with general-purpose creation; Pro focuses on complex visual tasks, brand consistency, and fine control. The former can be considered for frequent routine content, while key visuals with text-heavy layouts, complex asset relationships, or a need for careful adjustments are better suited to Pro.
Get Started
From a small-scale task to full integration.
01
Prepare Tasks and Materials
Define the goal, required inputs, and output requirements, using real business examples as a starting point.
02
Try It in the API Testing Area
Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to review the results.
03
Integrate According to the API Documentation
Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.
Usage Limitations
Before formal use, understand the scope of output quality and capabilities.
Stronger text generation does not mean character-perfect accuracy. Longer paragraphs, small fonts, multilingual mixed text, and brand marks may still have omissions or distortions. Important copy should be explicitly provided in the original text and checked item by item after image generation; text in images should not be treated directly as proofread final copy.
Character and product consistency is a generative capability, not pixel-level locking. When changing poses, camera angles, or lighting, faces, clothing textures, and small marks may change accordingly. Areas that must be strictly preserved should be explicitly listed in the instructions, and key commercial assets should still be compared with the original references.
Native reference image specifications do not mean every entry point accepts the same quantity; select assets directly relevant to the task when editing. Knowledge reasoning also does not mean every generation automatically accesses the internet; time-sensitive information such as weather and prices should be provided in the input, and structural diagrams cannot replace engineering validation.
Frequently Asked Questions
Answers to common questions about using nano-banana-pro.
What is the relationship between Nano Banana Pro and the original Nano Banana?
They are different models: Pro corresponds to Gemini 3 Pro Image, while the original corresponds to Gemini 2.5 Flash Image. Pro focuses on complex compositions, text within images, and fine-grained control; when calling it, enter nano-banana-pro. The service name nano-banana does not mean the original model is being used.
How do I generate 2K or 4K images?
When using /nano-banana/images, specify model as nano-banana-pro, action as generate, and select 2K or 4K through resolution. Set the image ratio separately with aspect_ratio; resolution tiers and landscape or portrait composition are two different choices.
How do I edit using an existing image?
For the dedicated endpoint, use action=edit, provide image URLs through image_urls, and describe the changes in prompt. You can also use /openai/images/edits, explicitly specifying the model and image. It is recommended to also state what to preserve, for example, preserve the product shape and change only the background and lighting.
Can I keep modifying the same design?
Yes, you can continue editing based on an existing result, for example, first finalize the composition, then adjust the text or lighting. In practice, you can use the previous image as the reference for the next round and clearly specify what to modify this time. Saving satisfactory intermediate results helps prevent later changes from affecting details that have already been confirmed.
How can generated results be used by an application?
The dedicated endpoint provides image_url in data and returns task_id and trace_id for associating tasks. The compatible image endpoint can return either an image URL or a Base64 result; when asynchronous processing is needed, you can use async and callback_url to integrate completed results into asset management or content workflows.
Model information · Updated: 2026-10-01. For calling parameters and billing rules, see the API and pricing sections.