Turn text requirements and reference images into usable visual assets
gpt-image-1 is OpenAI's native multimodal image model, designed for text-to-image generation and image-based visual editing. Its strengths include not only generating images in different styles, but also understanding composition requirements, following creative guidelines, and rendering text in images. It is suitable for posters, product visuals, illustrations, and design drafts, using dedicated image generation or editing endpoints for integration.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Creation method
Text-to-image generation; editing reference images with text instructions
Editing reference input
This platform's editing endpoint: a single image URL or an array of up to 16 URLs
Image file formats
PNG, JPEG, WebP
Background control
transparent、opaque、auto
Reference fidelity control
The editing endpoint provides input_fidelity: high、low
Result delivery
The image endpoint supports URL、b64_json; asynchronous responses can return task_id
Batch requests
Image endpoint n parameter range: 1–10
Model capabilities include image generation, text rendering, and visual editing; the input quantities and result formats above are invocation specifications for this platform's image endpoint.
Core capabilities
Bring creative requirements to life
gpt-image-1 can create images by combining requirements for theme, style, and composition, making it suitable for organizing scattered creative briefs into visual concepts. Prompts can separately describe the subject, environment, lighting, and negative space, then add brand colors and elements that must be avoided, making the generation goal clearer than simply specifying an art style.
Text is also part of visual design
Text rendering in images is a clear strength of this model, making it useful for headline posters, promotional graphics, and visual drafts with labels. When creating, list the text that needs to appear separately and specify the title hierarchy and placement; before final delivery, still check spelling, punctuation, and small text to avoid treating generated results directly as finalized layout.
Continue designing from reference images
In addition to generating from scratch, gpt-image-1 is also suitable for making visual modifications based on original images, such as changing the background, adjusting the color palette, or developing a sketch into a complete graphic. When editing, clearly specify what should be retained and what needs to change, and choose reference fidelity control appropriately; this conveys design intent more effectively than simply writing “optimize it.”
Applicable Scenarios
Event Posters and Marketing Drafts
Enter the event theme, headline copy, brand color palette, and layout requirements to generate posters or social media graphics for discussion. This is suitable for comparing different visual directions first, then selecting a concept for further revision. When dates, prices, and event terms are involved, they should be checked item by item, and the necessary design review should be completed before official publication.
Product Scenes and Asset Redesigns
Provide a product image URL, describe the desired background, lighting, and visual atmosphere, and emphasize the appearance features that should be retained to explore product promotional images. After generation, focus on checking whether logos, materials, structure, and colors are accurate; when transparent assets are needed, choose an output format that supports transparency for further processing.
Sketch-to-Illustration and Visual Proposals
Submit sketches or reference images together with written instructions, describing the line style, color relationships, and intended use to create illustrations, graphic elements, or concept proposals. Each reference image should serve a clear purpose, for example, one explaining the composition and another explaining the color palette, reducing conflicting visual requirements and making subsequent selection and revision easier.
How to Choose This Model
Choose It When You Need to Continue Editing After Generation
If a task includes both “generate an image first” and “continue modifying it based on the original image,” the generation and editing workflow of gpt-image-1 is worth prioritizing. Compared with DALL·E 3, which is primarily positioned for text-to-image generation, this is better suited to organizing modification requests around reference images. For text analysis or general Q&A alone, a conversational model should be selected.
Compare Relevant Models by Specific Task
gpt-image-1, gpt-image-1.5, and gpt-image-2 are different models, so you should not replace calls based on their names alone or assume their parameters are exactly the same. Existing gpt-image-1 workflows can continue to be optimized around text rendering and reference image editing; when considering a switch, it is recommended to compare results using the same set of posters, product images, and modification instructions before deciding whether to migrate.
Getting Started
Choose text-to-image or image editing
Provide a prompt when generating; provide both an image and editing instructions when editing. Clearly specify the original text, subject preservation requirements, and target aspect ratio.
Specify the model and parameter format
Use the image generation or image editing endpoint, explicitly specifying model=gpt-image-1; use auto or WIDTHxHEIGHT for size, and generate one image first before evaluating. Set masks, quality, and file format according to the guide for this endpoint.
Check images and cost records
Read the URL or Base64 image according to the response format, asynchronously save the task_id before querying results; check text, reference details, and the alpha channel, and record usage according to the current Pricing rules.
Trial suggestion: event poster and background whitespace
Input and goal
Create a horizontal community workshop poster, titled “Build Something,” featuring tools and materials on a desk, with space reserved on the right for registration information, in a blue-and-white color scheme with clear text.
Acceptance and next steps
First establish the visual design through the image generation API, then refine it with the image editing API; proofread the title and whitespace, and integrate according to the image responses for generation and editing.
Usage boundaries
Text rendering capability does not mean complex layouts can be completed accurately every time. Dense text, small labels, and strictly aligned layouts should be checked separately after image generation; for contract content, prices, or product specifications that must be accurate, it is more appropriate to use the generated image as a visual draft and then complete the text layout.
Reference image editing does not mean pixel-level locking. Even when preserving the subject is requested, details, edges, lighting, shadows, or logos should still be reviewed; for tasks with strict requirements for product appearance and brand assets, clearly specify what must be preserved and check the result, rather than interpreting high-fidelity control as meaning every detail remains absolutely unchanged.
Safety filtering still applies to image creation, and moderation=low does not mean arbitrary content is allowed. The output format must also match the intended use: transparent backgrounds should use a format that supports transparency, and JPEG is not suitable for preserving an alpha channel; batch generation should not be treated as a guarantee of cross-image consistency.
Frequently Asked Questions
Are gpt-image-1 and regular gpt-4o the same API call?
No. gpt-image-1 is designed for image generation and editing; when calling it, you should select the corresponding image endpoint and model ID. It has technical ties to the ChatGPT image experience, but this does not mean that regular gpt-4o conversation calls are the same workflow, nor does it automatically include features such as voice or web access.
How can I modify an existing image?
Call /openai/images/edits, explicitly specify model=gpt-image-1, and provide the image URL and prompt modification instructions. You can submit multiple reference images, but it is best to explain the purpose of each image while distinguishing what should be kept and what should be changed, making the editing goal easier to express.
How can I make titles more accurate when generating posters?
Write the title text separately from the image description, clearly specify the exact text to appear, its position, and hierarchy, and try to avoid crowding long passages into complex backgrounds. gpt-image-1 excels at rendering text in images, but the completed image still needs to be proofread character by character; for formal posters, refine the typography after confirming the visual design.
Can I generate assets with transparent backgrounds?
The image endpoint provides background=transparent, which can be used to request a transparent background; PNG or WebP is recommended for output. The prompt should also describe requirements for the subject and edges. After downloading the result, check the alpha channel and edge quality, especially thin lines, semi-transparent materials, and shadow areas.
How do I integrate image generation and editing into an application?
Use the image generation endpoint for creation from scratch; use the editing endpoint to modify existing images, and explicitly specify model=gpt-image-1. The application reads the image URL or Base64 result; for asynchronous requests, save task_id and query it later. Do not process image responses as ordinary conversation text.