How should images be submitted when GPT-4o analyzes images?
When using Chat Completions, include a text question and an image_url image block in the content array of the user message, and put the image address in image_url.url. detail can be auto, low, or high. It is recommended to also specify whether you want to identify, compare, or explain something, rather than submitting only an image without a task objective.
Can GPT-4o directly return generated images?
The gpt-4o on this page is used to understand text and images and generate text responses; it is not an image-generation entry point. If you need text-to-image generation, style changes based on reference images, or combined assets, choose gpt-4o-image or a dedicated image model. Understanding images and creating images are different delivery workflows, so first clarify whether your final output should be text or an image.
How should I choose between GPT-4o and gpt-4o-mini?
For combined text-and-image understanding, content explanation, and multilingual expression, prioritize evaluating GPT-4o; for simple and repetitive text tasks, try gpt-4o-mini first. It is recommended to compare key information retention, format stability, and number of revisions using the same set of real samples, then decide whether GPT-4o is needed rather than judging performance solely by the model name.
How can I make GPT-4o continue the previous discussion?
When using Chat Completions, include the necessary user and assistant history in messages; when using Responses, provide the relevant conversation in input. If you want to simplify history management, use AI Chat v2's stateful conversations and include the returned id in subsequent requests to keep the discussion focused on the same task.
Does GPT-4o include web access and code execution?
GPT-4o's language and programming capabilities do not mean that the model itself has web access or a code execution environment. When using custom function tools with Chat Completions, the application needs to validate parameters, perform operations, and return the results. AI Chat v2 provides built-in tools such as web search, web page fetching, and file reading, and can also use authorized connections; these are tool workflows provided by the entry point, not capabilities of the model itself. Running code still requires configuring the appropriate execution tools and implementing permission and security checks.