A balanced reasoning model for coding iterations and multi-step tasks
Gemini 3.5 Flash is Google's stable multimodal reasoning model, focused on coding agents, subtask collaboration, and long-running workflows that require continuous progress. It delivers analysis, code, and structured results in text, making it suitable for balancing reasoning capability and response efficiency during repeated planning, tool calls, and feedback checks, while also leveraging image understanding for complex tasks.
Clarify capacity, input/output, and invocation methods before choosing a model.
Native input limit
1,048,576 tokens
Native output limit
65,536 tokens
Native input and output
Text, image, video, audio, and PDF input; text output
Reasoning and result organization
Native support for Thinking, function calling, and structured output
Text and image invocation
/gemini/chat/completions: messages accepts text and image_url content blocks
Conversation invocation
/aichat2/conversations: supports stateful conversations and file link content blocks
Generation limitations
Does not support image generation, audio generation, or Live API
Native capacity and modality describe model capabilities; this platform's actual input formats, tool usage, and response methods vary according to the selected invocation endpoint.
Core Capabilities
Learn what gemini-3.5-flash can bring to your work.
Keep coding loops moving forward
Gemini 3.5 Flash focuses not on writing code in a single pass, but on adapting to a cycle of analyzing requirements, proposing changes, receiving test feedback, and continuing to refine. You can input code snippets, error logs, and constraints together to have it generate patch suggestions and testing ideas; tests are executed by tools provided by the application, making it easy to retain records of each review round.
Joint analysis of long materials and images
Its million-scale native input capacity is suited to organizing lengthy code, documents, and task histories, while image understanding can supplement screenshots, charts, and interface information. Clearly specifying the objects to compare, evidence to focus on, and delivery format in the prompt can keep the output centered on specific issues instead of merely compressing large amounts of material into a generic summary.
Integrate reasoning into business workflows
Native function calling and structured output capabilities are suited to turning analysis results into task data that subsequent programs can process. The Chat Completions endpoint provides tool definitions and JSON formatting options; when managed history is needed, AI Chat v2 can be used to continue tasks with a session ID, keeping supplementary information and follow-up questions coherent.
Use Cases
Start with specific tasks to find where the model can be effective.
Development debugging and refactoring
Input related functions, failing cases, logs, and expected behavior, and have the model first list possible causes, then propose modification plans and a regression test checklist. After the application runs the tests, feed the results back into the next round to ultimately deliver reviewable change suggestions, risk explanations, and test items, making it suitable for development assistance that requires multiple iterations.
Document review and difference extraction
Submit report text, clauses, or multiple versions of materials together with review objectives, and request differences, supporting evidence, and items requiring confirmation. When using AI Chat v2, file links can be submitted through file_url, combining file reading with summaries and follow-up questions; deliverables can be designed as comparison tables, action lists, or structured records.
Screenshot-driven task breakdown
Combine product screenshots, charts, and text requirements into image-and-text messages, allowing the model to identify key content and break down to-dos, such as organizing interface issues, explaining chart changes, or generating acceptance steps. Output can be organized as text or JSON for easy entry into ticketing systems; screenshot analysis and actual interface operations are different tasks, and clicks will not be completed automatically.
How to choose this model
Choose based on task complexity, input materials, and expected results.
Choosing between the stable and preview versions
gemini-3.5-flash is the stable version call name, while gemini-3-flash-preview is the preview version name; clearly specify the target during integration. Existing applications that have established workflows around the output format, tool loops, and test sets of 3.5 Flash can continue evaluating with this model. When comparing newer Flash versions, validate them using the same tasks rather than judging migration benefits solely by version number.
Choose the model tier based on task difficulty
When multimodal analysis, coding iteration, and continuous planning are needed, 3.5 Flash is a balanced choice. For simple classification, fixed-field extraction, and retryable batch processing, compare Flash Lite; for difficult reasoning and complex analysis, compare Pro. Focus on task success rate, number of tool rounds, and total output volume, rather than treating the length of a single response as a proxy for reasoning quality.
Get started
From a small-scale task to production integration.
01
Prepare tasks and materials
Define the objective, required inputs, and output requirements, using real business examples as a starting point.
02
Try it in the API testing area
Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to review the results.
03
Integrate according to the API documentation
Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.
Usage boundaries
Before formal use, understand output quality and capability scope.
This is a text-output model. It can understand multimodal materials, but it does not generate images, speech, or real-time bidirectional audio and video. Audio analysis and audio generation should be considered separately; when finished images, voice-overs, or real-time conversations are needed, choose the corresponding generation or real-time model.
Long inputs do not mean all content will be used with equal accuracy. Code repositories and long documents should include objectives, key locations, and acceptance criteria; when generating longer answers, reserve a reasoning budget as well and avoid setting output limits too low, which may cause the final text to be empty or incomplete.
Native code execution and preview Computer use capabilities do not mean ordinary chat requests will automatically run programs or operate a computer. Tool execution for Chat Completions must be connected by the application; file reading and authorized tools in AI Chat v2 belong to the conversational workflow, and write actions still require the appropriate permissions.
Frequently Asked Questions
Answers to common questions about using gemini-3.5-flash.
How do I choose between the two Gemini 3.5 Flash endpoints?
Choose /gemini/chat/completions when you need to manage messages yourself, handle tool calls, and parse choices. Choose /aichat2/conversations when you want to save conversation history, continue asking follow-up questions by id, or use file content blocks; its standard JSON response includes answer and id.
How can I make Gemini 3.5 Flash analyze images?
In Chat Completions, write message content as an array of content blocks, combining text and image_url, and put an accessible image address or Base64 data URI in image_url.url. The text should specify the object of interest and output requirements, such as extracting fields, comparing differences, or explaining charts.
Can I have it read PDFs?
Gemini 3.5 Flash natively supports PDF understanding. When using this platform, you can provide a PDF link through file_url in an AI Chat v2 message and combine it with file reading for analysis; do not submit a PDF link as a regular image block, and do not equate native capabilities with the upload methods of all endpoints.
Why is there token usage but no final text?
Gemini 3.5 Flash uses reasoning tokens, and an output budget that is too small may be exhausted before a final answer is produced. The usage guide recommends setting max_tokens to 512 or higher; complex tasks should allow more room, and you should use finish_reason, usage, and the actual text to determine whether the budget needs to be increased.
Is it the same invocation name as gemini-3-flash-preview?
No. Google lists gemini-3.5-flash as Stable and gemini-3-flash-preview as Preview. To invoke this model, use the full ID gemini-3.5-flash; when replacing models, especially retest structured results, tool loops, and long-task performance to avoid deploying it directly after only changing the name.
Model information · Updated: 2026-10-01. For invocation parameters and billing rules, see the API and pricing sections.