Multimodal reasoning model for long-horizon coding and multi-step tasks
Gemini 3.8 Flash is Google's Flash model built for software engineering, agents, and complex enterprise workflows. It excels at reasoning over code, requirements, images, and analytical materials within a single task, making it suitable for work that requires continuous decomposition, tool use, and deliverable organization. Compared with 3.7 Flash, it focuses more on deeply solving complex tasks and also requires sufficient budget for thinking and responding.
Clarify capacity, input/output, and invocation methods before selecting a model.
Native input limit
1,048,576 tokens
Native output limit
65,536 tokens
Native inputs and outputs
Text, image, video, audio, and PDF input; text output
Native thinking levels
low, medium, high; minimal is not supported
Text and image invocation
Chat Completions supports mixed text and image_url content
Structured output and tools
JSON, JSON Schema output; function tool calling
Responses and sessions
Streaming responses; AI Chat supports managed multi-turn sessions
Native specifications describe model capabilities; this platform's image-text, file reading, and tool workflows are used separately according to the selected interface.
Core Capabilities
Learn what gemini-3.8-flash can bring to your work.
Continuously reason around engineering goals
Its focus is not just generating a piece of code, but continuously analyzing requirements, existing implementations, and feedback. It is suitable to provide related modules, error logs, and acceptance criteria together, and ask the model to propose modifications, explain the scope of impact, and add testing recommendations, making the output closer to a complete software engineering task.
Analyze images and text together
With mixed image-and-text messages, the model can generate analysis results based on screenshots, charts, and textual requirements. For example, compare interface screenshots against requirements to check implementation differences, or explain relationships in a chart. Deliverables remain text, code, or structured data; it does not directly generate images or audio.
Connect controlled multi-step tool workflows
Function calling is suitable for incorporating retrieval, queries, and business operations into task steps, while structured output makes it easier for systems to continue processing results. When orchestrating on your own, retain records of tool calls and returns; when using AI Chat v2, you can combine authorized tools for multi-turn tasks and monitor execution progress through events.
Use Cases
Start with specific tasks to find where the model can be effective.
Cross-module fixes and refactoring
Provide the relevant source code, error logs, interface contracts, and behaviors that must not change, then let the model map dependencies and propose modifications. Request patch suggestions, change descriptions, and a regression testing checklist, then continue asking questions using real test results. This is suitable for maintenance tasks that require understanding across files.
Material comparison and report organization
Provide business materials, analysis questions, and the report structure together, and ask the model to distinguish facts, inferences, and information still needed, producing summaries, difference lists, or analysis reports. When PDF files and other documents need to be read, use the file-link workflow in AI Chat v2 instead of submitting files as images.
Stateful task assistants
Suitable for developing engineering or business assistants that continuously follow up on requirements: first describe the goal, then add materials and feedback to gradually form the final deliverable. The AI Chat API can use a session ID to continue context; when tool collaboration and progress display are needed, v2's structured streaming events are more convenient.
How to choose this model
Choose based on task complexity, input materials, and expected results.
How to choose between it and 3.7 Flash
When a task involves a longer coding process, repeated verification, or multi-step reasoning in specialized domains, prioritize 3.8 Flash. If it mainly involves simple processing, short answers, and greater emphasis on computational efficiency, 3.7 Flash remains a valuable option. When comparing, look at whether the final deliverable passes acceptance and the total token consumption, rather than only the length of a single response.
Do not confuse it with Cyber or generative models
3.8 Flash is a model for general engineering and enterprise tasks; 3.8 Flash Cyber is a security-specific variant for authorized defensive personnel. Its vulnerability discovery or patching metrics must not be applied to this model. When image generation, speech synthesis, or real-time audio-video interaction is needed, choose the corresponding generative or Live model rather than this model.
Get started
From a small-scale task to production integration.
01
Prepare tasks and materials
Clarify the objective, required inputs, and output requirements, and use real business examples as a starting point.
02
Try it in the API testing area
Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to view the results.
03
Integrate according to the API documentation
Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.
Usage boundaries
Before formal use, understand output quality and capability scope.
Higher reasoning levels may increase token consumption. Complex tasks in particular need enough space for both reasoning and the final answer. An output budget that is too small may result in empty content; when calling, it is recommended to set max_tokens above 512 and continue adjusting based on task complexity, rather than treating this value as a sufficient budget.
Long input capacity does not mean material selection can be ignored. Code and documents should be organized around the current issue, with versions, constraints, and acceptance criteria clearly specified; irrelevant history increases processing burden. Long tasks should check intermediate conclusions in stages, and generated patches should also be tested in a real environment.
Support for function calling does not mean the model will automatically execute local code or click computer interfaces. Tools in Chat Completions must be executed by the application and have their results returned; tool operations in AI Chat v2 depend on connection and authorization. Tasks involving writing, sending, or publishing should have clear permission and confirmation boundaries.
Frequently Asked Questions
Answers to common questions about using gemini-3.8-flash.
Which ID should I use to call Gemini 3.8 Flash?
Use gemini-3.8-flash. When maintaining message history and the tool loop yourself, you can use /gemini/chat/completions; if you want to continue a conversation with a session ID, you can use /aichat2/conversations or /aichat/conversations. v2 is more suitable for creating new tool-based assistants.
How do I submit an image for analysis?
In the Chat Completions message content, combine text and image_url content blocks, and place the image address in image_url.url. You can use a publicly accessible image URL or a Base64 data URI, and specify in the text the area, question, and expected output to analyze.
How do I choose the reasoning effort?
This model natively supports low, medium, and high, but does not support minimal. Lower reasoning effort can be used for simple tasks, while higher effort can be used for complex analysis; when calling through the platform, set the corresponding parameter only if the selected interface explicitly supports reasoning effort configuration for this model. For Chat Completions, it is recommended to set max_tokens above 512 and increase the budget according to task complexity, leaving room for both reasoning and the final answer; above 512 does not guarantee that the budget is sufficient.
Can it read PDFs and generate speech?
The model can natively understand PDFs, but its output is text and it does not support speech generation. When processing PDFs, you can use the file_url file-link content block in AI Chat v2 to complete summarization and Q&A through the file-reading workflow; Chat Completions image_url should not be used as a general file upload field.
How can I make the results easier for programs to process?
You can use response_format to select a JSON object or JSON Schema, and clearly specify field meanings, required fields, and how missing values should be handled. The client should still validate the returned structure and business rules; when business functions need to be called, use tools to define functions, then process the call parameters and fill in the execution results.
Model information · Updated: 2026-10-01. For calling parameters and billing rules, see the API and pricing sections.