All models

gemini-3.8-flash ★

GoogleChatReasoningVision
Get your API key
gemini-3.8-flash

Multimodal reasoning model for long-horizon coding and multi-step tasks

Gemini 3.8 Flash is Google's Flash model built for software engineering, agents, and complex enterprise workflows. It excels at reasoning over code, requirements, images, and analytical materials within a single task, making it suitable for work that requires continuous decomposition, tool use, and deliverable organization. Compared with 3.7 Flash, it focuses more on deeply solving complex tasks and also requires sufficient budget for thinking and responding.

GoogleModel brand
ChatModel type
Reasoning, visual understandingTask capabilities

Specifications and interface features

Clarify capacity, input/output, and invocation methods before selecting a model.

Native input limit
1,048,576 tokens
Native output limit
65,536 tokens
Native inputs and outputs
Text, image, video, audio, and PDF input; text output
Native thinking levels
low, medium, high; minimal is not supported
Text and image invocation
Chat Completions supports mixed text and image_url content
Structured output and tools
JSON, JSON Schema output; function tool calling
Responses and sessions
Streaming responses; AI Chat supports managed multi-turn sessions

Native specifications describe model capabilities; this platform's image-text, file reading, and tool workflows are used separately according to the selected interface.

Core Capabilities

Learn what gemini-3.8-flash can bring to your work.

Continuously reason around engineering goals

Its focus is not just generating a piece of code, but continuously analyzing requirements, existing implementations, and feedback. It is suitable to provide related modules, error logs, and acceptance criteria together, and ask the model to propose modifications, explain the scope of impact, and add testing recommendations, making the output closer to a complete software engineering task.

Analyze images and text together

With mixed image-and-text messages, the model can generate analysis results based on screenshots, charts, and textual requirements. For example, compare interface screenshots against requirements to check implementation differences, or explain relationships in a chart. Deliverables remain text, code, or structured data; it does not directly generate images or audio.

Connect controlled multi-step tool workflows

Function calling is suitable for incorporating retrieval, queries, and business operations into task steps, while structured output makes it easier for systems to continue processing results. When orchestrating on your own, retain records of tool calls and returns; when using AI Chat v2, you can combine authorized tools for multi-turn tasks and monitor execution progress through events.

Use Cases

Start with specific tasks to find where the model can be effective.

Cross-module fixes and refactoring

Provide the relevant source code, error logs, interface contracts, and behaviors that must not change, then let the model map dependencies and propose modifications. Request patch suggestions, change descriptions, and a regression testing checklist, then continue asking questions using real test results. This is suitable for maintenance tasks that require understanding across files.

Material comparison and report organization

Provide business materials, analysis questions, and the report structure together, and ask the model to distinguish facts, inferences, and information still needed, producing summaries, difference lists, or analysis reports. When PDF files and other documents need to be read, use the file-link workflow in AI Chat v2 instead of submitting files as images.

Stateful task assistants

Suitable for developing engineering or business assistants that continuously follow up on requirements: first describe the goal, then add materials and feedback to gradually form the final deliverable. The AI Chat API can use a session ID to continue context; when tool collaboration and progress display are needed, v2's structured streaming events are more convenient.

How to choose this model

Choose based on task complexity, input materials, and expected results.

How to choose between it and 3.7 Flash

When a task involves a longer coding process, repeated verification, or multi-step reasoning in specialized domains, prioritize 3.8 Flash. If it mainly involves simple processing, short answers, and greater emphasis on computational efficiency, 3.7 Flash remains a valuable option. When comparing, look at whether the final deliverable passes acceptance and the total token consumption, rather than only the length of a single response.

Do not confuse it with Cyber or generative models

3.8 Flash is a model for general engineering and enterprise tasks; 3.8 Flash Cyber is a security-specific variant for authorized defensive personnel. Its vulnerability discovery or patching metrics must not be applied to this model. When image generation, speech synthesis, or real-time audio-video interaction is needed, choose the corresponding generative or Live model rather than this model.

Get started

From a small-scale task to production integration.

01

Prepare tasks and materials

Clarify the objective, required inputs, and output requirements, and use real business examples as a starting point.

02

Try it in the API testing area

Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to view the results.

03

Integrate according to the API documentation

Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.

Usage boundaries

Before formal use, understand output quality and capability scope.

  • Higher reasoning levels may increase token consumption. Complex tasks in particular need enough space for both reasoning and the final answer. An output budget that is too small may result in empty content; when calling, it is recommended to set max_tokens above 512 and continue adjusting based on task complexity, rather than treating this value as a sufficient budget.
  • Long input capacity does not mean material selection can be ignored. Code and documents should be organized around the current issue, with versions, constraints, and acceptance criteria clearly specified; irrelevant history increases processing burden. Long tasks should check intermediate conclusions in stages, and generated patches should also be tested in a real environment.
  • Support for function calling does not mean the model will automatically execute local code or click computer interfaces. Tools in Chat Completions must be executed by the application and have their results returned; tool operations in AI Chat v2 depend on connection and authorization. Tasks involving writing, sending, or publishing should have clear permission and confirmation boundaries.

Frequently Asked Questions

Answers to common questions about using gemini-3.8-flash.

Which ID should I use to call Gemini 3.8 Flash?

Use gemini-3.8-flash. When maintaining message history and the tool loop yourself, you can use /gemini/chat/completions; if you want to continue a conversation with a session ID, you can use /aichat2/conversations or /aichat/conversations. v2 is more suitable for creating new tool-based assistants.

How do I submit an image for analysis?

In the Chat Completions message content, combine text and image_url content blocks, and place the image address in image_url.url. You can use a publicly accessible image URL or a Base64 data URI, and specify in the text the area, question, and expected output to analyze.

How do I choose the reasoning effort?

This model natively supports low, medium, and high, but does not support minimal. Lower reasoning effort can be used for simple tasks, while higher effort can be used for complex analysis; when calling through the platform, set the corresponding parameter only if the selected interface explicitly supports reasoning effort configuration for this model. For Chat Completions, it is recommended to set max_tokens above 512 and increase the budget according to task complexity, leaving room for both reasoning and the final answer; above 512 does not guarantee that the budget is sufficient.

Can it read PDFs and generate speech?

The model can natively understand PDFs, but its output is text and it does not support speech generation. When processing PDFs, you can use the file_url file-link content block in AI Chat v2 to complete summarization and Q&A through the file-reading workflow; Chat Completions image_url should not be used as a general file upload field.

How can I make the results easier for programs to process?

You can use response_format to select a JSON object or JSON Schema, and clearly specify field meanings, required fields, and how missing values should be handled. The client should still validate the returned structure and business rules; when business functions need to be called, use tools to define functions, then process the call parameters and fill in the execution results.

Model information · Updated: 2026-10-01. For calling parameters and billing rules, see the API and pricing sections.

Use gemini-3.8-flash for your next task

Start with a clear goal and assess whether it suits your work based on real results.