All models

gemini-3.7-flash

GoogleChatReasoningVision
Get your API key
gemini-3.7-flash

Multimodal reasoning model for long documents and code collaboration

Gemini 3.7 Flash is Google's stable multimodal reasoning model in the Gemini 3 series, suitable for analyzing code, documents, and images within the same task. It combines long input capacity, adjustable thinking levels, function calling, and structured output for coding assistance, information organization, and multi-step assistants. This platform provides two usage methods: image-and-text chat and managed sessions.

GoogleModel brand
ChatModel type
Reasoning, visual understandingTask capabilities

Specifications and interface features

Clarify capacity, input and output, and calling methods before selecting a model.

Native input limit
1,048,576 tokens
Native output limit
65,536 tokens
Native input and output
Text, image, video, audio, and PDF input; text output
Thinking levels
low, medium, high; minimal is not supported
Image-and-text chat
Mixed text and image_url messages, with streaming responses supported
Structured output and tools
JSON, JSON Schema output, and function calling
Managed sessions
Continue chats with session IDs, file-link input, SSE and NDJSON events

Native capacity and modality describe model capabilities; this platform's image-and-text chat and file-tool workflows are provided by their respective selected entry points.

Core capabilities

Learn what gemini-3.7-flash can bring to your work.

Reason over related materials together

Its large input capacity is suitable for including requirements, related code, API contracts, and historical discussions at the same time, comparing and synthesizing them around the same issue. When organizing materials in practice, grouping them by file or section and clearly stating the problem to solve is more useful than simply piling in all content; outputs can be organized into a change list, summary, or items requiring confirmation.

Combine image understanding with text tasks

You can submit charts, page screenshots, and textual descriptions together, allowing the model to explain the image, compare information, or organize observations around a specified question. Multimodal messages use text blocks and image_url blocks, and images can use public links or Base64 data URIs; the deliverable is text analysis, not regenerated images.

Extend from answers to tool collaboration

Function calling is used to describe the operations and parameters the model wants to execute, while structured output makes it easier to connect results to business applications. When orchestrating yourself, you should preserve complete tool calls and results; when using managed sessions, you can combine file reading and authorized tools to complete multi-step tasks, while execution permissions are still determined by the application and tool configuration.

Applicable scenarios

Start with specific tasks to find where the model can be effective.

Code modification and review assistance

Provide relevant source code, error logs, requirements, and test conditions, allowing the model to first identify the scope of impact and then propose modification plans, code snippets, and testing recommendations. This is suitable for routine maintenance tasks requiring cross-file understanding; when delivering results, require explanations of the basis for changes and unverified assumptions, while actual compilation and testing are still completed in the development environment.

Long-document organization and Q&A

Organize report text, clauses, or meeting notes by topic, and request summaries, comparison tables, and lists of questions. When you need to process PDF, CSV, or TXT links, you can choose the managed-session file_url workflow, allowing file tools to read content before analysis, and continue follow-up questions in the same session.

Chart and interface analysis

Submit chart screenshots or product pages and specify the analysis goal, such as explaining trends, checking whether text and graphics are consistent, or organizing interface improvement items. Output can be an item-by-item explanation or predefined JSON fields; when precise values are involved, providing the original table also helps with cross-checking.

How to choose this model

Choose based on task complexity, input materials, and expected results.

Existing 3.7 workflow: decide whether to migrate by task

Gemini 3.7 Flash is a stable model, but not the latest Flash; Gemini 3.8 Flash is now available. If you already have validated 3.7 prompts and tool workflows, you can keep the current configuration first, then compare the new version on the same tasks. Whether switching from 3.6 is worthwhile should be based on answer usability, tool parameter correctness, and the number of rework cycles, rather than assuming every task will be more efficient.

Make trade-offs based on complexity and integration method

When you need to connect multiple materials, inspect images, and produce structured results, 3.7 Flash is worth considering; simple classification may not require long context and a higher reasoning level, while complex reasoning can be compared with the Pro series. If you manage message history and tool loops yourself, choose Chat Completions; if you want to continue conversations by session ID and use file tools, choose AI Chat v2.

Get started

From a small-scale task to production integration.

01

Prepare tasks and materials

Define the goal, required inputs, and output requirements, using real business examples as a starting point.

02

Try it in the API testing area

Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to review the results.

03

Integrate according to the API documentation

Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.

Usage boundaries

Before production use, understand output quality and capability scope.

  • It is a comprehension and text generation model and does not provide image generation, audio generation, or the Live API. Native support for analyzing video and audio does not mean media can be placed arbitrarily in image fields; text-and-image entry points should use text and images, while file materials can use the file-reading method of managed sessions.
  • The reasoning level should be low, medium, or high; minimal cannot be treated as a valid native option. The response budget must also leave room for reasoning, as max_tokens that is too low may cause empty content or truncated answers; long reports should be generated by chapter, and the finish reason should be checked.
  • A long input limit does not mean the model will pay equal attention to all materials, nor does it mean it can directly run projects or operate a computer. Prioritize relevant excerpts, specify acceptance criteria, and preserve tool state in multi-turn tasks; code execution and external writes require the appropriate environment and authorization.

Frequently Asked Questions

Answers to common questions about using gemini-3.7-flash.

Which invocation ID should I use for Gemini 3.7 Flash?

Use gemini-3.7-flash, keeping the dot between the numbers. For text-and-image conversations, call /gemini/chat/completions; for hosted sessions, call /aichat2/conversations. The former submits model and messages, while the latter can start a conversation with model and question.

How much input and output does it support?

The native input limit is 1,048,576 tokens, and the output limit is 65,536 tokens; they are not the same budget. When organizing long tasks, you should still filter relevant materials and set response length according to the deliverable; these native figures do not mean every call should use the full capacity.

How do I set its thinking intensity?

You can choose low, medium, or high; minimal is not natively supported. Simple summaries can start with low, while multi-condition analysis can try medium or high. Do not treat an extremely small response budget as a way to disable thinking; the Gemini 3.x Flash guide recommends setting max_tokens above 512.

Can it read PDFs and generate images or audio?

It natively supports PDF understanding, but does not generate images or audio. Hosted sessions on this platform can submit PDF links through file_url and analyze them with the file-reading tool; standard text-and-image messages use text and image_url content blocks, and these two input methods should be handled separately.

How do I continue a conversation after calling a tool?

When orchestrating function calls yourself, retain the call ID, function parameters, and execution result, then return the tool result to the model. Hosted sessions can continue using the returned id; if a task pauses while awaiting additional information, fill it back through tool_results and keep tool_use_id consistent with the tool ID in the pause event.

Model information · Updated: 2026-10-01. For call parameters and billing rules, see the API and pricing sections.

Use gemini-3.7-flash for your next task

Start with a clear goal and use real results to determine whether it fits your work.