Multimodal reasoning model for long documents and complex code
Gemini 2.5 Pro is Google's long-context thinking model, focused on complex programming, mathematics, and scientific problems, and is also suitable for cross-document analysis and large codebase review. It can combine text and images to understand tasks and produce analysis, code, or structured text. On this platform, you can choose chat calls with self-organized messages or use a session workflow that preserves history.
Clarify capacity, input/output, and invocation methods before selecting a model.
Native input limit
1,048,576 tokens
Native output limit
65,536 tokens
Native input and output
Text, image, audio, video, and PDF input; text output
Reasoning and structured capabilities
Thinking, function calling, structured output
Version and knowledge cutoff
Stable gemini-2.5-pro; model updated in June 2025, knowledge cutoff January 2025
Chat invocation
/gemini/chat/completions; text and image messages, streaming responses
Session invocation
/aichat2/conversations; preserved history, file URL content blocks, and session management
Native capacity and modality descriptions reflect model capabilities; specific input methods, tool execution, and session features are provided separately by the selected platform entry point.
Core Capabilities
Learn what gemini-2.5-pro can bring to your work.
Bring long materials into a single analytical framework
Long context is suitable for organizing specifications, code, reports, and historical discussions at the same time, comparing constraints and conclusions around the same issue. Rather than simply asking for a summary, it is better suited to specifying the contradictions, dependencies, and exceptions that need to be checked, and requiring evidence identified by chapter or file, so that long-material analysis becomes a deliverable that can be reviewed and verified.
Focus on code and multi-step reasoning
Gemini 2.5 Pro's reasoning capabilities are designed for code, mathematics, and STEM problems, and can be used to explain algorithms, analyze failure conditions, or derive solutions. When submitting tasks, include input conditions, previous attempts, and acceptance criteria, and request conclusions, assumptions, and validation methods separately to facilitate subsequent testing rather than merely receiving an explanation.
From image-and-text understanding to structured results
Images can be submitted together with text questions to explain screenshots, charts, or diagrams; results are still delivered as text. Combined with structured output and function-calling capabilities, analysis can be organized into records with clearly defined fields or used to request tool calls. Tool execution and result validation should be handled by the application or an authorized session workflow.
Use Cases
Start with specific tasks to find where the model can make an impact.
Cross-file code review
Provide relevant modules, interface contracts, error logs, and reproduction steps, and ask the model to inspect call chains, boundary conditions, and potential regressions. Deliverables can include a file-by-file issue list, modification recommendations, and test-case drafts. Retain file paths and version information so developers can locate and verify recommendations rather than accepting changes directly.
Multi-document clause comparison
Organize contracts, policies, and business requirements into text, or submit accessible file links through the session interface, and ask for item-by-item comparison of scope, exceptions, and conflicting clauses. Output a difference table, questions requiring confirmation, and a summary while retaining chapter references, making it suitable for material preparation before review and follow-up questions.
Chart interpretation and technical review
Submit architecture diagrams, product screenshots, or experimental charts together with background information, clearly specifying the relationships, anomalies, or design trade-offs that need to be explained. The model can generate written analysis, review outlines, and validation checklists. For critical values, provide clear original images or text data first to help distinguish observations from the image from further inferences.
How to choose this model
Choose based on task complexity, input materials, and expected results.
Prioritize complex analysis; weigh trade-offs for simple tasks
When a task requires maintaining constraints across long materials, analyzing code relationships, or completing multi-step reasoning, Gemini 2.5 Pro is a suitable candidate. For short text classification, simple rewriting, or fixed-field extraction, include the Flash series in comparative testing and decide based on actual accuracy, response experience, and usage; there is no need to choose Pro for every task.
Evaluate stable and preview versions separately
gemini-2.5-pro is the stable version code and is not the same version as gemini-3.1-pro-preview. Existing projects can continue evaluating 2.5 Pro around current prompts, structured results, and tool workflows; when considering a version switch, run regression tests using the same materials, and do not directly regard the parameters or performance of the preview version as capabilities of this model.
Get started
From a small-scale task to production integration.
01
Prepare tasks and materials
Define the goal, required inputs, and output requirements, using real business examples as a starting point.
02
Try it in the API testing area
Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to review the results.
03
Integrate according to the API documentation
Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.
Usage boundaries
Before formal use, understand output quality and capability scope.
Native multimodal input does not mean every entry point accepts the same attachments. Chat Completions can organize text and images using text and image_url blocks; file links can use file_url in the session entry point. Do not treat PDF, recording, or video links directly as a general use of image fields.
This model outputs text and does not generate images or speech, nor is it a Live API real-time audio-video model. You can ask it to explain visual materials, write voice-over scripts, or generate drawing instructions, but actual media generation requires selecting an appropriate creative model.
Knowledge is current through January 2025, and recent facts should be supplemented with new materials. Long context also does not mean it automatically finds every detail; for important terms, code conclusions, and mathematical derivations, request clear evidence and then verify through human review, testing, or calculation.
Frequently Asked Questions
Answers to common questions about using gemini-2.5-pro.
How much material can Gemini 2.5 Pro process at once?
The native input limit is 1,048,576 tokens, and the output limit is 65,536 tokens; they are calculated separately. When organizing long material, retain chapter and file identifiers and clearly define the scope of analysis; these are the model's native limits and do not mean that every call will return the maximum length.
How do I submit screenshots together with questions?
In /gemini/chat/completions, write the message content as an array of content blocks, including both text and image_url. The text should describe the location and task to inspect, while the image provides visual information; read replies from choices, and for streaming calls, concatenate delta.content.
Which endpoint should I use to analyze PDFs?
When you need to submit a file link, you can choose /aichat2/conversations, provide the file using a file_url block in the message array, and add a text task. You can also extract the PDF text first and then send a regular message. File reading is part of the conversation workflow and should be considered separately from native PDF understanding.
Can I make Gemini 2.5 Pro run code automatically?
The model natively supports code execution and function calling, but generating code or tool parameters in a regular conversation does not mean they have already been executed. When managing tools yourself, the application must execute the function and return the result; when using the conversation tool workflow, actual operations depend on the available tools and granted permissions.
Do I need to resend all history during follow-up questions?
When using Chat Completions, the client organizes relevant history in messages. When using /aichat2/conversations, you can enable stateful and return the same id in subsequent requests so the conversation continues to retain context; set stateful: false when you do not want to retain that turn.
Model information · Updated: 2026-10-01. See the API and pricing sections for request parameters and billing rules.