All models

gemini-3.5-flash-lite

GoogleChatReasoningVision
Get your API key
gemini-3.5-flash-lite

Multimodal model for high-frequency document parsing and lightweight agents

Gemini 3.5 Flash-Lite is a multimodal model from Google designed for low-latency, cost-sensitive tasks, with a focus on document parsing, simple data extraction, and clearly scoped sub-agent tasks. It combines long-input processing, image understanding, reasoning, and structured output capabilities, making it suitable for organizing large amounts of repetitive work into usable data rather than turning every request into deep analysis with long chains.

GoogleModel brand
ChatModel type
Reasoning, visual understandingTask capabilities

Specifications and interface features

Clarify capacity, input and output, and invocation methods before selecting a model.

Native input limit
1,048,576 tokens
Native output limit
65,536 tokens
Native input and output
Text, image, video, audio, and PDF input; text output
Native task capabilities
Thinking, function calling, structured output, caching
Image and text invocation method
Combine text and image_url in messages; streaming responses supported
Files and continuous sessions
The session endpoint supports file_url, stateful, and session id

Native specifications describe model capabilities; image-and-text requests and file sessions on this platform use their respective endpoints, and capacity figures do not represent the available quota for each request.

Core Capabilities

Learn what gemini-3.5-flash-lite can bring to your work.

Turn long materials into clear fields

Flash-Lite is optimized for document parsing and simple extraction, organizing reports, records, or explanatory text into fields, categories, and summaries. Clearly specifying field definitions, missing-value rules, and output formats in the task is better aligned with its positioning than broadly requesting comprehensive analysis; when results need to be consumed by machines, structured output and type validation can be used together.

Understand text and images together, with text results

It can understand images in combination with text instructions, making it suitable for extracting specified information from screenshots, forms, or charts. For multimodal requests, place text and image_url in the same content array to keep questions aligned with visual materials. Deliverables can be explanations or structured text, but it does not directly generate images or synthesize answers into speech.

Handle focused sub-agent tasks with clear responsibilities

Native function calling and reasoning capabilities enable it to handle focused tasks such as information verification and parameter organization. When designing workflows, provide tool purposes, required parameters, and completion criteria, and break complex goals into verifiable steps. Function calling expresses operational intent; actual execution still needs to be completed by the application or configured tools.

Use Cases

Start with specific tasks to find where the model can be effective.

Customer service record classification and organization

Input customer service conversations or ticket text, require issue types to be identified according to predefined labels, and extract products, requests, and information that needs to be supplemented, delivering records with consistent fields. Labels lacking sufficient evidence may be returned as pending confirmation and then reviewed manually for difficult cases, making this suitable for embedding frequent, repetitive organization work into business processes.

Document and form information entry

Input document text, form images, or provide file links through a conversation endpoint, specify the required names, dates, amounts, and corresponding source text, and generate structured results suitable for database entry. First limit the extraction scope, then check field completeness and correspondence with the source text to avoid allowing summary content to replace precise data records.

Continue asking questions about the material

First submit the material and generate a summary of key points, then ask follow-up questions about differences, items, or missing information in the same material. Using a conversation endpoint allows you to save the returned id to continue the discussion; existing custom chat systems can carry history through messages. Final deliverables can include summaries, question lists, and items requiring further confirmation.

How to choose this model

Choose based on task complexity, input materials, and expected results.

Prioritize Lite for routine organization; evaluate difficult tasks separately

When the task involves classification, simple extraction, document parsing, or clearly scoped sub-agent work, the low-latency, cost-friendly positioning of Flash-Lite is more suitable. If requirements shift toward complex coding, long-chain reasoning, or continuous autonomous execution, consider comparing the full Flash or Pro. Evaluate field accuracy, task completion rate, and rework volume using the same set of samples, rather than choosing based solely on the series name.

Differentiate models, and differentiate workflows

gemini-3.5-flash-lite, gemini-3.5-flash, and gemini-3.1-flash-lite are different models and should not be treated as spelling aliases. When directly managing history, output formats, and tool callbacks, choose Chat Completions; when you need to save sessions, submit file links, and manage conversations, choose the sessions entry point. When migrating existing applications, retain representative samples to verify results.

Get started

From a small-scale task to production integration.

01

Prepare tasks and materials

Define objectives, required inputs, and output requirements, using real business examples as a starting point.

02

Try it in the API testing area

Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to review the results.

03

Integrate according to the API documentation

Keep the full model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.

Usage boundaries

Before formal use, understand output quality and capability boundaries.

  • Large input capacity does not mean every piece of material can be extracted with the same accuracy. For cross-document fields, similar names, and scattered entries, require the corresponding original text to be included; excessively long and irrelevant materials also increase processing burden, so prioritize retaining content needed for the task and split and summarize it when necessary.
  • This is a comprehension and text output model and does not support native image generation, audio generation, or the Live API. Native support for understanding video and audio does not mean arbitrary media links can be placed in image fields; file tasks should use appropriate submission methods and avoid mixing different content types.
  • A lightweight sub-agent should not be treated as a fully autonomous execution system without supervision. For multi-step tasks, verify whether tools actually completed their actions, whether returned results are valid, and whether stop conditions are met; when writing, sending, or publishing is involved, set clear permissions and retain confirmation and recovery steps after failures.

Frequently Asked Questions

Answers to common questions about using gemini-3.5-flash-lite.

Are Gemini 3.5 Flash-Lite and 3.5 Flash the same model?

No. Flash-Lite focuses on low-latency, cost-sensitive document parsing, extraction, and lightweight sub-agent tasks. You should not directly reuse 3.5 Flash configurations for Lite. Use gemini-3.5-flash-lite when calling it, and recheck output quality and task completion after migration.

How can I make it return JSON that is easy for programs to process?

Clearly require it to return only JSON in the prompt, and specify field meanings, required fields, enums, and rules for missing values. You can include an example of the expected output. The model natively supports structured output; if you use response_format to configure a JSON object or JSON Schema, rely on the configurations actually supported by this model through that interface. After receiving the result, you still need to parse and validate the fields, especially distinguishing between correct formatting and content that is actually supported by the source material.

Which interface should I use to analyze PDFs?

For file conversations, you can use /aichat2/conversations and provide file_url and task instructions in the message. The multimodal content in Chat Completions uses text and image_url; do not submit a PDF as an image. You can also extract the text first, then pass it to the model for classification, summarization, or field organization.

Does it support image understanding, and can it also generate images or read answers aloud?

It supports image understanding and can answer questions or extract information based on images, but its native output is text and it does not support image generation or audio generation. Images can be submitted together with text instructions; if you ultimately need an image or voice asset, send the text result to the appropriate generation service.

Do I need to save the entire history myself for follow-up questions?

When using /gemini/chat/completions, provide the conversation history you need to retain through messages. When using /aichat2/conversations, you can enable stateful and include the conversation id in subsequent requests. The former allows fine-grained control over context, while the latter is suitable for hosted continuous conversations.

Model information · Updated: 2026-10-01. For call parameters and billing rules, see the API and pricing sections.

Use gemini-3.5-flash-lite for your next task

Start with clear goals and determine from real-world results whether it fits your work.