Are Gemini 3.5 Flash-Lite and 3.5 Flash the same model?
No. Flash-Lite focuses on low-latency, cost-sensitive document parsing, extraction, and lightweight sub-agent tasks. You should not directly reuse 3.5 Flash configurations for Lite. Use gemini-3.5-flash-lite when calling it, and recheck output quality and task completion after migration.
How can I make it return JSON that is easy for programs to process?
Clearly require it to return only JSON in the prompt, and specify field meanings, required fields, enums, and rules for missing values. You can include an example of the expected output. The model natively supports structured output; if you use response_format to configure a JSON object or JSON Schema, rely on the configurations actually supported by this model through that interface. After receiving the result, you still need to parse and validate the fields, especially distinguishing between correct formatting and content that is actually supported by the source material.
Which interface should I use to analyze PDFs?
For file conversations, you can use /aichat2/conversations and provide file_url and task instructions in the message. The multimodal content in Chat Completions uses text and image_url; do not submit a PDF as an image. You can also extract the text first, then pass it to the model for classification, summarization, or field organization.
Does it support image understanding, and can it also generate images or read answers aloud?
It supports image understanding and can answer questions or extract information based on images, but its native output is text and it does not support image generation or audio generation. Images can be submitted together with text instructions; if you ultimately need an image or voice asset, send the text result to the appropriate generation service.
Do I need to save the entire history myself for follow-up questions?
When using /gemini/chat/completions, provide the conversation history you need to retain through messages. When using /aichat2/conversations, you can enable stateful and include the conversation id in subsequent requests. The former allows fine-grained control over context, while the latter is suitable for hosted continuous conversations.