Technical Specifications of Gemini 3.5 Flash-Lite
| Item | Gemini 3.5 Flash-Lite |
|---|---|
| Provider | Google DeepMind |
| Model ID | gemini-3.5-flash-lite |
| Model family | Gemini 3.5 |
| Availability | General Availability (GA) |
| Input types | Text, Image, Video, Audio, PDF |
| Output types | Text |
| Context window | 1,048,576 tokens |
| Maximum output | 65,536 tokens |
| Thinking | Supported (default: minimal; also medium and high) |
| Function calling | Yes |
| Code execution | Yes |
| File Search | Yes |
| URL Context | Yes |
| Search Grounding | Yes |
| Structured Output | Yes |
| Computer Use | Not supported |
| Caching | Supported |
| Primary focus | High-throughput, low-latency, low-cost inference |
What is Gemini 3.5 Flash-Lite?
Gemini 3.5 Flash-Lite is Google's fastest and most cost-effective model in the Gemini 3.5 family. It is optimized for workloads where latency, throughput, and API cost are more important than maximum reasoning performance. Typical applications include document parsing, structured data extraction, routing, lightweight coding, and autonomous subagents operating at scale. Google positions it as the recommended upgrade path from Gemini 3.1 Flash-Lite and Gemini 2.5 Flash for production deployments.
Main Features of Gemini 3.5 Flash-Lite
- Optimized for high-volume, low-cost inference.
- Supports a 1 million-token context window for long documents and repositories.
- Native multimodal inputs including text, images, video, audio, and PDF.
- Supports Function Calling, Code Execution, File Search, URL Context, Search Grounding, and Structured Outputs.
- Configurable thinking levels (
minimal,medium,high) to balance speed and reasoning quality. - Significantly improves coding, reasoning, and agentic workflows compared with Gemini 3.1 Flash-Lite while maintaining very low latency.
Benchmark Performance of Gemini 3.5 Flash-Lite
Google states that Gemini 3.5 Flash-Lite substantially outperforms previous Flash-Lite generations in coding, document understanding, and agentic workflows while remaining the lowest-cost model in the Gemini 3.5 lineup. Gemini 3.5 Flash-Lite scored 54% in the Terminal-Bench 2.1 test, a significant improvement over its predecessor's 31%.
Its strengths lie in classification, extraction, chat replies, and simple retrieval-enhanced answers. These are precisely where fast, low-cost models can excel.
Gemini 3.5 Flash-Lite vs Gemini 3.6 Flash vs Gemini 3.5 Flash
| Aspect | Gemini 3.5 Flash-Lite | Gemini 3.6 Flash | Gemini 3.5 Flash |
|---|---|---|---|
| Positioning | Fastest & cheapest for high-volume everyday tasks (e.g., document processing, search) | Current flagship Flash: Best balance of intelligence, efficiency & agentic performance | Strong agentic/coding model (now largely superseded by 3.6) |
| Best For | High-throughput, low-cost workflows | Coding, multimodal, complex agents, knowledge work | General agentic & coding tasks |
| Intelligence / Performance | Good (outperforms older Lites; competitive on many agentic tasks) | Highest among the three (noticeable gains in coding & agentic) | Strong (near-frontier for Flash series) |
| Key Benchmarks | - Terminal-Bench: ~54%- Strong on document & high-volume tasks | - DeepSWE: 49% (vs 37%)- MLE-Bench: 63.9% (vs 49.7%)- OSWorld: 83% (vs 78.4%) | - Terminal-Bench: 76.2%- Lower than 3.6 on most coding/agentic |
| Speed | Fastest (~350 output tokens/sec) | Very fast (~280 t/s) | Fast |
| Efficiency | Excellent for volume | Best (17% fewer output tokens on average; up to 65% on coding) | Good |
| Multimodal | Text + Image + Video + Audio | Text + Image + Video + Audio (stronger reasoning) | Text + Image + Video + Audio |
| Context Window | ~1M tokens | ~1M tokens | ~1M tokens |
| Max Output Tokens | ~64k | ~64k | ~64k |
| Pricing (per 1M tokens) | Cheapest: ~$0.30 input / $2.50 output | $1.50 input / $7.50 output | $1.50 input / $9.00 output |
| Effective Cost | Lowest per task for simple work | Lower than 3.5 due to efficiency gains | Higher than 3.6 |
Summary Recommendations
- Pick 3.5 Flash-Lite — when you need maximum speed and minimum cost for high-volume or simple tasks.
- Pick 3.6 Flash — for most users and developers (best overall choice right now).
- Pick 3.5 Flash — only if you have existing workflows tied to it (plan to upgrade to 3.6).
Limitations of Gemini 3.5 Flash-Lite
- Optimized for efficiency rather than frontier-level reasoning.
- Computer Use is not supported.
- Native image and audio generation are unavailable.
- Fine-tuning is not currently supported.
- Complex multi-step reasoning tasks are generally better suited to Gemini 3.5 Flash or Gemini 3.6 Flash.