Gemini 3.1 Flash Lite Preview
retiredGoogle · Multimodal · Released Mar 2026
Google's lightweight model designed for high-volume inference with a 1M token context window.
- Strengths
- Delivers faster response times and lower overhead than larger Gemini models while maintaining competitive quality for general tasks.
- Best for
- High-throughput applications, real-time interactions, and workloads where latency matters more than maximum reasoning capability.
- Limitations
- May struggle with complex multi-step reasoning, nuanced analysis, or specialized domain tasks that benefit from larger model capacity.
Input / 1M
$0.25
Output / 1M
$1.50
Cached input / 1M
$0.025
Context window
1.05M
Price history
Input (solid)Output (dashed)
Price change
- 30d
- in decreased 0.0% out decreased 0.0%
- 90d
- in decreased 0.0% out decreased 0.0%
- 1y
- in decreased 0.0% out decreased 0.0%
- Since launch
- in decreased 0.0% out decreased 0.0%
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 29 Jun 2026 | $0.25 | $1.50 | $0.025 | Imported from OpenRouter | openrouter.ai |
| 11 Jun 2026 | $0.25 | $1.50 | $0.025 | Imported from OpenRouter | openrouter.ai |