Gemini 2.5 Flash Lite Preview 09-2025
retiredGoogle · Multimodal · Released Sep 2025
A lightweight model in Google's Gemini 2.5 family optimized for low-latency and cost-efficient inference with a 1M token context window.
- Strengths
- Delivers fast token generation and high throughput, making it suitable for real-time applications where response speed matters.
- Best for
- Latency-sensitive workloads and high-volume inference where a smaller, faster model is preferred over maximum capability.
- Limitations
- Being a lite variant, it sacrifices some reasoning depth and complex task handling compared to the full Gemini 2.5 model.
Input / 1M
$0.1
Output / 1M
$0.4
Cached input / 1M
$0.01
Context window
1.05M
Price history
Input (solid)Output (dashed)
Price change
- 30d
- in decreased 0.0% out decreased 0.0%
- 90d
- in decreased 0.0% out decreased 0.0%
- 1y
- in decreased 0.0% out decreased 0.0%
- Since launch
- in decreased 0.0% out decreased 0.0%
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 29 Jun 2026 | $0.1 | $0.4 | $0.01 | Imported from OpenRouter | openrouter.ai |
| 11 Jun 2026 | $0.1 | $0.4 | $0.01 | Imported from OpenRouter | openrouter.ai |