prices synced 2026-09-11

Gemini 3.1 Flash Lite

Google · Multimodal · Released May 2026

Compare

Google's lightweight text-only model with a 1M token context window, built on the Gemini 3.1 Flash architecture for efficient inference.

Strengths
Delivers fast response times and low latency while maintaining a large context window, making it suitable for applications that prioritize speed without sacrificing document-handling capacity.
Best for
High-volume text processing, real-time applications, and latency-sensitive workloads where quick throughput matters more than maximum reasoning capability.
Limitations
Text-only with no multimodal support; trades reasoning depth for speed compared to larger or more capable Gemini models, making it less suited for complex analytical tasks.

Input / 1M

$0.25

Output / 1M

$1.50

Cached input / 1M

$0.025

Context window

1.05M

Also billed

Cache write
$0.0833 / 1M
Audio input
$0.5 / 1M

Price history

Input (solid)Output (dashed)
Price per 1M tokens over timeOutput went from $1.50 to $1.50; Input went from $0.25 to $0.25 between Jun 2026 and Jun 2026.$0$0.5$1.00$1.50Output on 11 Jun 2026: $1.50Output on 29 Jun 2026: $1.50Out $1.50Input on 11 Jun 2026: $0.25Input on 29 Jun 2026: $0.25In $0.25Jun 2026Jun 2026

Price change

30d
in out
90d
in decreased 0.0% out decreased 0.0%
1y
in decreased 0.0% out decreased 0.0%
Since launch
in decreased 0.0% out decreased 0.0%

Snapshots

Effective Input Output Cached in Note Source
29 Jun 2026 $0.25 $1.50 $0.025 Imported from OpenRouter openrouter.ai
11 Jun 2026 $0.25 $1.50 $0.025 Imported from OpenRouter openrouter.ai

More from Google

Report a problem