GPT-3.5 Turbo (batch)
OpenAI · Released May 2023
GPT-3.5 Turbo optimized for batch processing through OpenAI's Batch API, processing requests asynchronously with a 16K token context window.
- Strengths
- Delivers lower latency and reduced cost for high-volume processing compared to the standard GPT-3.5 Turbo endpoint when latency tolerance allows for queued execution.
- Best for
- Large-scale batch jobs like bulk summarization, data classification, or content generation where results can be retrieved asynchronously rather than in real time.
- Limitations
- Requires asynchronous batch submission and retrieval rather than synchronous API calls, and sits below GPT-4o-mini and later models in capability; for synchronous chat workloads or real-time interaction, use the standard GPT-3.5 Turbo endpoint.
Input / 1M
$0.25
Output / 1M
$0.75
Cached input / 1M
—
Context window
16K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 6 Aug 2026 | $0.25 | $0.75 | — | Imported from OpenRouter | openrouter.ai |