prices synced 2026-09-21

GPT-3.5 Turbo (batch)

OpenAI · Released May 2023

Compare

GPT-3.5 Turbo optimized for batch processing through OpenAI's Batch API, processing requests asynchronously with a 16K token context window.

Strengths
Delivers lower latency and reduced cost for high-volume processing compared to the standard GPT-3.5 Turbo endpoint when latency tolerance allows for queued execution.
Best for
Large-scale batch jobs like bulk summarization, data classification, or content generation where results can be retrieved asynchronously rather than in real time.
Limitations
Requires asynchronous batch submission and retrieval rather than synchronous API calls, and sits below GPT-4o-mini and later models in capability; for synchronous chat workloads or real-time interaction, use the standard GPT-3.5 Turbo endpoint.

Input / 1M

$0.25

Output / 1M

$0.75

Cached input / 1M

Context window

16K

Price history

Price per 1M tokens over timeOutput $0.75; Input $0.25 as of Aug 2026.$0$0.2$0.4$0.6$0.8Output on 6 Aug 2026: $0.75Out $0.75Input on 6 Aug 2026: $0.25In $0.25Aug 2026

Snapshots

Effective Input Output Cached in Note Source
6 Aug 2026 $0.25 $0.75 Imported from OpenRouter openrouter.ai

More from OpenAI

Report a problem