Skip to content

OpenAI

38 modelsModel creator

Models published by OpenAI, available through the AnyRouter API. Each can route across multiple upstream providers for availability and price.

200K

Codex Mini Latest is a fast reasoning model optimized for the Codex CLI, fine-tuned from o4-mini and kept current with regular snapshots. It supports a 200,000-token input context with 100,000 output tokens, accepts text and image inputs, and exposes streaming, function calling, structured outputs, and reasoning tokens across Chat Completions and Responses endpoints. Priced at $1.50 / $6.00 per 1M input/output tokens.

1M

GPT-4.1 Mini is the smaller, faster version of GPT-4.1, retaining the same 1M+ token context window with strong instruction following and tool calling at lower cost and latency. It supports a 1,047,576-token input context with 32,768 output tokens, text and image inputs, and exposes streaming, function calling, structured outputs, fine-tuning, and predicted outputs across Chat Completions, Responses, and Realtime APIs.

Terms4
TTFT 250ms150 tok/s
1M

GPT-4.1 Nano is the fastest, most cost-efficient version of GPT-4.1, designed for ultra-low-latency instruction-following and tool-calling at high volume. It supports a 1,047,576-token input context with 32,768 output tokens, accepts text and image inputs, and exposes streaming, function calling, structured outputs, fine-tuning, and predicted outputs. Priced at $0.10 / $0.40 per 1M input/output tokens.

$2.00 in
$8.00 out
1M

GPT-4.1 is OpenAI's smartest non-reasoning model, designed to excel at instruction following and tool calling with broad knowledge across domains. It supports a 1,047,576-token input context with 32,768 output tokens, accepts text and image inputs, and exposes streaming, function calling, structured outputs, and Responses-API tools including web search, file search, image generation, code interpreter, and MCP. Fine-tuning is supported.

Terms4
TTFT 500ms80 tok/s
128K

GPT-4o Mini is OpenAI's fast, affordable small model for focused tasks. It processes text and image inputs to generate text outputs (including structured outputs) and is particularly well suited for fine-tuning and distilling larger model results into cost-effective production solutions. Supports a 128,000-token input context with 16,384 output tokens, plus streaming, function calling, and predicted outputs. Priced at $0.15 / $0.60 per 1M input/output tokens ($0.075 cached input).

$2.50 in
$10.00 out
1M

GPT-4o ("o" for "omni") is OpenAI's versatile, high-intelligence flagship multimodal model. It processes text and image inputs to generate text outputs and serves as the primary choice for most tasks outside the o-series reasoning models. The model supports a 128,000-token input context with 16,384 output tokens, and exposes streaming, function calling, structured outputs, fine-tuning, and predicted outputs across Chat Completions, Responses, Realtime, and Assistants APIs.

128K

GPT-5 Chat Latest is the snapshot of GPT-5 that powered ChatGPT, kept current with the model deployed in the consumer product. It supports a 128,000-token input context with 16,384 output tokens, accepts text and image inputs, and exposes streaming, function calling, and structured outputs along with web search, file search, image generation, code interpreter, and MCP tools.

400K

GPT-5 Codex is a version of GPT-5 optimized for agentic coding in Codex or similar environments, available exclusively through the Responses API with regular underlying snapshot refreshes. It supports a 400,000-token input context with 128,000 output tokens, text and image inputs, and exposes reasoning tokens, streaming, function calling, and structured outputs.

400K

GPT-5 Mini delivers near-frontier intelligence for cost-sensitive, low-latency, high-volume workloads. It supports a 400,000-token input context with 128,000 output tokens, text and image inputs, and exposes streaming, function calling, structured outputs, reasoning tokens, and tool integrations including web search, file search, code interpreter, and MCP. OpenAI recommends GPT-5.4 Mini for most new low-latency workloads.

400K

GPT-5 Nano is the fastest, most cost-efficient member of the GPT-5 family, optimized for summarization, classification, and other simple high-volume tasks. It supports a 400,000-token input context with 128,000 output tokens, text and image inputs, and exposes streaming, function calling, structured outputs, and tool integrations including web search, file search, image generation, code interpreter, and MCP. OpenAI recommends GPT-5.4 Nano for new speed- and cost-sensitive workloads.

400K

GPT-5.1 Codex Max is a maximum-effort variant of GPT-5.1 Codex optimized for long-horizon agentic coding inside the Codex environment, with extended context and reasoning support. Served through OpenAI's Codex surfaces.

400K

GPT-5.1 Codex Mini is a smaller, more cost-effective version of GPT-5.1 Codex, tuned for agentic coding workloads where efficiency matters more than peak capability. It supports a 400,000-token input context with 128,000 output tokens, text and image inputs, and exposes streaming, function calling, and structured outputs with reasoning-token support.

400K

GPT-5.1 Codex is a version of GPT-5.1 optimized for agentic coding inside the Codex environment, with regular underlying model refreshes. It supports a 400,000-token input context with 128,000 output tokens, accepts text and image inputs, and exposes reasoning tokens, streaming, function calling, and structured outputs. Available exclusively through the Responses API.

$1.25 in
$10.00 out
400K

GPT-5.1 is OpenAI's flagship model for coding and agentic tasks, with configurable reasoning effort (none, low, medium, high) and non-reasoning alternatives. It supports a 400,000-token input context with 128,000 output tokens, text and image inputs, and exposes streaming, function calling, and structured outputs across Chat Completions, Responses, and Realtime APIs.

400K

GPT-5.2 Codex is the Codex-environment variant of GPT-5.2, tuned for agentic coding workloads with streaming, function calling, and structured outputs. Available through OpenAI's Responses API.

$1.75 in
$14.00 out
400K

GPT-5.2 is the previous-generation frontier model for professional work, with configurable reasoning effort across the GPT-5 family. It supports a 400,000-token input context with 128,000 output tokens, accepts text and image inputs, and exposes streaming, function calling, structured outputs, and batch processing. OpenAI recommends upgrading new workloads to GPT-5.4, but GPT-5.2 remains a strong choice for cost-sensitive reasoning and coding tasks.

200K

GPT-5.3 Codex Spark is a low-latency, lower-cost variant of GPT-5.3 Codex tuned for fast iterative coding inside the Codex environment, with streaming and function calling. Served through OpenAI's Responses API.

400K

GPT-5.3 Codex is the Codex-environment variant of GPT-5.3 for agentic coding with streaming, function calling, and structured outputs.

400K

GPT-5.4 Mini is OpenAI's strongest mini model yet for coding, computer use, and subagents, delivering GPT-5.4-class capability in a faster, more cost-efficient package optimized for high-volume workloads. It supports a 400,000-token input context with 128,000 output tokens, accepts text and image inputs, and exposes reasoning tokens alongside streaming, function calling, and structured outputs.

Terms2
TTFT 250ms200 tok/s
400K

GPT-5.4 Nano is OpenAI's cheapest GPT-5.4-class model, purpose-built for simple, high-volume tasks where speed and cost matter most — classification, data extraction, ranking, and subagents. It supports a 400,000-token input context with 128,000 output tokens, text and image inputs, and ships with streaming, function calling, structured outputs, and broad tool support including web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, and MCP.

Terms2
TTFT 150ms300 tok/s
$30.00 in
$180.00 out
1.1M

OpenAI's GPT-5.4 Pro, a high-capacity model in the GPT-5 family with advanced reasoning, broad knowledge, coding proficiency, and tool-use capabilities.

$2.50 in
$15.00 out
1.1M

GPT-5.4 is OpenAI's frontier model for complex professional workloads, with the highest reasoning capability in the GPT-5 family and configurable reasoning effort (none, low, medium, high, xhigh). It features a 1M+ token context window (1,050,000 input, 128,000 output) with support for text and image inputs, enabling large-scale coding, agentic, and multimodal workflows. Strong tool integration includes web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP, and tool search.

Terms2
TTFT 600ms60 tok/s
922K

GPT-5.5 Pro uses OpenAI's Responses API with built-in tools, improved reasoning, and stateful context management. Based on GPT-5.5 with enhanced tool integration for agentic workflows.

Terms2
TTFT 600ms50 tok/s
$2.50 in
$15.00 out
1M

GPT-5.5 is OpenAI's flagship model with strong coding, reasoning, and multimodal capabilities. It features configurable reasoning effort (none, low, medium, high, xhigh) and a 922k token context window with support for text and image inputs.

Terms3
TTFT 500ms55 tok/s
1.1M

OpenAI's GPT-5.6 Luna, a reasoning-focused model in the GPT-5.6 family optimized for cost-sensitive workloads, with a 1M+ token context window and tool-use support via the Responses API.

1.1M

OpenAI's GPT-5.6 Sol, a reasoning-focused model in the GPT-5.6 family with a 1M+ token context window, image understanding, configurable reasoning effort, and tool-use support via the Responses API.

1.1M

OpenAI's GPT-5.6 Terra, a reasoning-focused model in the GPT-5.6 family that balances intelligence and cost, with a 1M+ token context window and tool-use support via the Responses API.

$1.25 in
$10.00 out
400K

GPT-5 is the previous flagship reasoning model in the GPT-5 family, built for coding and agentic tasks with configurable reasoning effort (minimal, low, medium, high). It supports a 400,000-token input context with 128,000 output tokens, text and image inputs, and exposes streaming, function calling, structured outputs, and tool integrations including web search, file search, image generation, code interpreter, and MCP. OpenAI recommends GPT-5.1 for new workloads.

Terms4
TTFT 800ms50 tok/s
1.1M

OpenAI GPT-6 Astra reasoning/chat model with ~1.05M context; Experiential Cloud promotional free daily tier.

1.1M

OpenAI's GPT-6 Luna is the efficient GPT-6 tier for focused, high-throughput tasks, with configurable reasoning, tool use, image input, and a 1.05M-token context window.

$2.00 in
$10.00 out
1.1M

OpenAI's GPT-6 Sol is built for complex coding and agentic workflows, with configurable reasoning, tool use, image input, and a 1.05M-token context window. It is the high-end GPT-6 tier below GPT-6 Astra.

128K

OpenAI gpt-oss-120b open-weight MoE via SiliconFlow free tier — reasoning and tool-use support.

131K

OpenAI's open-weight model designed for powerful reasoning, agentic tasks, and versatile developer use cases - optimized for lower latency and specialized use-cases at the edge

Terms6
TTFT 180ms150 tok/s
131K

OpenAI's open-weight safety reasoning model built on gpt-oss-20b (Apache 2.0). A 21B-parameter MoE tuned for safety tasks: content classification, LLM input/output filtering, and trust & safety judgments with reasoning traces. Use it as a guardrail/judge hop in front of or behind chat models.

$1.10 in
$4.40 out
128K

OpenAI o1-mini is a small, faster, and more affordable o-series reasoning model. It is text-only and supports a 128,000-token input context with 65,536 output tokens, exposing streaming and reasoning tokens across Chat Completions, Responses, Batch, Fine-tuning, and Assistants APIs. Note that o1-mini does not support function calling, structured outputs, or predicted outputs — OpenAI recommends o3-mini for higher intelligence at comparable latency and cost.

$15.00 in
$60.00 out
200K

OpenAI o1 is a reasoning model trained with reinforcement learning to handle complex tasks. It thinks before it answers, producing a long internal chain of thought before responding, and is well suited to math, science, coding, and multi-step analytical reasoning. Supports a 200,000-token input context with 100,000 output tokens, accepts text and image inputs, and exposes streaming, function calling, and structured outputs. Priced at $15 / $60 per 1M input/output tokens ($7.50 cached input).

$1.10 in
$4.40 out
200K

OpenAI o4-mini is a fast, cost-efficient o-series reasoning model with strong performance across coding and visual tasks, succeeded by GPT-5 Mini for new workloads. Supports a 200,000-token input context with 100,000 output tokens, accepts text and image inputs, and exposes streaming, function calling, structured outputs, and reasoning tokens across Chat Completions, Responses, Realtime, Batch, Fine-tuning, and Assistants APIs. Priced at $1.10 / $4.40 per 1M input/output tokens.

400K

GPT-5.4 Mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding, and tool use.