Skip to content

Qwen

20 modelsModel creator

Models published by Qwen, available through the AnyRouter API. Each can route across multiple upstream providers for availability and price.

128K

Qwen2.5 72B instruct — mature, broadly capable general chat model.

131K

Alibaba's specialized coding model with 32B parameters, excelling at code generation, completion, and debugging across 92 programming languages.

Terms2
TTFT 400ms80 tok/s
$0.12 in
$0.24 out
128K

Mid-size Qwen3 dense model balancing quality and latency.

262K

Qwen3-235B-A22B-Instruct-2507 is Qwen's updated non-thinking MoE model for multilingual instruction following, coding, tool use, and long-context work. It activates 22B of 235B parameters and provides a native 262K context window.

$0.12 in
$0.46 out
128KFree

Small Qwen3 dense model — fast, efficient, good for lightweight chat and tool use.

33K

Qwen3 Embedding 4B is the 4B size in the Qwen3 embedding and reranking family, between the 0.6B and 8B cuts. It embeds text over a 32,768-token window at 2,560 dimensions, supports 100+ languages, is instruction-aware for task prefixes, and supports Matryoshka truncation from 32 to 2,560 dimensions. It ranks below the 8B cut on MTEB multilingual but well above the 0.6B cut, at a much lower cost.

33K

Qwen3 8B embedding model for text embedding tasks with a 32,768 token context window.

262K

Qwen3.5-27B is Alibaba's largest dense Qwen3.5 model, delivering near-frontier quality across reasoning, coding, and instruction following. Features a 262K token context window (extensible to 1M), thinking/reasoning mode, tool calling, multi-token prediction, and support for 201 languages. Best suited for production deployments and complex enterprise tasks requiring top-tier performance.

131K

Alibaba's Qwen 3.5 is a 397B-parameter mixture-of-experts model with 17B active parameters, offering strong reasoning capabilities with efficient inference.

Terms5
TTFT 400ms80 tok/s
$0.10 in
$0.15 out
262K

Qwen3.5-9B is a high-performance model from Alibaba's Qwen3.5 series with a hybrid Gated Delta Networks and sparse MoE architecture. Features a 262K token context window (extensible to 1M), thinking/reasoning mode, tool calling, multi-token prediction, and support for 201 languages. Excels at reasoning, coding, instruction following, and long-context tasks.

256K

Qwen3.5 Plus is Alibaba's mid-tier "plus" model from the Qwen3.5 generation, with a 256K context window and strong reasoning, coding, and tool use.

1M

Qwen3.6 Plus is Alibaba's "plus" model from the Qwen3.6 generation with multimodal (text + image) support, a 1 million token context window, and strong reasoning, coding, and tool use.

1M

Qwen3.7 Flash is Alibaba's fast multimodal model with strong spatial understanding, real-world visual perception, and efficient inference. Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world visual perception. Available through BYOK with a user-owned API key (BYOK).

1M

Qwen3.7 Max is Alibaba's flagship "max" tier from the Qwen3.7 generation with multimodal (text + image) support, a 1 million token context window, and top-tier reasoning, coding, and tool use.

1M

Qwen3.7 Plus is Alibaba's advanced multimodal modeltext and image input support, a 1 million token context window, and strong reasoning capabilities. Available

1M

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen, the open-weight variant of Qwen3.8 Max, with 95 billion active parameters out of 2.4 trillion total and a 1,010,000-token context window.

1M

Qwen3.8-27B is Alibaba's dense Qwen3.8 model for agentic coding and complex reasoning, with route-dependent context up to 1M and native thinking mode.

1M

Alibaba Qwen native vision-language model for coding, office, long-context reasoning, and agents. 1M context.

1M

Qwen3.8 Max is Alibaba's flagship 2.4-trillion parameter MoE model from the Qwen3.8 generation with multimodal (text, image, video) support, a 1 million token context window, and top-tier reasoning, coding, and tool use.

1M

Qwen3.8 Omni Flash is Alibaba Cloud Qwen's native multimodal model for text, image, audio, and video input at 1M context. It is optimized for agentic programming, knowledge work, GUI operation, and audio/video-centered workflows. AIHubMix serves it as a distinct product from Qwen3.8 Flash.