Alibaba / Qwen AI Models

Explore 45 curated AI models from Alibaba / Qwen and compare context, API pricing, modalities and capabilities.

45models tracked

Model directory

45 models in this view

Compare

Alibaba / Qwen

Qwen3.8 27B

Active

An open-weight vision-language model listed for coding, research, multimodal interaction and agent tasks.

Context
262K
Input
$0.45
Output
$3.2
TextImageVideo
View model

Alibaba / Qwen

Qwen3 Coder Next

Active

An open-weight coding model listed for coding agents and local development workflows.

Context
262K
Input
$0.12
Output
$0.80
Text
View model
Active

Qwen Image 3 is a unified image generation and editing model from Qwen. It supports precise rendering of text and details as small as 10px, along with a richer world…

Context
66K
Input
$0.00
Output
$0.00
TextImage
View model
Active

Qwen Image 3 Pro is an image generation and editing model from Qwen. It supports precise rendering of text and details as small as 10px, along with richer world knowledge…

Context
66K
Input
$0.00
Output
$0.00
TextImage
View model
Active

Qwen Plus 0728, based on the Qwen3 foundation model, is a 1 million context hybrid reasoning model with a balanced performance, speed, and cost combination.

Context
1M
Input
$0.26
Output
$0.78
Text
View model

Qwen-Audio-3.0-TTS Flash is Alibaba's fast, cost-efficient text-to-speech model, generating spoken audio from text via the DashScope Speech Synthesizer API.

Context
0
Input
$15
Output
$0.00
Text
View model

Qwen-Audio-3.0-TTS Plus is Alibaba's higher-quality text-to-speech model, generating spoken audio from text via the DashScope Speech Synthesizer API.

Context
0
Input
$20
Output
$0.00
Text
View model

Over the past three months, we have continued to scale the **thinking capability** of Qwen3-30B-A3B, improving both the **quality and depth** of reasoning. We are pleased to introduce **Qwen3-30B-A3B-Thinking-2507**, featuring the following key enhancements:

Context
82K
Input
$0.20
Output
$2.4
Text
View model
Active

The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects. Both leverage large-scale speech training data and the strong audio understanding capability of their foundation model, Qwen3-Omni. Experiments show that the 1.7B version achieves state-of-the-art performance among open-source ASR models and is competitive with the strongest proprietary commercial APIs. Here are the main features:

Context
0
Input
$3.33
Output
$0.00
Audio
View model
Active

The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects. Both leverage large-scale speech training data and the strong audio understanding capability of their foundation model, Qwen3-Omni. Experiments show that the 1.7B version achieves state-of-the-art performance among open-source ASR models and is competitive with the strongest proprietary commercial APIs. Here are the main features:

Context
0
Input
$7.5
Output
$0.00
Audio
View model
Active

Qwen3-ASR-Flash is Alibaba's automatic speech recognition service, built on the Qwen3-Omni foundation and trained on tens of millions of hours of multimodal speech data. The model handles 11 languages —…

Context
0
Input
$35
Output
$0.00
Audio
View model
Active

Qwen3 Coder Flash is Alibaba's fast and cost efficient version of their proprietary Qwen3 Coder Plus. It is a powerful coding agent model specializing in autonomous programming via tool calling…

Context
1M
Input
$0.195
Output
$0.975
Text
View model
Active

Qwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B. It is a powerful coding agent model specializing in autonomous programming via tool calling and…

Context
1M
Input
$0.65
Output
$3.25
Text
View model
Active

The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. Building upon the dense foundational models of the Qwen3 series, it provides a comprehensive range of text embeddings and reranking models in various sizes (0.6B, 4B, and 8B). This series inherits the exceptional multilingual capabilities, long-text understanding, and reasoning skills of its foundational model. The Qwen3 Embedding series represents significant advancements in multiple text embedding and ranking tasks, including text retrieval, code retrieval, text classification, text clustering, and bitext mining.

Context
33K
Input
$0.02
Output
$0.00
Text
View model
Active

The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. Building upon the dense foundational models of the Qwen3 series, it provides a comprehensive range of text embeddings and reranking models in various sizes (0.6B, 4B, and 8B). This series inherits the exceptional multilingual capabilities, long-text understanding, and reasoning skills of its foundational model. The Qwen3 Embedding series represents significant advancements in multiple text embedding and ranking tasks, including text retrieval, code retrieval, text classification, text clustering, and bitext mining.

Context
33K
Input
$0.01
Output
$0.00
Text
View model
Active

Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and long-tail knowledge coverage compared to the January 2025 version. It…

Context
262K
Input
$0.78
Output
$3.9
Text
View model
Active

Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi-step reasoning. By significantly scaling model capacity and reinforcement learning compute, it…

Context
262K
Input
$0.78
Output
$3.9
Text
View model

Over the past few months, we have observed increasingly clear trends toward scaling both total parameters and context lengths in the pursuit of more powerful and agentic artificial intelligence (AI). We are excited to share our latest advancements in addressing these demands, centered on improving scaling efficiency through innovative model architecture. We call this next-generation foundation models **Qwen3-Next**.

Context
262K
Input
$0.10
Output
$1.1
Text
View model

Over the past few months, we have observed increasingly clear trends toward scaling both total parameters and context lengths in the pursuit of more powerful and agentic artificial intelligence (AI). We are excited to share our latest advancements in addressing these demands, centered on improving scaling efficiency through innovative model architecture. We call this next-generation foundation models **Qwen3-Next**.

Context
131K
Input
$0.15
Output
$1.2
Text
View model
Active

Qwen3 Reranker 8B is a text reranking model from Alibaba Cloud built on the Qwen3 architecture. It evaluates query-document pairs to produce relevance scores for use in retrieval and RAG…

Context
41K
Input
$0.00
Output
$0.00
Text
View model

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

Context
131K
Input
$0.21
Output
$1.9
TextImage
View model

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

Context
131K
Input
$0.40
Output
$4
TextImage
View model

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

Context
131K
Input
$0.13
Output
$0.52
TextImage
View model

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

Context
131K
Input
$0.20
Output
$2.4
TextImage
View model

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

Context
131K
Input
$0.104
Output
$0.416
TextImage
View model

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

Context
131K
Input
$0.117
Output
$0.455
ImageText
View model

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

Context
131K
Input
$0.18
Output
$2.1
ImageText
View model
Active

> [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

Context
262K
Input
$0.39
Output
$2.34
TextImageVideo
View model

The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of…

Context
1M
Input
$0.26
Output
$1.56
TextImageVideo
View model

Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text output, with a 1M token context window. This…

Context
1M
Input
$0.30
Output
$1.8
TextImageVideo
View model

> [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

Context
262K
Input
$0.26
Output
$2.08
TextImageVideo
View model
Active

> [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

Context
262K
Input
$0.195
Output
$1.56
TextImageVideo
View model
Active

> [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

Context
262K
Input
$0.225
Output
$1.8
TextImageVideo
View model
Active

> [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

Context
262K
Input
$0.10
Output
$0.15
TextImageVideo
View model
Active

The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the…

Context
1M
Input
$0.065
Output
$0.26
TextImageVideo
View model
Active

> [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

Context
262K
Input
$0.60
Output
$3.6
TextImageVideo
View model
Active

> [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

Context
262K
Input
$0.14
Output
$1
TextImageVideo
View model
Active

Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context window. Tiered pricing kicks in…

Context
1M
Input
$0.1875
Output
$1.13
TextImageVideo
View model

Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and…

Context
262K
Input
$1.03
Output
$6.16
Text
View model
Active

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers…

Context
1M
Input
$0.325
Output
$1.95
TextImageVideo
View model
Active

Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world…

Context
1M
Input
$0.03
Output
$0.13
TextImageVideo
View model
Active

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,…

Context
1M
Input
$1.48
Output
$4.43
Text
View model
Active

Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its…

Context
1M
Input
$0.32
Output
$1.28
TextImage
View model
Active

> [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with vLLM, SGLang, TokenSpeed, etc.

Context
1M
Input
$2
Output
$6
Text
View model
Active

Qwen3.8 Max is the flagship model in Alibaba's Qwen3.8 series, the general-availability successor to the Qwen3.8 Max Preview. It is a multimodal reasoning model intended for complex reasoning, visual understanding,…

Context
1M
Input
$2
Output
$6
TextImageVideo
View model
OUR METHOD

How AIToolly handles model data

Model-native facts and provider endpoint facts stay separate. Third-party catalogs support discovery and provider snapshots, while official documentation and verified model cards take priority for model facts.

01Sources and verification
02Verified
03Capability