Qwen3
Qwen3 is the large language model series from Alibaba Cloud's Qwen team. It offers dense and Mixture-of-Experts models with thinking and non-thinking modes, multilingual, long-context, reasoning, coding, and agent capabilities.
Visit QwenLM/Qwen3Overview
Qwen3 is the large language model series developed by the Qwen team at Alibaba Cloud. It includes dense and Mixture-of-Experts models in sizes from 0.6B to 235B-A22B, along with Qwen3-2507 updates in Instruct and Thinking variants. The series supports thinking and non-thinking modes for reasoning and chat, multilingual use, long-context understanding, and deployment or local execution across multiple frameworks.
Key Features
- Dense and Mixture-of-Experts models in sizes including 0.6B, 1.7B, 4B, 8B, 14B, 32B, 30B-A3B, and 235B-A22B.
- Qwen3-2507 provides Instruct and Thinking variants in 235B-A22B, 30B-A3B, and 4B sizes.
- Switching between thinking mode for complex logical reasoning, math, and coding, and non-thinking mode for efficient general-purpose chat.
- Support for 100+ languages and dialects with multilingual instruction following and translation.
- Long-context understanding of 256K tokens, extendable up to 1 million tokens for Qwen3-2507.
- Agent capabilities for precise integration with external tools and complex agent-based tasks.
Use Cases
- Complex logical reasoning, mathematics, science, and coding tasks.
- Efficient general-purpose chat, creative writing, role-playing, and multi-turn dialogue.
- Multilingual instruction following and translation across 100+ languages and dialects.
- Tool use and complex agent-based tasks requiring integration with external tools.
Getting Started
- Install Transformers with transformers>=4.51.0, load a Qwen3 checkpoint with AutoModelForCausalLM and AutoTokenizer, apply the chat template, and generate text.
- ModelScope is recommended for users in mainland China and provides a Python API similar to Transformers plus the modelscope download CLI.
- Other local run options include llama.cpp, Ollama, LMStudio, ExecuTorch, MNN, MLX LM, and OpenVINO.
Deployment & Requirements
- Transformers>=4.51.0 is required for Transformers usage, and the latest version is recommended.
- SGLang>=0.4.6.post1 is required for SGLang serving.
- vLLM>=0.9.0 is recommended for vLLM serving.
- TensorRT-LLM>=0.20.0rc3 is recommended for TensorRT-LLM serving.
- llama.cpp>=b5401 is recommended for full Qwen3 support.
Before You Adopt
- No recognizable license was detected; verify permission to use, modify, and distribute the project before adoption.
- Qwen3-Instruct-2507 only supports non-thinking mode.
- Qwen3-Thinking-2507 only supports thinking mode.
- Ollama naming may not be consistent with Qwen original naming.
- Ollama default context and prediction settings can cause trouble for Qwen3 models.
- SGLang and vLLM may have suboptimal multi-step tool use quality with Qwen3 thinking models due to dropped reasoning content; the suggested workaround is to pass content as-is while fixes are pending.