Qwen3-VL
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud, offering dense and MoE architectures, advanced visual reasoning, long context, and expanded OCR.
Visit QwenLM/Qwen3-VLOverview
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud. It delivers comprehensive upgrades in text understanding and generation, visual perception and reasoning, extended context length, spatial and video dynamics comprehension, and agent interaction capabilities. It is available in Dense and MoE architectures that scale from edge to cloud, with Instruct and reasoning-enhanced Thinking editions.
Key Features
- Visual Agent: Operates PC/mobile GUIs—recognizes elements, understands functions, invokes tools, completes tasks.
- Visual Coding Boost: Generates Draw.io/HTML/CSS/JS from images/videos.
- Advanced Spatial Perception: Judges object positions, viewpoints, and occlusions; provides stronger 2D grounding and enables 3D grounding.
- Long Context & Video Understanding: Native 256K context, expandable to 1M; handles books and hours-long video with full recall and second-level indexing.
- Expanded OCR: Supports 32 languages (up from 10); robust in low light, blur, and tilt; better with rare/ancient characters and jargon.
- Enhanced Multimodal Reasoning: Excels in STEM/Math—causal analysis and logical, evidence-based answers.
Use Cases
- Omni recognition for animals, plants, people, scenic spots, cars, and merchandise.
- Document parsing with layout position information and Qwen HTML format.
- Video understanding including video OCR, long video understanding, and video grounding.
- Mobile and computer-use agents for GUI control.
Getting Started
- Install transformers version 4.57.0 or later: pip install "transformers>=4.57.0".
- Use ModelScope for checkpoint downloads, especially recommended for users in mainland China.
- Load models such as Qwen3-VL-235B-A22B-Instruct with Transformers AutoModelForImageTextToText and AutoProcessor, then apply the chat template for inference.
Before You Adopt
- License: Apache-2.0. Review its terms before using, modifying, or distributing the project.
- Requires transformers version 4.57.0 or later.
- Users in mainland China are advised to use ModelScope for downloading checkpoints.