Qwen3.7-Plus is a multimodal agent model released by Alibaba on June 2, 2026. Positioned as a unified vision-language agent foundation, it ranks among the top 5 globally and No. 1 in China on the prestigious Vision Arena benchmark.
Qwen3.7-Plus is a multimodal agent model released by Alibaba on June 2, 2026. Positioned as a unified vision-language agent foundation, it ranks among the top 5 globally and No. 1 in China on the prestigious Vision Arena benchmark.
| Plan | Price | Description |
|---|---|---|
| API International (Standard Context) | Input: $0.40 / Output: $1.60 per 1M tokens | Supports a 1M-token context window with multimodal capabilities for text, image, and video inputs. Includes Implicit Cache support with cached input pricing at $0.08 per 1M tokens. Ideal for everyday AI applications and multimodal workloads. |
| API International (Long Context) | Input: $1.20 / Output: $4.80 per 1M tokens | Designed for extended context workloads from 256K to 1M tokens, using tiered pricing. Suitable for long-document analysis, complex reasoning, and large-scale knowledge processing. |
| New User Free Credits | Free (1M tokens) | New Alibaba Cloud users receive 1M free token credits, available for 90 days after registration. Ideal for testing, prototyping, and early development. |
• GUI Agent
Understands graphical interfaces on mobile and desktop devices, and performs actions like clicking and typing just like a human. For example, it can open a stock trading app, observe its interface, and then replicate a fully functional clone from scratch.
• Visual Programming
Generates runnable web pages, SVG animations, or application code directly from images, sketches, or even videos. Simply provide a design mockup, and it can generate a complete frontend page for you.
• Visual Agent
Combines tools to solve complex problems. It can invoke a code interpreter to analyze images for "spot the difference" games or maze solving, and leverage web search to accurately identify equipment functions and parameters from a mechanical drawing.
• Multimodal Reasoning & Understanding
Demonstrates strong comprehension of images and video, scoring above Gemini 3.1 Pro on visual reasoning benchmarks. It also understands complex real-world scenes such as autonomous driving scenarios.
• Long-Horizon Task Autonomy
Integrates "see, think, write, do, and verify" into a unified workflow. In official demonstrations, it ran autonomously for over 11 hours, completing the entire development lifecycle of an English learning app—from requirements analysis and coding to testing and release.
• Individual Developers & Programming Enthusiasts: Low-cost API access for daily coding, prototyping, and learning new technologies.
• SMEs & Startups: Affordable AI integration for automation, customer support, and content production.
• Business Professionals & Content Creators: Boosting daily productivity in writing, report processing, and multimedia material handling.
• Students & Researchers: Assisting with studies, paper polishing, and data organization for moderately complex logical tasks.
| Model | Context | Pricing | API | Released | Global Heat |
|---|---|---|---|---|---|
![]()
Qwen3.7 Plus
|
1M | Paid | YES | 2026-06 |
81/100
|
| 1M | — | YES | 2026-08 |
88/100
|
|
| 262K | — | YES | 2026-08 |
90/100
|
|
| 262K | — | YES | 2026-08 |
60/100
|
|
| 1M | — | YES | 2026-07 |
79/100
|
Comments (0)