Models
59 modelsBrowse the full range of AI models, covering text, reasoning, vision, coding, image, video, embedding and more
qwen3.7-plusQwen3.7 series cost-effective Plus model, with fully upgraded visual-language capabilities on top of strong text abilities, retaining complete agent capabilities for coding, tool use, and productivity workflows. Supports multimodal interactive hybrid agents: perceiving real-world scenes, reading screens and operating GUIs, generating code based on visual references, and end-to-end navigation of mobile apps. Equivalent to snapshot qwen3.7-plus-2026-05-26.
qwen3.7-maxQwen3.7 flagship model, built for the agent era with comprehensive improvements in coding, office productivity, and long-cycle autonomous execution. Supports thinking mode toggle, function calling, and web search. 1M token context window.
qwen3-maxQwen3 most powerful flagship model, supports thinking mode toggle, excels in complex reasoning, code generation, and mathematics. 262K context window.
qwen3.6-max-previewQwen3.6 most powerful preview model, designed for complex reasoning, code generation, and multi-step tool tasks, ideal for scenarios requiring stronger thinking capabilities.
qwen3.6-plusQwen3.6 balanced flagship model, supports 1M context window, function calling, and built-in tools, ideal for large codebases and general production scenarios.
qwen3.5-plusQwen3.5 enhanced model, best balance of quality, speed, and cost. Supports 1M context window, ideal for large-scale application scenarios.
qwen3.6-flashQwen3.6 flash model, ideal for simple tasks with fast speed and low cost. Supports 1M context window and context caching.
qwen3.5-flashQwen3.5 flash model, ideal for simple tasks with fast speed and low cost. Supports 1M context window and context caching.
qwen-plusQwen enhanced model, classic balance of quality and speed, ideal for large-scale application scenarios.
qwen-turboQwen high-speed model, extremely fast response and lowest cost, ideal for latency-sensitive application scenarios.
qwen-longQwen long-context model, supports ultra-long context windows up to 10M tokens, ideal for document analysis and long-text understanding.
qwen-flashQwen ultra-fast general model, 1M context window, extremely fast response and ultra-low cost, ideal for large-scale high-concurrency scenarios. Supports function calling and thinking mode.
qwen3-235b-a22bQwen3 open-source flagship, 235B parameter MoE architecture (22B active), supports dynamic switching between thinking and non-thinking modes.
qwen3.6-35b-a3bQwen3.6 open-source MoE model, 35B total parameters with only 3B active, excels at agent coding, STEM, and reasoning tasks. Apache 2.0 licensed. Supports thinking mode toggle.
qwen3-32bQwen3 open-source 32B parameter dense model, excels among medium-scale models.
qwq-plusQwen reasoning model, trained on Qwen2.5, excels at mathematics, logical reasoning, and complex problem analysis, displaying complete chain-of-thought.
qwen-vl-maxQwen vision flagship model, supports image understanding, visual-text dialogue, document OCR, and other multimodal tasks.
qwen-vl-plusQwen vision enhanced model, balanced performance and cost multimodal model.
qwen3-vl-plusQwen3 vision-language model, significantly improved image understanding, supports high-resolution image input. 262K context window.
qwen3-vl-flashQwen3 vision flash model, fast image understanding, ideal for real-time scenarios.
qwen3.5-omni-plusQwen3.5 flagship omni model, supports any combination of text, image, audio, and video input with text and voice output. Up to 3 hours audio / 1 hour video input, 113 input languages, 55 voice tones, supports web search and voice cloning.
qwen3.5-omni-flashQwen3.5 lightweight omni model, supports any combination of text, image, audio, and video input with text and voice output. Up to 3 hours audio / 1 hour video input, 113 input languages, 55 voice tones, supports web search. Best cost-efficiency choice.
qwen3-omni-flashQwen3 omni model, accepts text, image, audio, and video inputs with text and voice output. Supports thinking mode (text-only output in thinking mode). Ideal for short video analysis and cost-sensitive scenarios.
qwen3-coder-plusQwen3 exceptional code model, excels at tool calling and environment interaction, with outstanding code generation, completion, debugging, and refactoring capabilities. 1M context window.
qwen3-coder-flashQwen3 code flash model, fast code completion and generation, ideal for IDE integration scenarios.
qwen-math-plusQwen math-specialized model, excels at mathematical problem solving, proofs, and computation, supports LaTeX format output.
qwen-mt-plusQwen flagship translation model, supports 92 language pairs, outstanding translation quality, ideal for professional translation scenarios.
tongyi-intent-detect-v3Qwen intent understanding model, rapidly and accurately parses user intent within milliseconds, suitable for customer service routing, intelligent dialogue distribution, and instruction parsing scenarios.
text-embedding-v4Qwen latest text embedding model, supports 100+ languages and multiple programming languages, vector dimensions selectable from 2048, 1536, 1024, 768, 512, 256, 128, 64, suitable for semantic retrieval, clustering, recommendation, and RAG.
text-embedding-v3Qwen text embedding model, converts text into high-dimensional vector representations, suitable for semantic search, clustering, and recommendation scenarios.
qwen3-asr-flashQwen3 speech recognition model, supports automatic detection and transcription in 11 languages, word-level timestamps, emotion recognition, singing recognition, and speaker diarization. Supports both real-time and non-real-time modes.
qwen3-tts-flash-realtimeQwen3 real-time text-to-speech model, streaming synthesis via WebSocket, supports multiple languages and voice tones including Chinese and English, suitable for voice assistants and audiobooks.
wan2.6-t2iLatest generation text-to-image flagship model, supports mixed text-image output and image editing. Can process complex instructions, render Chinese and English text, and generate high-definition realistic images. Supports multiple resolutions and aspect ratios.
wan2.6-t2vLatest generation text-to-video flagship model, supports multi-shot narrative and intelligent storyboard. Can generate 2-15 second 1080P HD video, supports prompt rewriting. Generation time approximately 1-5 minutes.
wan2.6-i2vImage-driven video generation model, uses the input image as the first frame to generate coherent video. Supports multi-shot narrative, automatic dubbing, 720P/1080P resolution, 2-15 seconds duration. Excellent frame coherence and motion consistency.
wan2.6-i2v-flashFast image-to-video model, supports audio and silent video generation. Faster generation speed, ideal for latency-sensitive scenarios. Supports 720P/1080P, 2-15 seconds duration.
wan2.6-r2vMultimodal input video generation model, supports text/image/video as references. Can use characters or objects as protagonists to generate single-character or multi-character interaction videos. 2-10 seconds duration, supports intelligent storyboarding.
wan2.6-r2v-flashFast reference-to-video model, supports audio and silent output. Faster generation speed, ideal for rapid iteration scenarios. Supports 720P/1080P resolution.
pixverse-v6PixVerse latest flagship video generation model, supports text-to-video and image-to-video, with significantly improved visual quality and motion consistency. Supports 1-15 seconds duration, 360p/540p/720p/1080p multiple resolutions, various aspect ratios.
happyhorse-1.0-t2vAlibaba's 2026 latest AI video generation model, ranked #1 on benchmarks. Generates high-quality video from text, supports 720P/1080P, 3-15 seconds duration, various aspect ratios. Default audio included.
happyhorse-1.0-i2vGenerates coherent video using the input image as the first frame, supports 720P/1080P, 3-15 seconds duration. Excellent frame coherence and motion consistency. Default audio included.
happyhorse-1.0-r2vSupports 1-9 reference images input, can fuse characters/objects/scenes from images to generate video. Supports 720P/1080P, 3-15 seconds, various aspect ratios. Default audio included.
happyhorse-1.0-video-editAI video editing based on input video, supports 0-5 reference images for assisted editing. Input video 3-60 seconds (truncated beyond 15s), supports 720P/1080P, can preserve original audio.
deepseek-v4-flashDeepSeek V4 Flash high-speed model via DashScope, ideal for low-latency and high-concurrency online dialogue scenarios.
deepseek-v4-proDeepSeek V4 Pro flagship model via DashScope, designed for complex reasoning, code generation, and multi-step tasks.
deepseek-v3.2DeepSeek latest general-purpose LLM, MoE architecture, strong bilingual Chinese-English capabilities, powerful coding abilities.
deepseek-r1DeepSeek reasoning model with outstanding performance in mathematics, coding, and logical reasoning, displaying complete chain-of-thought.
deepseek-v3DeepSeek V3 general-purpose LLM, 671B parameter MoE architecture, excellent bilingual Chinese-English capabilities.
claude-opus-4-7Anthropic's most capable general-purpose model, designed for complex reasoning, agentic coding, and long-context tasks. Official pricing: $5 input / $25 output per 1M tokens.
claude-sonnet-4-6Anthropic's balanced speed and intelligence model, ideal for production-grade dialogue, code, tool use, and long-context workflows. Official pricing: $3 input / $15 output per 1M tokens.
claude-haiku-4-5Anthropic's fast and low-cost model with near-frontier intelligence, ideal for low-latency dialogue, classification, extraction, and batch tasks. Official pricing: $1 input / $5 output per 1M tokens.
glm-4.7Zhipu AI latest GLM-4.7 model with significantly improved overall capabilities and strong Chinese language understanding.
glm-5Zhipu AI GLM-5 flagship model with comprehensive capability improvements, outstanding performance in reasoning, coding, and long-text tasks.
glm-5.1Zhipu AI GLM-5.1 enhanced flagship model, further optimized over GLM-5, stronger complex reasoning and code generation capabilities.
kimi-k2.5Moonshot AI Kimi K2.5 model, excels at long-text understanding and multi-turn dialogue with outstanding Chinese language capabilities.
kimi-k2.6Moonshot AI Kimi K2.6 latest flagship model with significantly improved long-text understanding and creative writing, supports longer context windows.
MiniMax-M2.1MiniMax M2.1 model, outstanding performance in creative writing and multi-turn dialogue.
MiniMax-M2.5MiniMax M2.5 enhanced model with improved reasoning and coding capabilities, more stable multi-turn dialogue.
qwen3-8bQwen3 open-source 8B parameter lightweight model, ideal for edge deployment and low-cost inference scenarios.