Providers
Provider-specific identifiers, limits, and listed prices per million tokens. Every row links back to the provider's own documentation.
The model record exists, but no source-linked provider offer is available yet.
Submit a sourceQwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception...
Provider-specific identifiers, limits, and listed prices per million tokens. Every row links back to the provider's own documentation.
The model record exists, but no source-linked provider offer is available yet.
Submit a sourceRecorded from the source catalog and provider listings.
qwen/qwen3-vl-30b-a3b-instructEvery result stays attached to its source, version, metric and harness. Scores from different versions are never merged, and the profile below plots each benchmark against its own population rather than on a shared scale.
ModelBench does not infer quality from price, context size, or model name. When a source publishes a comparable result, it appears here with its version and link.
Read the methodologyField-level changes detected between successful source imports.
Every figure on this page traces back to one of these records.
Fetch the complete source-linked model record. No key, no account, no rate-limited tier.
GET https://model.kyssta.lol/api/v1/models/qwen/qwen3-vl-30b-a3b-instructcurl "https://model.kyssta.lol/api/v1/models/qwen/qwen3-vl-30b-a3b-instruct"Answered directly from the stored record — nothing here is generated beyond the catalog's own fields.
Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception. It is published by Alibaba Qwen and catalogued here from OpenRouter.
Qwen: Qwen3 VL 30B A3B Instruct accepts up to 262.144K tokens of context and returns up to 16.384K output tokens.
Provider catalogs list support for image input.
Open-weight experimental preview of the Qwen4 architecture: hybrid-attention MoE (125B total, 6B active) with vision encoder for coding, agent tasks, and image and video understanding
Qwen3.6 27BQwen vision-language model for visual reasoning, documents, and agent tasks
Qwen3 MaxFlagship Qwen3 model for coding agents, complex reasoning, and tool use
Qwen3.5 122B-A10BQwen vision-language model for visual reasoning, documents, and agent tasks