Qwen
Qwen: Qwen3 VL 32B Instruct
Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, and video. With 32 billion parameters, it combines deep visual perception with advanced text...
| Specifications | |
|---|---|
| Provider | Qwen |
| Input pricing | $0.10/M tokens |
| Output pricing | $0.42/M tokens |
| Context length | 131K tokens |
| Max output | 33K tokens |
| Input modalities | textimage |
| Output modalities | text |
| Tokenizer | Qwen |
| Content moderated | No |
| HuggingFace | Qwen/Qwen3-VL-32B-Instruct |
Supported Parameters
logprobsmax_tokenspresence_penaltyresponse_formatseedstructured_outputstemperaturetool_choicetoolstop_logprobstop_p