Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text, images, and video. It features improved multimodal fusion with Interleaved-MRoPE for long-horizon...
| Type | Model |
| Section | Models & platforms / text+image->text |
| Pricing | paid (от $0.12/mo) |
| Platform | Self-hosted |
| Systems | api, python, self-hosted |
| Hosting | cloud |
| Site language | en |
| Vendor | Qwen |
| Launched | 2025-10-14 |