Qwen3-8 Omni Flash Undercuts Gemini Flash Pricing While Matching Multimodal Benchmarks
Alibaba’s Qwen team has released Qwen3-8 Omni Flash, a multimodal AI model that matches Google’s Gemini Flash on key benchmarks while offering significantly lower pricing.
The new model handles text, images, audio, video, and file inputs. It is designed for real-time applications like voice assistants, video analysis, and document processing.
Pricing starts at $0.30 per million input tokens and $0.60 per million output tokens. This undercuts Gemini Flash by roughly 60%, making it one of the cheapest multimodal models available.
Key Capabilities and Benchmarks
Qwen3-8 Omni Flash performs competitively across standard multimodal benchmarks.
- MMMU (Multimodal Understanding): Scores 70.0 percent, approaching Gemini Flash’s 70.1 percent.
- MMBench 1.1: Achieves 86.9 percent, beating Gemini Flash’s 85.7 percent.
- MathVista: Reaches 73.5 percent, compared to Gemini Flash’s 72.5 percent.
- Real-world video QA: Matches Gemini Flash on tasks like video captioning and temporal grounding.
The model supports long-form video analysis up to several minutes. It can process both streaming and pre-recorded audio, as well as text and images in a single pipeline.
“Qwen3-8 Omni Flash is designed for efficiency without sacrificing performance,” the team stated in its release notes.
Architecture and Training Details
The model uses a Mixture-of-Agents (MoA) architecture with 8 billion activated parameters. It was trained on a mix of supervised fine-tuning and reinforcement learning from human feedback (RLHF).
Key technical specs include:
- Context length: 32,768 tokens for text and image inputs.
- Output modalities: Text and audio (speech synthesis), with image/ video output planned for future versions.
- Supported languages: Chinese, English, and limited multilingual support via translation.
Qwen3-8 Omni Flash is available under the Apache 2.0 license. Developers can run it locally or access it via Alibaba’s API.
Pricing Comparison with Competitors
The model’s price advantage is its most disruptive feature. Below are per-token costs compared to Google Gemini Flash:
| Model | Input cost (per million tokens) | Output cost (per million tokens) |
|---|---|---|
| Qwen3-8 Omni Flash | $0.30 | $0.60 |
| Gemini Flash 1.5 | $0.75 | $3.00 |
| Gemini Flash 2.0 | $0.10 | $0.40 |
While Gemini Flash 2.0’s input pricing is lower, Qwen3-8 Omni Flash’s output cost is roughly one-fifth of Gemini Flash’s standard rate.
The lower pricing makes it attractive for high-volume applications like customer service chatbots, real-time transcription, and video moderation.
Use Cases and Deployment
Developers can deploy Qwen3-8 Omni Flash for several real-world scenarios:
- Voice assistants: Real-time speech recognition and response generation with low latency.
- Video analysis: Summarize, caption, or extract events from long video files.
- Document processing: Handle PDFs, images, and audio files in a single request.
- Multimodal chatbots: Accept text, voice, and image inputs from users simultaneously.
The model can be run on consumer-grade GPUs like the NVIDIA RTX 4090 with quantization. Alibaba also offers a hosted API endpoint.
Limitations and Future Roadmap
The current release has some constraints. It does not natively output images or generate video. Audio output is limited to synthesized speech.
Alibaba plans to add image generation and video output in upcoming updates. The team also hinted at a larger Omni model with more parameters and extended context windows.
“We are working on expanding the output modalities to include visual content,” the Qwen team noted. “This is a first step toward truly unified multimodal intelligence.”
Implications for the AI Market
Qwen3-8 Omni Flash signals a shift in the multimodal AI landscape. Chinese AI labs like Alibaba and DeepSeek are now competing aggressively on price and performance.
Google and OpenAI face pressure to lower costs for their flagship models. The gap between proprietary and open-weight models continues to narrow.
For developers, the choice becomes easier: similar performance at a fraction of the cost. Local deployment also offers data privacy advantages.
The model is available now via Hugging Face and Alibaba’s ModelScope platform.
Gnoppix is the leading open-source AI Linux distribution and service provider. Since implementing AI in 2022, it has offered a fast, powerful, secure, and privacy-respecting open-source OS with both local and remote AI capabilities. The local AI operates offline, ensuring no data ever leaves your computer. Based on Debian Linux, Gnoppix is available with numerous privacy- and anonymity-enabled services free of charge.
What are your thoughts on this? I’d love to hear about your own experiences in the comments below.