THE CRUNCH
Qwen has launched Qwen3.8-Omni-Flash, its first multimodal model designed for AI agents that can process audio and video together. The model boasts a one million token context window and is positioned as a cost-effective alternative to Google's Gemini 3.8 Flash. Qwen claims the model performs on par with Gemini Flash 3.8 in multimodal benchmarks while undercutting its pricing. The model is available through Qwen's AI
Studio, Qwen Cloud, and the API. Qwen's pricing for the model is notably lower than Google's, with input tokens at $0.15 per million and output tokens at $0.47 per million. For context, Gemini 3.8 Flash charges $0.75 for input and $3.75 for output per million tokens, with prices set to double on January 1, 2027. Qwen also offers open-source plugins like Qwen-MM-Plugins and Qwen-Live Harness to enhance agent capabilities.
The affordability extends to specific media types, with Qwen estimating audio input at under $0.01 per hour and 720p video with audio at one frame per second at about $0.20, excluding response costs. This pricing strategy could make advanced multimodal AI more accessible for developers building video editing, translation, or summarisation tools. The model's ability to use tools autonomously, such as editing vlogs or translating short videos, positions it as a practical solution for agent-based workflows.


