THE CRUNCH

Alibaba's Qwen team has released Qwen-Audio-3.1, a suite of five models for speech recognition, text-to-speech, and real-time interaction. The new lineup includes an ASR model that cleans up filler words and a TTS system that can generate sound effects and background audio in a single pass. The company is also aggressively cutting prices, with TTS dropping by about 70 percent, real-time models by roughly 85 percent,

ASR up to 95 percent. The real-time model is designed to listen and speak simultaneously, with a feature that adjusts its response speed and tone based on the detected mood of the speaker. Qwen-Audio-3.1 is available on Qwen Cloud, with pricing details posted on the company's blog.

WHAT HAPPENS NEXT

Developers can now access the new models on Qwen Cloud, with the lower pricing making advanced audio features more affordable.