THE CRUNCH
Alibaba's Qwen team has released Qwen-Audio-3.1, a suite of five models for speech recognition, text-to-speech, and real-time interaction. The new lineup includes an ASR model that cleans up filler words and a TTS system that can generate sound effects and background audio in a single pass. The company is also aggressively cutting prices, with TTS dropping by about 70 percent, real-time models by roughly 85 percent,
ASR up to 95 percent. The real-time model is designed to listen and speak simultaneously, with a feature that adjusts its response speed and tone based on the detected mood of the speaker. Qwen-Audio-3.1 is available on Qwen Cloud, with pricing details posted on the company's blog.


