THE CRUNCH

Deepseek has released V4.1-Flash, a multimodal model with 552 billion parameters that cuts KV cache memory to a quarter of its predecessor. On the DeepSWE coding benchmark, it narrowly beats Opus 5 and GPT-5.6 Sol, even though only 16 billion parameters are active per token. The model ships under the MIT license and targets much cheaper AI agents.

The model is notable for its efficiency, using only 16 billion active parameters per token despite its massive size. This allows it to run with significantly reduced memory requirements, making it a more practical option for AI agents. The model is released under the MIT license, which is a permissive open-source licence.

The model's performance on the DeepSWE coding benchmark is a key differentiator. It narrowly beats Opus 5 and GPT-5.6 Sol, which are considered top-tier models. This suggests that the model is not only efficient but also highly capable in a specific domain.

The model targets much cheaper AI agents, which is a significant market opportunity. By reducing memory needs, the model can be deployed on more affordable hardware, making AI agents more accessible to a wider range of users.

The model is a multimodal model, meaning it can process and generate different types of data, such as text and images. This makes it a versatile tool for a variety of applications beyond just coding.

WHAT HAPPENS NEXT

It remains to be seen how quickly other AI developers will adopt the model and whether its performance on the DeepSWE benchmark translates to other coding tasks.