IT Home On October 9, a post by JetBrains yesterday (October 8) announced the launch of the Mellum2.1 model, focusing on enhancing AI programming capabilities. The model continues the 12B hybrid expert architecture of Mellum2, with 2.5B active parameters, and is still released under the Apache 2.0 license.
Mellum2.1 mainly upgrades the pre-trained reinforcement learning phase, expanding it from a short-term closing step into the main training process, and adds training data for tasks such as mathematics, algorithm competitions, science, tool usage, and software engineering.
JetBrains has built an internal reinforcement learning environment infrastructure for this model. During training, millions of sandboxes are launched to cover thousands of environments. In addition, before training, the team filtered open datasets to remove tasks with test flaws, unverifiable answers, and inappropriate difficulty levels.
After the upgrade, Mellum 2.1 can explore the code library, edit files, and check its own modifications. JetBrains says that the biggest improvement occurs in agent programming, where the model can identify the root cause of failed tests, draft repair solutions, and verify the results.
In terms of performance, the Mellum2.1 architecture is consistent with the Mellum2 architecture, and the speed remains unchanged. Multi-Token Prediction (MTP) further improves response speed. In a single request scenario, MTP increases its speed by approximately 1.6 times.


JetBrains compared Mellum2.1 with Mellum2, Qwen3.5-9B, and Gemma 4 E4B under the same evaluation settings. The results show that under high load, Mellum2.1’s reasoning throughput in Token count is nearly twice that of Qwen3.5-9B, and there are improvements in all areas such as programming, algorithm competitions, mathematics, tool calls, and general knowledge. IT Home has the relevant screenshots below:

Mellum2.1 has now been launched on Hugging Face. It supports deployment on local or proprietary infrastructure, allowing enterprises to keep their code and data within their own environment. Officially, versions for llama.cpp, Ollama, and LM Studio will also be released in GGUF format, along with the MTP inference decoding component of vLLM.
