IT之家 AI

JetBrains Programming AI Model Mellum2.1 Release: Near double inference throughput at high load compared to Qwen3.5-9B

IT Home reported on October 9 that JetBrains posted a blog post yesterday (October 8) announcing the launch of the Mellum2.1 model, with a focus on enhancing AI programming capabilities. The model continues the 12B…

Mellum2.1 compared with Mellum2, Qwen3.5-9B, and Gemma 4 E4B
Image source · IT之家 AI
Thank you to the IT Home netizen who submitted the clue! !

IT Home On October 9, a post by JetBrains yesterday (October 8) announced the launch of the Mellum2.1 model, focusing on enhancing AI programming capabilities. The model continues the 12B hybrid expert architecture of Mellum2, with 2.5B active parameters, and is still released under the Apache 2.0 license.

Mellum2.1 mainly upgrades the pre-trained reinforcement learning phase, expanding it from a short-term closing step into the main training process, and adds training data for tasks such as mathematics, algorithm competitions, science, tool usage, and software engineering.

JetBrains has built an internal reinforcement learning environment infrastructure for this model. During training, millions of sandboxes are launched to cover thousands of environments. In addition, before training, the team filtered open datasets to remove tasks with test flaws, unverifiable answers, and inappropriate difficulty levels.

After the upgrade, Mellum 2.1 can explore the code library, edit files, and check its own modifications. JetBrains says that the biggest improvement occurs in agent programming, where the model can identify the root cause of failed tests, draft repair solutions, and verify the results.

In terms of performance, the Mellum2.1 architecture is consistent with the Mellum2 architecture, and the speed remains unchanged. Multi-Token Prediction (MTP) further improves response speed. In a single request scenario, MTP increases its speed by approximately 1.6 times.

Mellum2.1 compared with Mellum2, Qwen3.5-9B, and Gemma 4 E4B

JetBrains compared Mellum2.1 with Mellum2, Qwen3.5-9B, and Gemma 4 E4B under the same evaluation settings. The results show that under high load, Mellum2.1’s reasoning throughput in Token count is nearly twice that of Qwen3.5-9B, and there are improvements in all areas such as programming, algorithm competitions, mathematics, tool calls, and general knowledge. IT Home has the relevant screenshots below:

Output tokens per second on one H200 for Mellum2.1, Qwen3.5-9B, and Gemma 4 E4B

Mellum2.1 has now been launched on Hugging Face. It supports deployment on local or proprietary infrastructure, allowing enterprises to keep their code and data within their own environment. Officially, versions for llama.cpp, Ollama, and LM Studio will also be released in GGUF format, along with the MTP inference decoding component of vLLM.

Original source

IT之家 AI

Content notes

Original publication and rights belong to the source.

Machine translation · Refer to the original