JetBrains has released downloadable Mellum 2.1 weights for developers who want to run a coding model on private infrastructure. Its training targets repository exploration, file editing, and checking changes, giving teams an agent-oriented LLM to evaluate without sending their source code to a hosted model.
The Mellum 2.1 announcement describes a 12-billion-parameter mixture-of-experts model that activates 2.5 billion parameters during inference. It retains Mellum 2âs architecture and Apache-2.0 license. The main change is post-training, particularly reinforcement learning in sandboxed environments.
The release is more interesting as a potential worker inside a coding system than as another general-purpose chatbot. JetBrains positions it for coding agents, fast subagents, and self-hosted deployment. Whether it deserves those jobs depends on how reliably it completes repository tasks and how efficiently a team can serve it. The launch supplies JetBrainsâ own evidence for both, with no independent validation.
Repository Work Became the Main Training Target
JetBrains says almost all the work on Mellum 2.1 went into post-training, primarily reinforcement learning. In Mellum 2, RL was a short final stage; for this release, the company says it became the main part of training.
The company reports millions of sandboxed runs across thousands of environments, using infrastructure it built in-house. Those environments trained the model for repository exploration, file editing, and checking its own changes.
A repository agent needs to find relevant files, interpret existing behavior, make a change, and use feedback to decide whether that change worked. Training around those interactions addresses a different problem from producing plausible code from a prompt.
JetBrains also reports adding RL tasks in mathematics, competitive programming, science, tool use, and software engineering. It combined open datasets with tasks developed internally, then filtered sources before training. The company identifies broken tests, unverifiable answers, and tasks that were too easy or impossible for the model as problems in open training data.
An environmentâs feedback is only useful when it measures the intended outcome: a broken test can reward or penalize the wrong behavior. JetBrainsâ description suggests it treated feedback quality as a training concern, although the announcement does not provide enough detail to independently assess the filtering process.
The resulting capabilities remain company claims. âChecks its own changesâ should mean that the model can participate in a verification workflow, not that its output is automatically correct. Passing an available test is useful evidence, but suitable test coverage and review are still needed.
The 2.5B Active Figure Is Not the Modelâs Memory Footprint
Mellum 2.1 has 12 billion total parameters, with 2.5 billion active during inference. In a mixture-of-experts architecture, computation routes through a subset of the modelâs expert components instead of using every parameter for every token.
The design aims to combine a larger pool of model weights with less per-token computation than a similarly sized dense model. For an agent making repeated model calls, latency and serving efficiency can be especially relevant.
The active-parameter figure is not a complete hardware specification, though. Developers still need to account for the modelâs full weights, their numerical precision, the serving implementation, and memory used to process context. A model with 2.5 billion active parameters should not be assumed to have the storage or memory requirements of a 2.5-billion-parameter dense model.
JetBrains does not specify consumer-hardware requirements in the announcement. There is no supported basis here for promising a particular laptop experience, GPU configuration, or minimum memory capacity.
With the architecture unchanged since Mellum 2, JetBrains attributes the new repository capabilities to training rather than increased model size. The release attempts to make the same compact architecture a more capable agent worker.
JetBrains Produced the Benchmark Comparisons
In JetBrainsâ evaluation, Mellum 2.1 was compared with Mellum 2, Qwen3.5-9B, and Gemma 4 E4B. The company says it used the same evaluation setup for all four.
The largest reported improvement over Mellum 2 came in agentic coding. JetBrains also describes gains in coding, competitive programming, mathematics, tool calling, and general knowledge. These vendor-reported findings offer useful signals about the intended scope of the post-training.

A common setup is preferable to comparing unrelated published scores. It does not establish how the models would behave inside a different agent framework with different tools and repository tasks, and the comparisons remain JetBrainsâ work, not an independent labâs.
The speed claims need to be read separately. Under heavy load, JetBrains says Mellum 2.1 served almost twice as many tokens as Qwen3.5-9B. It also reports an approximately 1.6-fold speedup for a single request using multi-token prediction, or MTP.
The first figure concerns throughput under load; it does not mean every individual response arrives twice as quickly. The second describes an MTP speedup and should not be rewritten as a 1.6-fold advantage over Qwen3.5-9B.
Neither figure establishes how quickly a complete coding task finishes. An agent may spend time searching files, running tests, processing long context, or retrying unsuccessful changes. A higher token rate helps only to the extent that model generation is a meaningful part of that workload.
For a deployment decision, the relevant measurement is closer to time and cost per correctly completed task. JetBrainsâ token-serving results justify testing Mellum 2.1, but an end-to-end comparison is still needed.
Hugging Face Weights Are Available, With Runtime Extras Pending
JetBrains has published the weights through its Mellum 2.1 Hugging Face collection. Combined with the Apache-2.0 license, that gives developers a path to running the model on infrastructure they control without depending on a JetBrains-hosted inference service.
Downloadable weights do not amount to a ready-made local deployment. At announcement time, JetBrains said GGUF builds for llama.cpp, Ollama, and LM Studio were coming soon. It also said the MTP head for speculative decoding in vLLM was forthcoming.
The announcement thus reports an acceleration benefit involving MTP while listing the relevant head as a future release. Teams should check the actual artifacts available before assuming their deployment can reproduce that configuration.
Sources
- Mellum 2.1 announcementblog.jetbrains.com
- Mellum 2.1 Hugging Face collectionhuggingface.co





