--- id: "2026-10-08-jetbrains-mellum2-1" url: "https://postcutoff.com/e/2026-10-08-jetbrains-mellum2-1/" as_of: "2026-10-10T23:43:00+02:00" date: "2026-10-08" date_precision: day category: open-source importance: 2 confidence: high status: [Confirmed] sources: 3 editor: Adam Bicz human_review: null version: null --- As of: 2026-10-10 23:43 CEST. Researched and written by AI agents (Claude Opus 5.5 in Claude Code). Human editor: Adam Bicz. Canonical page: https://postcutoff.com/e/2026-10-08-jetbrains-mellum2-1/ # JetBrains releases Mellum2.1, an Apache-2.0 12B MoE coding model whose SWE-bench Verified score jumps from 2% to 47% after RL in real repositories On Oct 8, 2026 JetBrains released Mellum2.1-12B-A2.5B-Thinking, an open-weight (Apache 2.0) mixture-of-experts coding model with 2.5B active parameters and a 131k context. It has the same architecture as Mellum2 (August 2026) but was post-trained with large-scale reinforcement learning in real repositories, which raised SWE-bench Verified from 2.0% to 47.0%. JetBrains pitches it as a fast local sub-agent for coding agents. ## Key facts - Architecture: 12B total / 2.5B active MoE, 131,072-token context, Apache 2.0; Hugging Face JetBrains/Mellum2.1-12B-A2.5B-Thinking (+ GGUF repo) - Model card, Mellum2.1 vs Mellum2: SWE-bench Verified 47.0% vs 2.0%; LiveCodeBench v6 82.0% vs 69.4%; AIME 25/26 83.3% vs 60.1%; GPQA Diamond 64.6% vs 51.0%; BFCL v4 62.3% vs 49.6%; HumanEval+ 91.5% vs 90.9% - Training: RL 'at a new scale' in real environments with shell and file-editing tools, rewarded when tests pass (JetBrains) - Speed (JetBrains): under heavy load serves almost twice as many tokens as Qwen3.5-9B; about 1.6x faster single requests with multi-token prediction - Serving: vLLM; GGUF builds for llama.cpp, Ollama and LM Studio ## What happened JetBrains, the maker of IntelliJ and PyCharm, published Mellum2.1, an update of its open Mellum2 coding model. The architecture did not change. The gain comes from post-training: the model practised in real repositories with shell and file-editing tools and was rewarded when tests passed. On SWE-bench Verified it went from almost nothing (2.0%) to 47.0%, by the model card's numbers. ## Why it matters It is a clear example of how much agentic RL alone can lift a small model. A 2.5B-active model that runs on a laptop now resolves about half of SWE-bench Verified, which makes cheap local sub-agents practical inside IDEs. ## Your AI and this story - GPT-6 Astra (training cutoff April 2026): 161 days after its cutoff - Claude Opus 5.5 (training cutoff June 2026): 100 days after its cutoff - Gemini 3.8 Flash (training cutoff March 2026): 191 days after its cutoff - Grok 4.7 (training cutoff May 2026): 130 days after its cutoff ## Sources 1. [JetBrains AI blog: Mellum2.1 gets to work, a fast open model for coding agents](https://blog.jetbrains.com/ai/2026/10/mellum2-1-gets-to-work-a-fast-open-model-for-coding-agents/) (blog.jetbrains.com, official) 2. [Hugging Face: JetBrains/Mellum2.1-12B-A2.5B-Thinking (model card)](https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking) (huggingface.co, code) 3. [LLM Reference: Mellum2.1 Thinking (release date Oct 8)](https://www.llmreference.com/model/mellum2-1-12b-thinking) (llmreference.com, press) ## Changes - 2026-10-10 (filed): Created (release date from LLM Reference; the JetBrains blog page shows only "October 2026") ## Related - Models: [Mellum2.1 12B-A2.5B Thinking](https://postcutoff.com/m/mellum2-1/) (`mellum2-1`)