PrismML releases Ternary Bonsai 2 27B: a 5.9 GB ternary-weight Qwen3.8 27B that keeps 98% of its performance
On Sept 17, 2026 PrismML released Ternary Bonsai 2 27B under Apache 2.0. It converts Qwen3.8 27B to ternary weights (−1, 0, +1 with FP16 group scaling, 1.76 effective bits per weight), shrinking it from about 54 GB to 5.9 GB while keeping 98.2% of the original's aggregate benchmark score (83.9 vs 85.4). A 27B reasoning model with image input and 262K context can then run on a laptop or one consumer GPU.
Key facts
- Base: Qwen3.8 27B; ternary weights, 1.76 effective bits/weight; 5.9 GB (~9x smaller)
- Aggregate score 83.9 vs 85.4 for full precision (98.2% retained); coding and math said to be level with the original
- 262K context, text + image input, reasoning on by default with configurable effort, tool calling
- Apache 2.0; weights on Hugging Face (prism-ml/bonsai-2 collection); also on OpenRouter
What happened
PrismML, a startup focused on extreme compression, released its largest ternary model, built on Alibaba's open Qwen3.8 27B.
Why it matters
Near-lossless ternary compression at 27B scale makes capable reasoning models practical on local, low-power hardware. That matters for on-device AI and for how freely open-weight capability spreads.
Changelog
- 2026-10-01: created (leads run, missed pre-window item)
Related events
Sources (3)
- officialPrismML: Introducing Bonsai 2 27B
- codeHugging Face: prism-ml Bonsai 2 collection
- pressMarkTechPost: PrismML releases Ternary Bonsai 2 27B
id: 2026-09-17-prismml-ternary-bonsai-2-27b · updated 2026-10-01 · open in the interactive timeline