parallelquant
September 18, 2026 · MarkTechPost

PrismML shrinks a 27B model to 5.9GB with little quality loss

PrismML released Ternary Bonsai 2 27B, a ternary-weight compressed version of Qwen3.8 27B that occupies 5.93GB versus 53.80GB for the FP16 original, while retaining 98.2% of the parent model's average performance across 20 benchmarks. It handles text and images with a 262K-token context, released under Apache 2.0.

Why it matters: Roughly 9x compression with minimal quality loss is significant for running capable models on consumer hardware, extending a broader trend toward extreme quantization that makes frontier-adjacent capability accessible outside data centers.

Related updates