PrismML has announced the launch of Bonsai 2 27B, its newest flagship multimodal AI model based on Qwen3.8 27B that brings advanced reasoning, coding, vision, and agentic capabilities into a significantly smaller, faster, and energy-efficient footprint. Offered under an open-source Apache 2.0 license as a 5.9 GB ternary model, Bonsai 2 27B reduces the memory footprint by over 9x compared to its full-precision counterpart while retaining over 98% of its aggregate benchmark performance. Using v5 Google TPUs and tuned for high-speed inference in consumer-grade devices up to 143 tokens/second on an NVIDIA GeForce RTX 5090 the 27.8-billion parameter model was designed to handle applications like local assistants, personal knowledge systems, and other long-range applications that once needed cloud computing.
Also Read: Microsoft Shifts AI Paradigm with Global Availability of Copilot Cowork
Highlighting this architectural evolution, Babak Hassibi, Founder and CEO of PrismML, stated, “Bonsai 27B proved that powerful models do not have to be confined to cloud infrastructure. With Bonsai 2 27B, we are closing the quality gap while keeping the same deployment advantages: a dramatically smaller footprint, strong local performance, and the ability to support more demanding workflows such as agentic coding, multimodal understanding, and long-horizon task execution.” Emphasizing the broader industry implications of low-bit model retention, Ion Stoica, Advisor to PrismML and Professor at UC Berkeley, added, “AI deployment has historically forced a choice between model quality and the resources required to run it. What is compelling about Bonsai 2 27B is how little capability is lost despite such a dramatic reduction in footprint. If that gap continues to close, it can fundamentally expand where highly capable models can be deployed.”


