PrismML has developed a reasoning large language model, Bonsai 2 27B, which is a mere 5.9 GB, small enough to run on your average PC or high-end smartphone. This marks a significant step in making AI more accessible, as it can now fit on devices without the need for cloud computing. The model, which is a compressed version of Qwen3.8 27B, retains 98% of its benchmark scores, showing that high performance can coexist with miniaturisation.
Founded by Caltech researchers and backed by Khosla Ventures, the startup is pushing the boundaries of compression technology. CEO Babak Hassibi claims that the model’s compression technique is unique and that it could potentially apply to even larger models in the future. The next models, expected to have several-hundred billion parameters, may require less compression, making them even more efficient.
PrismML’s approach, called “ternary” weights, simplifies the storage of model weights to +1, −1, or 0, drastically reducing the space needed. The company has already seen significant interest, with over 13.6 million downloads of their models. This development could lead to more intelligent devices, offering privacy and convenience by running on local devices rather than the cloud.
PrismML’s goal is to make AI more accessible and user-friendly. With the potential for these models to run on users’ devices, it could mean that the next generation of smartphones and PCs will be truly intelligent, without the need for constant internet access. As Hassibi suggests, the trend is towards larger models that can be compressed more easily, making them more practical for everyday use.
Professor Ion Stoica, advisor to PrismML, is excited about the potential of this technology. He believes that it will enable advanced models to run on users’ devices, providing intelligence that is both private and free. This could transform the way we use AI in our daily lives, making it an integral part of our devices.







