Back to list
Product LaunchPrismMLLarge Language ModelsMobile AI

PrismML Unveils Bonsai 27B: The First 27B-Class Multimodal AI Model to Run Locally on Smartphones

PrismML has announced the launch of Bonsai 27B, a groundbreaking multimodal AI model based on Qwen3.6 27B that marks the first time a model of this scale can operate on a mobile device. By utilizing advanced 1-bit and ternary weight quantization, PrismML has reduced the memory footprint of a 27B-parameter model from the standard 54GB to as little as 3.9GB. This allows the model to perform complex tasks like multi-step reasoning, vision processing, and agentic loops directly on hardware like the iPhone 17 Pro. The release includes two variants: a 5.9GB ternary version for laptops and a 3.9GB 1-bit version for mobile, both supporting a 262K-token context and multimodal capabilities.

Hacker News

Key Takeaways

  • Mobile Milestone: Bonsai 27B is the first 27B-class model capable of running on a smartphone, specifically optimized for devices like the iPhone 17 Pro.
  • Extreme Quantization: The model utilizes 1-bit and ternary weight representations to compress a 54GB model down to 3.9GB–5.9GB without relying on high-precision "escape hatches."
  • Advanced Capabilities: Despite its small footprint, it supports multi-step reasoning, structured tool calls, and computer-use agentic loops.
  • Multimodal Integration: Both variants include a compact 4-bit vision tower, enabling the processing of screenshots, documents, and camera input.
  • Large Context Support: The model features a full 262K-token context window and supports speculative decoding for improved performance.

In-Depth Analysis

Breaking the Memory Barrier with Low-Bit Weights

Historically, deploying a 27B-parameter model locally has been an impractical feat for consumer-grade mobile hardware. In standard 16-bit precision, such a model requires approximately 54GB of VRAM, and even traditional 4-bit quantization typically results in an 18GB footprint—still far exceeding the memory capacity of modern smartphones and many laptops. PrismML’s Bonsai 27B addresses this bottleneck through the implementation of ultra-low-bit representations.

The architecture is built on two primary quantization strategies. The Ternary Bonsai 27B variant uses weights restricted to {-1, 0, +1} with FP16 group-wise scaling. This results in an effective 1.71 bits per weight, bringing the total size to 5.9GB. This version is positioned as the quality-oriented variant, designed for everyday laptops while maintaining full reasoning and tool-calling capabilities.

The 1-bit Bonsai 27B variant pushes the boundaries further by using binary {-1, +1} weights. This achieves an effective 1.125 bits per weight, resulting in a 3.9GB footprint. This reduction is critical as it fits within the strict memory budgets of mobile devices, specifically enabling a 27B-class model to run on an iPhone 17 Pro for the first time. Notably, this low-bit representation is applied end-to-end across the entire network, including embeddings, attention mechanisms, MLPs, and the LM head.

Multimodal Functionality and Agentic Loops

Bonsai 27B is not merely a language model; it is a multimodal flagship designed for complex, real-world workflows. By incorporating a vision tower in a compact 4-bit form, the model can interpret visual data such as screenshots, physical documents, and live camera feeds. This integration allows for on-device workflows that go beyond text-based interaction.

Furthermore, the model is engineered for high-level cognitive tasks. PrismML highlights its ability to handle multi-step reasoning and structured tool calls, which are essential for creating autonomous agents. The model supports "computer-use agentic loops," which remain coherent across numerous steps. This suggests a level of stability and logic retention that was previously reserved for much larger, cloud-based models. With a 262K-token context window, Bonsai 27B can process and remember vast amounts of information within a single session, further supported by speculative decoding to enhance the speed of generation.

Industry Impact

The introduction of Bonsai 27B represents a significant shift in the AI industry's approach to local deployment. By proving that 1-bit and ternary weights can produce commercially viable and highly capable models, PrismML is challenging the assumption that high-tier reasoning requires massive hardware clusters.

This development has profound implications for privacy and accessibility. Bringing 27B-class intelligence to a phone allows users to perform complex data analysis, vision tasks, and agentic automation without sending sensitive data to the cloud. Moreover, it sets a new benchmark for mobile hardware utilization, potentially accelerating the demand for specialized AI chips in consumer electronics. As the first of its kind to bridge the gap between massive parameter counts and mobile memory constraints, Bonsai 27B paves the way for a new generation of truly portable, high-intelligence applications.

Frequently Asked Questions

Question: What are the specific hardware requirements for Bonsai 27B?

Bonsai 27B comes in two versions. The 1-bit variant (3.9 GB) is designed to fit the memory budget of an iPhone 17 Pro, making it the first 27B-class model to run on a phone. The ternary variant (5.9 GB) is intended for use on everyday laptops where higher quality reasoning and tool-calling are required.

Question: Does the model lose functionality due to its small size?

According to PrismML, Bonsai 27B maintains high-tier capabilities including multi-step reasoning, structured tool calls, and vision tasks. It operates end-to-end in low-bit representation across all components—embeddings, attention, and MLPs—without needing high-precision "escape hatches" to maintain its performance.

Question: Can Bonsai 27B process images and documents?

Yes, both variants are multimodal. They include a vision tower in a 4-bit form that allows the model to see and process screenshots, documents, and camera input, enabling comprehensive on-device visual workflows.

Related News

LangChain Introduces LangSmith Tuned Evaluators to Streamline AI Agent Error Detection and Production Trace Analysis
Product Launch

LangChain Introduces LangSmith Tuned Evaluators to Streamline AI Agent Error Detection and Production Trace Analysis

LangChain has officially unveiled LangSmith Tuned Evaluators, a sophisticated toolset aimed at enhancing the observability and reliability of AI agents. By integrating quality feedback directly into production traces—beginning with the "Perceived Error" metric—LangSmith provides developers with the necessary context to identify, analyze, and resolve agent-driven errors. This update represents a significant step forward in the LLMops space, offering a structured approach to feedback that bridges the gap between execution and evaluation. The primary goal of this release is to empower development teams to find and fix agent mistakes more efficiently, ensuring that production-level AI applications maintain high standards of accuracy and performance through continuous feedback loops.

fx: A Tiny Open-Source Native Coding Agent Built with Zig for High-Performance AI Workflows
Product Launch

fx: A Tiny Open-Source Native Coding Agent Built with Zig for High-Performance AI Workflows

fx is a newly released, experimental open-source coding agent harness and CLI (v0.0.3) designed for minimalism and extreme performance. Written in Zig, the tool features a remarkably small 6.39MB binary and a cold start time of just 10 microseconds. It is optimized for research, embeddability, and resource-constrained environments like agent sandboxes. Supporting WebAssembly (Wasm) and model-agnostic inference, fx offers a shell-like user interface rather than a heavy TUI. Its design focuses on context efficiency with minimal system prompts to reduce token costs and improve time-to-first-token (TTFT) performance. Currently available under the Apache-2.0 license, fx aims to provide a lightweight alternative for both local and cloud-based AI coding tasks.

Comcast Transforms Millions of Xfinity Routers into Wi-Fi Motion Detectors via Xfinity Shield Update
Product Launch

Comcast Transforms Millions of Xfinity Routers into Wi-Fi Motion Detectors via Xfinity Shield Update

Comcast has officially launched a significant update to its Xfinity Internet app, enabling Wi-Fi motion sensing capabilities across millions of existing customer routers. This new feature, integrated into the Xfinity Shield service, allows compatible routers to act as activity monitors by detecting disruptions in Wi-Fi signals caused by movement. Released on August 18, 2026, the update is being rolled out at no additional cost to customers with supported hardware. By repurposing existing networking equipment into home monitoring tools, Comcast is expanding the utility of its Xfinity ecosystem without requiring users to purchase new devices. This move highlights a growing trend in the telecommunications industry to provide value-added security and monitoring services through software-defined updates to hardware already present in the home.