1-Bit LLM in the Browser
Summary
A 1-bit LLM (Bonsai) is now runnable in the browser via WebGPU, enabling efficient on-device inference.
View Cached Full Text
Cached at: 07/20/26, 09:47 AM
Bonsai 1-bit WebGPU - a Hugging Face Space by webml-community
Source: https://huggingface.co/spaces/webml-community/bonsai-webgpu
Spaces
— https://huggingface.co/webml-community
webml-community / bonsai-webgpu Running
Refreshing
Similar Articles
Bonsai 27B: 1-bit dense LLM running locally in your browser using custom WebGPU kernels
Bonsai 27B is a 1-bit dense large language model that can run locally in a browser using custom WebGPU kernels, enabling efficient on-device inference.
@xenovacom: Bonsai 27B just changed the local LLM game forever. 1-bit quantization shrinks it from 54GB to just 3.8GB (-93%), while…
Bonsai 27B achieves 93% size reduction via 1-bit quantization while retaining 90% intelligence, enabling local browser inference with custom WebGPU kernels.
Bonsai 27B (1-bit LLM): The First 27B-Class Model to Run on a Phone
PrismML announces Bonsai 27B, a 1-bit and ternary quantized version of Qwen3.6 27B that runs on phones and laptops, retaining 90-95% of baseline performance with a 3.9GB footprint, enabling agentic and multimodal on-device AI.
LFM2.5 230M running in-browser at 1,400 tok/s using custom WebGPU kernels
LFM2.5 230M model achieves 1,400 tokens per second in-browser using custom WebGPU kernels, demonstrating efficient local inference.
WebLLM: high-performance in-browser LLM inference engine
WebLLM is a high-performance in-browser LLM inference engine that leverages WebGPU for hardware acceleration and is fully compatible with the OpenAI API, enabling local execution of open-source language models.