We need a 80-160B model urgently. The unified memory device market needs more Models.
Summary
The author argues that there is an urgent need for AI models in the 80-160B parameter range to support users with unified memory devices (e.g., high-RAM Apple/AMD systems), as recent models are either too small or too large for their hardware.
Similar Articles
Why there is a lack of new 100B-120B models?
Analysis of the trend in AI model sizes, noting a gap in the 100-120B parameter range with recent releases focusing on smaller (25-35B) or larger (200B+) models.
AMD & Intel, now onwards it's your turn to release your own models
NVIDIA has released a 550B parameter model (Nemotron-3-Ultra-550B), prompting commentary that AMD and Intel should similarly release their own AI models as model development becomes a commodity for hardware companies.
@awnihannun: It's very cool that Apple shipped a 20B parameter on-device. You can't put 20B parameters in RAM at any reasonable prec…
Apple shipped a 20B parameter on-device model using a MoE variant that selects experts once per query to fit in NAND, enabling inference despite RAM constraints.
Unified Memory, Explained: Why Mini PCs Can Run 70B Models a Big GPU Can't
Explains how unified memory in mini PCs allows them to run large 70B parameter AI models that exceed the VRAM capacity of high-end GPUs, though at slower speeds due to lower memory bandwidth.
@berryxia: Apple has been betting on on-device models all along! Unified architecture memory is the natural habitat for on-device models! Unified memory means memory is VRAM. We are seeing more and more excellent on-device models emerge. OpenBMB released MiniCPM-V 4.6, a 1.3B multimodal model. After reading it…
OpenBMB released MiniCPM-V 4.6, a 1.3B parameter multimodal model. Using high-resolution visual processing and efficient compression, it achieves fast inference on consumer hardware and mobile phones, outperforming larger models. It is fully open-source and supports multiple inference and quantization frameworks.