Tag
UltraViT is a latency-optimized vision encoder for large vision-language models, designed for on-device deployment with a pyramidal architecture and a two-stage generative pre-training strategy, achieving state-of-the-art performance at 1.7x speed.
This paper introduces Differentiable Logic Gate Networks (Diff-Logic) as a hardware-native alternative to conventional neural networks for real-time EEG classification on edge devices, achieving competitive performance with significantly lower latency and model size.
Sundar Pichai reminded that Google was built on open source and applies the same philosophy to AI, with Gemma models designed for edge devices, while noting that frontier models require massive capital investment.
Embodied.cpp is a portable C++ inference runtime that enables efficient deployment of vision-language-action and world-action models across heterogeneous edge devices and robots through modular execution layers and optimized inference.
The author explores what features would make local AI agents genuinely useful for developers, including working with files/repos, safe terminal use, hardware/robotics support, and offline capability.
SupraLabs announces its founding with a focus on training and releasing open-source small language models (SLMs) for edge devices, already publishing models like Supra-Mini-v4-2M on Hugging Face.
This paper presents Flavors of Moonshine, a suite of tiny specialized ASR models for edge devices. The authors show that monolingual models trained on a balanced mix of human-labeled, pseudo-labeled, and synthetic data outperform larger multilingual models like Whisper, achieving state-of-the-art error rates for small models and enabling on-device ASR for underrepresented languages.