Tag
Details a method to run a 13 million parameter ASR Conformer model directly on a microcontroller, highlighting advances in edge AI deployment.
NVIDIA released Cosmos 3 Edge, a 4-billion-parameter open world model for edge devices that helps robots and vision AI agents understand surroundings, reason in real time, and generate actions. It achieves best-in-class throughput and accuracy among similar-sized models.
NVIDIA announces the Jetson T3000 and T2000 modules based on the Thor architecture, delivering powerful AI compute for mainstream robotics and edge AI applications, with adoption by leading robotics companies.
NVIDIA released DeepStream 9.1, an update to its video analytics toolkit, featuring 13 agentic skills that allow users to describe pipelines in natural language, new multi-view 3D tracking, automatic camera calibration, and support for Jetson Orin and Thor. The update is open source on GitHub.
PrismML's Bonsai 27B model runs on the Jetson Orin Nano 8GB with 4.31 tokens/s and 27 t/s prompt processing, using 6.2GB RAM and about 25W power. It indicates surprisingly usable edge AI performance.
This paper explores using a compact Small Language Model (Qwen2.5-1.5B) retrained with GRPO and combined with a validator-guided correction loop for autonomous industrial control. The framework achieves high alignment accuracy and low latency, demonstrating practical viability for edge deployment.
This paper proposes an edge-aware online tracking pipeline for thermal infrared UAV swarm tracking, featuring the Adaptive Kinematic Kalman Filter (AKKF) that balances efficiency and robustness under challenging conditions.
EvoLP is a self-evolving latency predictor for neural network models on edge devices, designed to guide model compression while satisfying strict latency constraints. It outperforms prior methods across multiple edge devices and model variants.
An in-depth guide explaining NVIDIA's physical AI infrastructure stack, including Cosmos, Isaac, Jetson, Omniverse, and edge AI, positioning the company as the gravitational center of physical AI with both cloud and edge capabilities.
Google announces LiteRT.js, a high-performance JavaScript runtime for running AI models directly in the browser using WebAssembly and hardware acceleration, as an evolution from TensorFlow.js.
Applies federated learning to object detection for drone fleets, enabling collaborative training without centralizing aerial imagery, achieving performance close to centralized training while preserving privacy and reducing bandwidth.
Small AI models are proving valuable in regions with unreliable networks, enabling life-saving applications like counterfeit drug detection and disease identification in crops without needing constant internet connectivity.
SupraLabs releases Supra-Router-51M, a 51.7M parameter micro-LLM for multi-task infrastructure routing, designed to decide whether to process prompts locally on edge or send them to cloud-hosted models. Fine-tuned on a small dataset, it uses multi-task sequence generation for robust routing.
The PyTorch Foundation supported the ExecuTorch Hackathon in San Francisco, where over 100 participants built real-time on-device AI applications using PyTorch and ExecuTorch on Snapdragon-powered Samsung Galaxy S25 Ultra devices. Winning projects included SafeScreen AI, SixthSense, and Toddle AI, showcasing local execution benefits for responsiveness, privacy, and offline capability.
Ben Burtenshaw announces a community hackathon focused on porting AI models directly to bare silicon for edge applications like vision, speech, and robotics.
NVIDIA discusses three workflows using synthetic data and fine-tuning on its Omniverse and Metropolis platforms to improve vision AI agent accuracy for edge deployment.
Firefly Aerospace is leveraging NVIDIA Jetson edge AI on its Blue Ghost Mission 2 to perform AI inference in lunar orbit, enabling near real-time processing of imaging data and reducing reliance on slow downlinks.
A language model capable of running on a $5 chip with 12 AI applications, fully offline and open source, available on GitHub and Hugging Face.
NASA is testing Red Hat's RamaLama open source tool to run local LLM and VLM inference for a medical AI assistant on deep space missions, enabling autonomous real-time diagnostics without Earth communication.
A developer demonstrates running the Hermes AI model locally on an M5 Stackchan robot, with plans to do the same on a Reachy Mini. This showcases local AI robotics integration.