Tag
Netlify rebuilt Edge Functions infrastructure on Firecracker MicroVMs (with Unikraft), cutting warm-invocation latency from 25–40ms to ~5–6ms at p50 — roughly 5x faster — while improving security and reliability without changing how developers write functions.
This paper proposes EN-HMTFL, an encoder-sharing hierarchical multi-task federated learning framework for vehicular ad hoc networks that lets vehicles train heterogeneous perception tasks collaboratively via a shared encoder while keeping task-specific decoders local, improving accuracy by up to 24% and reducing communication rounds by up to 28.8%.
Liquid AI and Artificial Analysis release Pipette, an open-source benchmarking suite for on-device AI that measures quality, speed, latency, and memory use across models, quantizations, runtimes, and devices, shipping with 10k+ verified results and clients for macOS, Windows, iOS, and Android.
Forlinx Embedded has launched an M.2 AI accelerator card based on Rockchip's RK1820 and RK1828 processors, providing 20 TOPS of INT8 performance for local LLM inference on embedded Linux and Android systems.
HybridInfer introduces a thermal-aware reinforcement-learning router for multi-tier LLM inference, achieving higher quality and cost efficiency than heuristics on real Android devices.
PTC-Decoder is a training-free, plug-and-play framework that improves small language models on offline, resource-constrained edge devices by enforcing plan adherence through token-level constraints, enhancing step-level reliability in multi-step agent tasks.
The article explains when and how to finetune a specialized decision model classifier that outperforms Opus, is faster than Jev, and runs on user devices.
Cloudflare has updated its Workers service to directly accept TCP connections and run gRPC, enabling backend workloads that previously required dedicated servers to be deployed on its global edge network.
The author runs the Ling Tiny 3.0 AI model on a 2017 laptop without GPU, achieving 10 tokens per second and completing tasks like code generation, showcasing the potential for edge intelligence on existing hardware.
An individual built an AI system for factory floors that operates offline to assist technicians by providing historical solutions for machine error codes.
DualRes is a compact oscillatory state-space model for machine fault diagnosis from vibration data, achieving state-of-the-art performance with limited labels and reduced computational requirements for edge deployment.
The paper introduces ZO-COSMO, an index-free method for decentralized zeroth-order optimization that improves accuracy by avoiding sparse index transmission through one-hop mixing. Experiments demonstrate gains over baseline methods on synthetic agents and Qwen LoRA workers.
The article presents an open-source study on optimizing latency for local vision-language models through benchmarking and techniques like native batching and MLX quantization, achieving significant speedups while maintaining decision accuracy on Apple hardware.
Microsoft Research findings demonstrate that offloading AI inference from robots to edge or cloud systems enhances task success rates, efficiency, and battery life for physical AI applications.
This paper presents an affordable, offline AI-integrated smart cane designed to assist visually impaired users with multimodal mobility assistance, using edge computing on a Raspberry Pi Zero 2W for real-time obstacle detection and feedback.
TierKV proposes a predictive multi-tier KV caching framework to optimize memory usage and throughput for long-context LLMs on mobile devices, achieving significant performance improvements with minimal accuracy degradation.
Radio-frequency convolutional neural networks (RF-CNNs) repurpose existing wireless communication hardware for efficient AI inference on edge devices, demonstrating deep CNN performance with significant energy savings.
A NATO-backed startup, Scaleout Systems, is using small AI models and federated learning to enable drones to autonomously identify and attack targets in battlefield settings, leveraging edge computing for real-time updates.
This paper introduces agentic-eCAL, a metric for evaluating energy costs in multi-agent AI workflows across edge-cloud networks, demonstrating that transmission costs are minimal but inference costs are significant.
The user explores adding AI object detection to older security cameras without straining network bandwidth, considering edge processing with on-premises appliances to analyze video locally and only send alerts or short clips, while seeking advice on hybrid architecture and API integration.