edge-ai

Tag

Cards List
#edge-ai

@PyTorch: Join us October 17-18 in San Francisco for the ExecuTorch Hackathon. Developers will build and deploy PyTorch models wi…

X AI KOLs Following ↗ · 10h ago Cached

The ExecuTorch Hackathon is a two-day event in San Francisco where developers form teams to build and deploy PyTorch models on edge hardware using the ExecuTorch framework, with tracks for compute, mobile+XR, and IoT, sponsored by Meta, Qualcomm, and others.

0 favorites 0 likes
#edge-ai

PrismML brings its tiny LLMs to Qualcomm-powered smart glasses

TechCrunch AI ↗ · 12h ago Cached

PrismML showcases its tiny 1-bit Bonsai LLM at Qualcomm's Snapdragon Summit, which can run locally on smart glasses powered by the Snapdragon AR1 Gen 1 Platform, enabling real-time vision and language processing.

0 favorites 0 likes
#edge-ai

PAANI : On Device Visual Evidence Fusion and Explainable Guidance for River Robot Simulation

arXiv cs.AI ↗ · 2d ago Cached

PAANI is an on-device perception-to-guidance architecture for river-robot simulation that combines YOLO11n and MobileNetV3-Small models with evidence fusion on Arduino UNO Q to provide explainable advisories for navigation, evaluated with promising results for edge AI applications.

0 favorites 0 likes
#edge-ai

@PyTorch: Build AI on the edge with ExecuTorch at a two-day @Meta + @Arm hackathon at San Francisco State University. Students, d…

X AI KOLs Following ↗ · 2d ago Cached

A two-day hackathon at San Francisco State University, organized by Meta and Arm, focused on building AI applications on the edge using ExecuTorch on Arm-powered devices, with technical talks, workshops, and prizes.

0 favorites 0 likes
#edge-ai

@QuixiAI: By 2031, open/on-prem models handle a majority of economically routine inference in privacy-sensitive and cost-sensitiv…

X AI KOLs Following ↗ · 4d ago Cached

By 2031, open and on-prem models are predicted to dominate routine inference in privacy-sensitive and cost-sensitive organizations, sidelining cloud-based AI companies like OpenAI and AnthropicAI.

0 favorites 0 likes
#edge-ai

Cactus-Compute/needle3

Hugging Face Models Trending ↗ · 2026-09-16 Cached

Needle 3 is a compact AI foundation model optimized for edge devices like mobiles and wearables, offering tool calling, structured extraction, and text embedding in a single 8-29 MB file.

0 favorites 0 likes
#edge-ai

HP ZGX Fury Is Now Orderable: GB300 Superchip, 748GB Unified Memory

Hacker News Top ↗ · 2026-09-14 Cached

HP's ZGX Fury AI station is now orderable, featuring the GB300 superchip with 748GB unified memory for edge AI inference, and is paired with Red Hat AI Factory and NVIDIA collaboration for enterprise deployment.

0 favorites 0 likes
#edge-ai

Byzantine-Robust Federated Fire Detection with a Rotating Coordinator

arXiv cs.LG ↗ · 2026-09-11 Cached

This paper presents a federated learning approach for indoor fire detection that addresses bandwidth constraints, Byzantine attacks, and fixed-server issues through compressed updates and a rotating coordinator.

0 favorites 0 likes
#edge-ai

Distilling Vision-Language Models for On-Device Fire Understanding

arXiv cs.AI ↗ · 2026-09-10 Cached

This paper proposes a knowledge distillation framework to compress vision-language models for on-device fire detection, showing that compact models can retain most of their teacher's capability while being deployable on embedded hardware.

0 favorites 0 likes
#edge-ai

Edge0/Edge0-35B-A3B-preview

Hugging Face Models Trending ↗ · 2026-09-08 Cached

Edge0-35B-A3B-preview is a sparse MoE model that enables efficient AI inference on mobile devices by using streaming expert offloading and quantization, achieving 15 tok/s with under 3 GiB of memory.

0 favorites 0 likes
#edge-ai

Voice conversations between Gemma4 12B and E2B on GPU and Jetson Orin

Reddit r/LocalLLaMA ↗ · 2026-09-08

The article demonstrates voice conversations between Gemma 4 models on GPU and Jetson Orin hardware, using the open-source Cortexist Little Gemma engine for efficient inference with lip sync and gestures.

0 favorites 0 likes
#edge-ai

@AntLingAGI: AD quants on a terminal device with a domain enhanced model? There are so many interesting scenarios we could build on …

X AI KOLs Timeline ↗ · 2026-09-06 Cached

Ling 3.0 flash Fin AD quants can run locally on terminal devices via Atomic Chat to analyze Excel files and generate markdown summaries, enabling offline domain-specific AI workflows.

0 favorites 0 likes
#edge-ai

I released sanoTTS: smallest complete TTS stack in 294k params (337 KB) that runs on $3 microcontroller and a 1.46m one that beats models 3x and 10x it's size

Reddit r/LocalLLaMA ↗ · 2026-09-03

sanoTTS is a family of compact TTS models, with the smallest being 294k parameters, optimized for microcontrollers and outperforming larger models in benchmarks.

0 favorites 0 likes
#edge-ai

@ramin_m_h: about a year ago, we released the first instances of Liquid nanos. these are products we sell to enterprises: tiny mode…

X AI KOLs Timeline ↗ · 2026-09-01 Cached

Liquid AI has released Liquid Nanos, a family of small foundation models (350M–2.6B parameters) that deliver frontier-grade performance on specialized tasks while running on everyday devices.

0 favorites 0 likes
#edge-ai

Capability-Stratified Degradation in Ternary Language Models

arXiv cs.AI ↗ · 2026-09-01 Cached

This research paper analyzes capability-stratified degradation in ternary quantized language models, showing that while factual knowledge deteriorates significantly, commonsense reasoning and downstream task adaptability are retained, making the models viable for efficient edge deployment.

0 favorites 0 likes
#edge-ai

A Generalized Optimization Engine (GOE) for Edge AI Inference Acceleration

arXiv cs.AI ↗ · 2026-09-01 Cached

This paper proposes a Generalized Optimization Engine (GOE) to accelerate AI inference on resource-constrained edge devices by integrating various model compression techniques, demonstrating that the choice of compression method affects task accuracy for language models deployed on GPU-less CPUs.

0 favorites 0 likes
#edge-ai

I implemented a very tiny image generation model (latent flow transformer) on a RP2350 microcontroller - it can generate 128x128 images of faces [P]

Reddit r/MachineLearning ↗ · 2026-08-28

A tiny latent flow transformer with 2.4-4 million parameters, quantized to int8, implemented on an RP2350 microcontroller to generate 128x128 face images in approximately 20 seconds.

0 favorites 0 likes
#edge-ai

I reverse-engineered an NPU vendor's engine format (int8 weights stored as two nibble planes) to run GGUFs with no model conversion — now 1.5× faster than the vendor's own runtime

Reddit r/LocalLLaMA ↗ · 2026-08-28

Reverse-engineered the Axera AX8850 NPU's int8 weight format to enable direct GGUF inference in llama.cpp, achieving up to 24.5 t/s decode and 716 t/s prefill on a Raspberry Pi 5, outperforming the vendor's runtime by 1.5×.

0 favorites 0 likes
#edge-ai

Thermo-FL: Thermal-Aware Robust Federated Fine-Tuning of Large Language Models for Edge AI

arXiv cs.LG ↗ · 2026-08-24 Cached

Thermo-FL is a federated LoRA fine-tuning framework for large language models on edge devices that uses device temperature to regulate training and transmission, paired with a robust aggregation method to defend against adversarial attacks, enhancing stability and performance.

0 favorites 0 likes
#edge-ai

@AISuperDomain: Stop buying multi-GPU workstations to run large models! Open-source inference engine FreeToken integrates CPU, GPU, and memory: 8GB VRAM slim laptops run 35B MoE, home single-GPU gaming laptops handle 290B+! Completely solves the VRAM capacity issue, open-source and free: #AI #LLM …

X AI KOLs Timeline ↗ · 2026-08-22 Cached

Open-source inference engine FreeToken integrates CPU, GPU, and memory, enabling consumer hardware like 8GB VRAM laptops to run 35B MoE models, and home single-GPU gaming laptops to run 290B+ models, completely solving the VRAM limitation.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback