Forlinx 20-TOPS M.2 AI accelerator supports PCIe cascading for local LLM inference
Summary
Forlinx Embedded has launched an M.2 AI accelerator card based on Rockchip's RK1820 and RK1828 processors, providing 20 TOPS of INT8 performance for local LLM inference on embedded Linux and Android systems.
Similar Articles
@LinusEkenstam: Quite a big deal. 1.35x faster inference than MLX-LM 1.23x faster at prefill having the hybrid option to pick from insi…
Perplexity has open-sourced Lily, a local inference engine optimized for Qwen3.6-35B-A3B on Apple Silicon, achieving 1.35x faster inference than MLX-LM.
I tested the CMP170HX
A hands-on benchmark of Nvidia CMP170HX mining cards repurposed as 64GB VRAM AI inference accelerators, showing they can run large local LLMs like DeepSeek V4-Flash and gpt-oss-120B at useful speeds, with caveats around Ampere-class throughput and PCIe Gen2 x4 connectivity.
Tiny company steals AMD's thunder and challenges Nvidia with old-tech PCIe AI accelerator that runs 700B LLMs locally, sipping just 240W thanks to decade-old DDR4 and 28nm chips
Taiwanese startup Skymizer unveiled the HTX301, a PCIe AI accelerator that uses older 28nm chips and DDR memory to run 700B parameter LLMs locally at just 240W, challenging high-power GPU solutions from Nvidia and AMD.
@liquidai: Introducing LFM2.5-230M: our smallest model yet, built to run fast anywhere (CPUs, NPUs, and GPUs) to enable agentic ta…
Liquid AI releases LFM2.5-230M, a small 230M parameter model optimized for fast inference on CPUs, NPUs, and GPUs, targeting agentic tasks on devices like phones and robots.
OpenAI and Broadcom unveil LLM-optimized inference chip
OpenAI and Broadcom unveiled Jalapeño, a custom LLM-optimized inference chip that promises substantially better performance per watt than current state-of-the-art, designed from the ground up for current and future AI models.