Forlinx 20-TOPS M.2 AI accelerator supports PCIe cascading for local LLM inference

Reddit r/LocalLLaMA Products

Summary

Forlinx Embedded has launched an M.2 AI accelerator card based on Rockchip's RK1820 and RK1828 processors, providing 20 TOPS of INT8 performance for local LLM inference on embedded Linux and Android systems.

Forlinx Embedded has listed an M.2 AI accelerator card based on Rockchip’s RK1820 and RK1828 processors, providing 20 TOPS of INT8 computing performance and up to 5GB of integrated DRAM. The module uses an M.2 2280 interface and is designed to handle local AI inference, including large language models, vision-language models, and computer vision workloads on embedded Linux and Android systems. https://linuxgizmos.com/forlinx-20-tops-m-2-ai-accelerator-supports-pcie-cascading-for-local-llm-inference/
Original Article

Similar Articles

I tested the CMP170HX

Reddit r/LocalLLaMA

A hands-on benchmark of Nvidia CMP170HX mining cards repurposed as 64GB VRAM AI inference accelerators, showing they can run large local LLMs like DeepSeek V4-Flash and gpt-oss-120B at useful speeds, with caveats around Ampere-class throughput and PCIe Gen2 x4 connectivity.

OpenAI and Broadcom unveil LLM-optimized inference chip

OpenAI Blog

OpenAI and Broadcom unveiled Jalapeño, a custom LLM-optimized inference chip that promises substantially better performance per watt than current state-of-the-art, designed from the ground up for current and future AI models.