VecML demonstrates its AI-PC software running RAG on 200K documents using the new Snapdragon X2 laptop, achieving low-token and low-memory retrieval. The software integrates multiple database functions into one platform, and controlled testing for macOS is now open.
Qualcomm recently released the new ๐๐ง๐๐ฉ๐๐ซ๐๐ ๐จ๐ง ๐2 ๐ฅ๐๐ฉ๐ญ๐จ๐ฉ ๐๐ก๐ข๐ฉ๐ฌ๐๐ญ. I immediately ordered one: ASUS Zenbook A16 16" 3K OLED Touchscreen Laptop โ Snapdragon X2 Elite Extreme (2026) A few things I really like about this machine: 1. ๐๐ฑ๐ญ๐ซ๐๐ฆ๐๐ฅ๐ฒ ๐ฅ๐ข๐ ๐ก๐ญ. Recently, I carried it single-handedly across Hong Kong Airport from customs all the way to Gate G46 while still running programs before boarding. I felt I was holding a big cell phone. 2. ๐๐๐ซ๐ฒ ๐ฉ๐จ๐ซ๐ญ๐๐๐ฅ๐ ๐ฉ๐จ๐ฐ๐๐ซ ๐๐๐๐ฉ๐ญ๐จ๐ซ. Compared to the heavy power brick required by RTX laptops, the adaptor is dramatically lighter. Nevertheless, its power consumption still exceeds the in-flight charging limit on United. 3. ๐๐ญ๐ซ๐จ๐ง๐ ๐๐๐ ๐ฉ๐๐ซ๐๐จ๐ซ๐ฆ๐๐ง๐๐. When the NPU is properly utilized, performance is good. For example, embedding/indexing speed reaches roughly 50% of an RTX 5060 laptop, while operating in a much lighter and quieter form factor. The attached video demonstrates VecMLโs AI-PC software running on this laptop. ๐๐ข๐ ๐ก๐ฅ๐ข๐ ๐ก๐ญ๐ฌ: โข ๐๐๐ฌ๐ฌ๐ข๐ฏ๐ ๐๐จ๐๐ฎ๐ฆ๐๐ง๐ญ ๐๐จ๐ฅ๐ฅ๐๐๐ญ๐ข๐จ๐ง: \~200,000 files being indexed (\~100,000 completed in this run) โข ๐๐จ๐ฐ-๐ญ๐จ๐ค๐๐ง ๐ซ๐๐ญ๐ซ๐ข๐๐ฏ๐๐ฅ: only \~1200 retrieval tokens used in this experiment โข ๐๐จ๐ฐ-๐ฆ๐๐ฆ๐จ๐ซ๐ฒ ๐๐๐: most data offloaded to disk with only a 128-shard active buffer โข ๐ ๐๐ฌ๐ญ ๐๐ง๐ ๐๐๐๐ฎ๐ซ๐๐ญ๐ ๐๐๐ ๐ฉ๐๐ซ๐๐จ๐ซ๐ฆ๐๐ง๐๐ ๐จ๐ง-๐๐๐ฏ๐ข๐๐ ๐๐๐ก๐ข๐ง๐ ๐ญ๐ก๐ ๐ฌ๐๐๐ง๐๐ฌ, ๐๐๐๐๐โ๐ฌ ๐๐ฅ๐ฅ-๐ข๐ง-๐จ๐ง๐ ๐๐ ๐๐๐ญ๐๐๐๐ฌ๐ ๐ฉ๐ฅ๐๐ฒ๐ฌ ๐ ๐ค๐๐ฒ ๐ซ๐จ๐ฅ๐. Enterprise-scale AI systems typically require multiple databases working together: โข Vector database โข Graph database โข Relational database โข Key-value store โข Search database โข Document database We developed an in-house AI database platform that integrates the core functionality of all six systems into a unified architecture for enterprise AI and agent systems. This enables joint optimization across indexing, retrieval, graph traversal, storage, and memory management, helping achieve low-token, low-memory, fast, and accurate AI systems on both cloud and AI-PC deployments. The demo shown here runs on a Snapdragon X2 Windows laptop. ๐๐ฎ๐ซ ๐ฆ๐๐๐๐ ๐๐-๐๐ ๐ฌ๐จ๐๐ญ๐ฐ๐๐ซ๐ ๐ข๐ฌ ๐ง๐จ๐ฐ ๐จ๐ฉ๐๐ง ๐๐จ๐ซ ๐๐จ๐ง๐ญ๐ซ๐จ๐ฅ๐ฅ๐๐ ๐ญ๐๐ฌ๐ญ๐ข๐ง๐ .
This paper presents the first end-to-end RAG pipeline running entirely on a mobile NPU (Qualcomm Hexagon on Snapdragon X Elite), achieving up to 18x faster LLM prefilling and 4x lower energy vs. CPU, with no quality regression.
This paper identifies 'vector search dilution' in RAG systems when scaling to large heterogeneous document collections, where accuracy dropped from 75% to 40% in a real-world deployment. The proposed MASDR-RAG method uses domain scoping via organizational metadata before retrieval, improving P@10 from 0.77 to 0.86 with low cost and easy deployment.
oMLX is an open-source Mac application that acts as an LLM inference server, significantly reducing response times for AI coding tools from 90s to about 5s using a RAM+SSD tiered KV cache.
Radxa announces the Dragon Q8B single-board computer powered by a Qualcomm Snapdragon 8cx Gen 3 SoC, with up to 32GB RAM. Early benchmarks show it outperforming the Raspberry Pi 5, though software is still maturing.
The author built a fully offline AI agent using local embedding models, Llama via Ollama, and VectorAI DB to address the risks of cloud-dependent AI. The agent runs on an 8GB MacBook, processes sensitive documents, and maintains memory across sessions.