Tired of RAM prices, so I have been pooling spare RAM across old devices to run bigger local models

Reddit r/LocalLLaMA Tools

Summary

This article demonstrates how to pool RAM from old devices to run a larger local AI model and build a private chatbot using custom knowledge from PDFs, ensuring all data remains local.

No content available
Original Article
View Cached Full Text

Cached at: 09/25/26, 03:17 AM

**TL;DR:** By pooling the idle RAM from older devices, run larger local models to build a private chatbot that uses confidential data. ## Background & Objectives In the video, the presenter introduces an excellent use case for RAM deck—how to build a private chatbot. This chatbot is trained using confidential data owned by the user, with all data retained on local hardware and never uploaded to the cloud. Applicable scenarios include researchers processing academic papers, scholars managing large databases, or business owners who need to input proprietary business information into a local model. ## Distributed Model Execution The current demo uses a 30B model. Since the mini PC’s memory is insufficient to host such a large model independently, the model is distributed across the mini PC and a Mac Mini. The specific allocation is as follows: - The mini PC is assigned 12GB of RAM. - The Mac Mini is assigned approximately 5.5GB of RAM. This pooling delivers decent performance: generating around 30 to 30+ tokens per second with low latency of about 1 second—quite impressive. ## Training the Chatbot The presenter wants to query the model about a recently approved drug for a rare disease, such as one named “Aurora.” Initially, the model has no knowledge of this information. Next, the presenter downloads relevant research papers from the internet—for example, a PDF article on “corometus” treating amyloid cardiomyopathy. The goal is to feed these PDFs to the model so it can learn about the disease or drug information. The specific process is: 1. The PDF is submitted to the RAM deck coordinator. 2. The coordinator shards the PDF and hosts the shards in the RAM of compute devices across the cluster (not on SSDs or storage devices), ensuring the chatbot can access knowledge quickly and efficiently. 3. The mini PC is designated as the storage node, but the PDF shards are distributed across all devices’ RAM. Once sharding is complete, the knowledge base context is enabled in the chat interface, allowing the chatbot to access the PDF content. After confirming that metrics like latency return to normal, the same question is asked again. ## Performance & Results The model now recognizes that “Aurora” is used to treat conditions like amyloid cardiomyopathy and provides the correct answer. This demonstrates that by pooling RAM, users can build a private chatbot trained with custom knowledge while keeping all data localized. In summary, this approach allows users to cluster the RAM of all compute devices to host a massive knowledge base, enabling efficient and private local AI applications. Source: https://youtu.be/H29ASZhpTKQ

Similar Articles

Running local models on an M4 with 24GB memory

Hacker News Top

A guide on running local AI models like Qwen 3.5-9B on an M4 MacBook with 24GB RAM using tools like LM Studio, Ollama, and pi, including specific configuration tips for optimal performance.

Are the rich RAM /poor GPU people wrong here?

Reddit r/LocalLLaMA

Discusses the trade-off between dense and Mixture-of-Experts (MoE) models for local AI, noting that high-RAM users have limited MoE options beyond Qwen 3.5 122B, and questioning if large GPU is the only viable path.