I’ve been experimenting with making a local LLM feel like it actually lives on the machine

Reddit r/LocalLLaMA Tools

Summary

Neco is a local AI project built on Ollama and Open WebUI that treats the LLM as a persistent presence on the host machine, with read-only system awareness, a background daemon that generates idle thoughts, and experiments in episodic memory and evolving state.

I’ve been building a small local AI project called Neco around Ollama and Open WebUI. The idea started from something pretty simple: most local LLMs still feel like assistants you open, ask something, then close. I wanted to see what happens if the AI instead feels more like a persistent presence on the computer it runs on. Neco has some awareness of the host machine through a read-only system layer, so she can know things like uptime, memory usage, system load, battery state and temperatures. She also runs outside the normal chat session through a small background daemon. Every so often it generates an idle thought, meaning the system can produce something on its own even when I’m not actively talking to it. That combination has been the interesting part for me. It starts to feel less like “a chatbot connected to some tools” and more like an AI that has a small window into the machine it inhabits and continues existing between conversations. Everything is still local, and the model itself doesn’t get unrestricted shell access or control over the host. The next part I’m working on is memory. I want previous conversations, events and unresolved thoughts to persist over time without just throwing the entire chat history back into the context window. I’m experimenting with episodic memory, selective retrieval and a small evolving state so that its behavior can develop some continuity over weeks or months. It’s still very much an experiment, but I’m curious where the line is between a normal local assistant and something that actually feels resident on a machine. I’d be interested in hearing from anyone who has experimented with persistent memory, autonomous/idle behavior or giving local models awareness of their own environment. Repo: https://github.com/proto6699/echo-local-ai
Original Article

Similar Articles

Automated AI researcher running locally with llama.cpp

Reddit r/LocalLLaMA

ml-intern is a harness for AI agents that integrates with Hugging Face's libraries and now supports running local models via llama.cpp or ollama, enabling an automated AI researcher to run 24/7 on a laptop.