@yishan: Meta really fumbled this guy.
Summary
John Carmack comments on memory cost and capacity issues for AI accelerators, noting that model inference can have deterministic memory access patterns, contrasting with game rendering.
View Cached Full Text
Cached at: 07/07/26, 05:25 AM
Meta really fumbled this guy.
John Carmack (@ID_AA_Carmack): Memory cost and capacity are significant issues for AI accelerators.
Unlike game rendering, model inference can have a deterministic memory access pattern. You don’t need “random access memory” at all for model weights, and you could tolerate cold-start latencies in the multiple
Similar Articles
@hwchase17: memory has been a hot topic for the past ~2 years every time we bring it up or do anything there, gets ton of interest,…
A tweet discusses challenges in AI memory integration, such as application specificity and limited utility in general-purpose agents, despite ongoing interest.
Memory layer for AI agents is totally FUCKED
A developer tested 4 memory SDKs for AI agents and found they fail to handle changing facts and entity resolution, indicating critical flaws in current memory tools.
@yoheinakajima: https://x.com/yoheinakajima/status/2081741659260477666
This thread explores how the brain's dual memory systems (hippocampus and neocortex) offer lessons for building long-running AI agents, arguing that agents need a fast episodic capture and slow consolidation mechanism to avoid catastrophic interference, rather than relying solely on frozen models with temporary scaffolding.
@mvanhorn: https://x.com/mvanhorn/status/2070966613994795489
The author argues that AI agent memory bloat degrades performance, and recommends keeping memory and CLAUDE.md files under 200 lines, using on-demand retrieval instead of loading everything into context.
Agentic AI memory isn't a hoarding problem. It's a pruning problem.
The author argues that AI agent memory should focus on pruning data rather than hoarding, drawing parallels to human memory types (sensory, short-term, long-term) and suggesting that modeling after human memory can reduce token usage while maintaining high-quality context.