@_akhaliq: LongMINT Evaluating Memory under Multi-Target Interference in Long-Horizon Agent Systems

X AI KOLs Following Papers

Summary

LongMINT is a benchmark for evaluating memory under multi-target interference in long-horizon agent systems.

LongMINT Evaluating Memory under Multi-Target Interference in Long-Horizon Agent Systems https://t.co/LOxiZFHtQd
Original Article
View Cached Full Text

Cached at: 05/22/26, 11:46 AM

LongMINT

Evaluating Memory under Multi-Target Interference in Long-Horizon Agent Systems https://t.co/LOxiZFHtQd

Similar Articles

Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents

Hugging Face Daily Papers

This paper introduces a proactive memory agent that runs alongside an action agent to prevent behavioral state decay in long-horizon tasks, achieving significant improvements on Terminal-Bench2.0 and τ^2-Bench. The authors also train Qwen3.5-27B using SFT and GRPO as an early step toward open-weight memory policies.

MemGym: a Long-Horizon Memory Environment for LLM Agents

arXiv cs.CL

MemGym is a benchmark for evaluating memory formation in LLM agents over long-horizon tasks, unifying existing agent gyms and synthetic pipelines with memory-isolated scores. It spans tool-use dialogue, multi-turn search, coding, and computer use, and includes a lightweight reward model (MemRM) for efficient evaluation.