标签
介绍了Belief Entropy和Metacognitive Memory Policy Optimization (MMPO),以提高长周期LLM代理的记忆质量,优于现有方法,并在长上下文中保持性能。