Tag
This paper investigates failures in a 2B model for dialogue games and introduces a diagnosis-guided post-training recipe using SFT, DPO, and LoRA to boost performance while maintaining general capabilities.
DelveRL is an open-source, turn-based roguelike game designed as a benchmark for training game-playing agents, including a Godot-based interface and a recurrent PPO trainer for research and development.
Introduces an automated prompt optimization framework for LLM game agents that decomposes the observation-to-action pipeline into two agents and iteratively refines prompts via an evolutionary loop guided by environment returns. Evaluated on BabyAI tasks, it significantly improves success rates (e.g., from 0% to 72.5% on PutNext) without updating model weights.
OmniGameArena introduces a unified benchmark for evaluating VLM agents in diverse Unreal Engine 5 game environments, featuring an Improvement Dynamics Curve for tracking skill evolution across reflection rounds.