balrog-benchmark

Tag

Cards List
#balrog-benchmark

Environment-Grounded Automated Prompt Optimization for LLM Game Agents

arXiv cs.CL ↗ · 2026-06-17 Cached

Introduces an automated prompt optimization framework for LLM game agents that decomposes the observation-to-action pipeline into two agents and iteratively refines prompts via an evolutionary loop guided by environment returns. Evaluated on BabyAI tasks, it significantly improves success rates (e.g., from 0% to 72.5% on PutNext) without updating model weights.

0 favorites 0 likes
← Back to home

Submit Feedback