Tag
The article discusses how a fixed evaluator in an agent loop can still be targeted by agent adaptation, based on the AQuA preprint, and raises questions about the isolation of validation feedback.
This paper introduces AHD Agent, a framework using agentic reinforcement learning to enable LLMs to autonomously design heuristics for combinatorial optimization problems by dynamically interacting with the solving environment.