Tag
A paper from Yale University built a large-scale evaluation framework to compare the distribution gap between LLMs and human researchers in generating research ideas. It found that LLM ideas are highly concentrated in bridge and synthesis types, while human ideas are more broadly distributed. This reveals differences in 'research taste' and poses a challenge to the diversity of Research Agents.
This paper investigates execution-grounded automated AI research by building an automated executor that implements LLM-generated ideas and runs experiments. It shows that execution-guided evolutionary search can find methods that significantly outperform baselines in both pre-training and post-training tasks.