Tag
This paper introduces Think-Probe-Respond, a method to improve large language models' judgment of research idea novelty by probing latent judgments and reducing bias towards medium novelty ratings, achieving a 22.30% performance improvement.
This paper evaluates how well LLMs (ChatGPT, Claude, DeepSeek) can generate one-page project plans in physics, astrophysics, and cosmology, and how human and AI reviewers assess them. Results show that human reviewers rate AI and human proposals similarly, while AI reviewers prefer AI-written proposals and can perfectly distinguish them from human-written ones.
This paper from Yale and the University of Chicago finds that the main difference between LLM-generated and human research ideas is not quality but range: LLMs produce narrower ideas, with 47-64% of their ideas focusing on connecting separate work, compared to only 12.1% for humans.
SoundnessBench is a benchmark of 1,099 machine-learning research proposals that evaluates LLMs' ability to assess methodological validity, finding a pervasive optimism bias in current models.