research-ideas

Tag

Cards List
#research-ideas

Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty

arXiv cs.CL · yesterday Cached

This paper introduces Think-Probe-Respond, a method to improve large language models' judgment of research idea novelty by probing latent judgments and reducing bias towards medium novelty ratings, achieving a 22.30% performance improvement.

0 favorites 0 likes
#research-ideas

AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation

arXiv cs.CL · 2026-07-29 Cached

This paper evaluates how well LLMs (ChatGPT, Claude, DeepSeek) can generate one-page project plans in physics, astrophysics, and cosmology, and how human and AI reviewers assess them. Results show that human reviewers rate AI and human proposals similarly, while AI reviewers prefer AI-written proposals and can perfectly distinguish them from human-written ones.

0 favorites 0 likes
#research-ideas

@rohanpaul_ai: This Yale + University of Chicago paper shows that real gap between LLM generated research ideas vs humans is not idea …

X AI KOLs Timeline · 2026-07-04 Cached

This paper from Yale and the University of Chicago finds that the main difference between LLM-generated and human research ideas is not quality but range: LLMs produce narrower ideas, with 47-64% of their ideas focusing on connecting separate work, compared to only 12.1% for humans.

0 favorites 0 likes
#research-ideas

SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones?

Hugging Face Daily Papers · 2026-05-28 Cached

SoundnessBench is a benchmark of 1,099 machine-learning research proposals that evaluates LLMs' ability to assess methodological validity, finding a pervasive optimism bias in current models.

0 favorites 0 likes
← Back to home

Submit Feedback