Tag
IdeaTrail is a dataset of multi-turn process trajectories for scientific ideation, synthesizing research processes from evidence gathering to proposal construction using a Generator–Advisor loop to ensure grounding.
Proposes Agentic-Ideation, a framework for efficient synthesis of agentic trajectories to train LLMs for scientific ideation, achieving over 10x improvement in sample efficiency and outperforming existing workflow-based baselines.
This paper introduces RQ-Bench, a benchmark to evaluate LLMs' ability to assess the novelty of scientific research questions. It finds that LLM judges consistently rate generated questions as more novel than human experts do, raising concerns about the reliability of using LLMs for scientific novelty evaluation.
This paper systematically evaluates human creativity tests for LLMs and finds they fail to predict scientific ideation. It introduces the DRAT, a new test that combines convergent and divergent thinking to reliably predict scientific ideation ability in language models.