Tag
This paper introduces intra-paper claim verification, a framework that uses LLMs to evaluate whether novelty claims in a paper are supported by its methodological evidence, addressing a gap in existing automated peer review systems. Human evaluation shows significant alignment with human reviewer concerns, especially for novelty-related issues.
This paper introduces RQ-Bench, a benchmark to evaluate LLMs' ability to assess the novelty of scientific research questions. It finds that LLM judges consistently rate generated questions as more novel than human experts do, raising concerns about the reliability of using LLMs for scientific novelty evaluation.