Tag
This paper investigates the reliability of LLM judges for complex professional tasks using patent drafting as a testbed, finding that judge-guided revision improves quality but exhibits metric-dependent agreement with human expert evaluation.
Discovery Foundation Models are proposed as general-purpose systems for enabling open-ended scientific discovery through iterative problem formulation, hypothesis testing, and evidence-based revision across dry and wet lab settings. The paper introduces a framework with capabilities like problem discovery and continual improvement, instantiated with systems like Zetema and GALILEO.