Tag
This paper investigates the reliability of LLM judges for complex professional tasks using patent drafting as a testbed, finding that judge-guided revision improves quality but exhibits metric-dependent agreement with human expert evaluation.