Tag
The paper evaluates local open-weight LLM judges against human ratings, finding high self-consistency but limited agreement with human judgments, highlighting the need for dual assessment.