@rohanpaul_ai: https://x.com/rohanpaul_ai/status/2061959891036885027
Summary
A Stanford Law School study found that law professors rated LLM-generated answers higher than peer answers in a blinded evaluation of short-answer tutoring in contracts courses, with LLMs winning 75.33% of comparisons and being flagged as harmful less often.
View Cached Full Text
Cached at: 06/03/26, 11:56 PM
Law Professors Prefer AI Over Peer Answers | Stanford Law School
Source: https://law.stanford.edu/publications/law-professors-prefer-ai-over-peer-answers/
Abstract
Large language models (LLMs) are increasingly promoted as educational tutors, yet most evaluations focus on domains with a single ground truth. Many disciplines, however, hinge on judgment: reasoning, weighing ambiguity, and reaching defensi- ble conclusions. Law provides a sharp test. We conducted a blinded evaluation of short-answer tutoring in contracts courses with sixteen U.S. law professors. Partici- pants created 40 representative questions, wrote answers, and judged 2,918 anonymized comparisons between human and LLM responses. Professors rated LLMs far higher than their peers (average win rate = 75.33%), with models performing similarly to the best instructor. LLM responses were also rarely flagged as harmful (3.53%, vs 12.06% for professors). Preferences for LLM answers were consistent across evaluators and reflected shared professional standards. Our evaluation can be reliably extended to ad- ditional models by employing a separate LLM as a judge, rendering expert agreements an effective, scalable method to evaluate AI tutors in judgment-rich domains.
Similar Articles
AI Outperforms Law Professors in Stanford Law Study
A Stanford Law study found that law professors prefer AI-generated answers over those written by fellow instructors in blind evaluations, with AI winning 75% of head-to-head matchups, indicating potential for AI to transform legal education.
@rohanpaul_ai: LLMs often cannot tell when an attack made them say something unsafe. Asking an LLM whether its own previous answer was…
This paper investigates whether LLMs can reliably self-report when their outputs have been compromised by adversarial prefills, finding that models often cannot distinguish between compromised and intentional outputs, and their limited recognition stems from normal refusal behavior rather than true self-awareness.
LLM Judges Can Be Too Generous When There Is No Reference Answer
This paper shows that LLM judges tend to over-credit incorrect answers when no reference answer is provided, and adding a reference can flip verdicts by up to 85%, aligning more with human judgments. The authors propose calibration steps for using LLM judges in reference-free settings.
Distinguishing Artificial from Authentic: Evaluating LLMs for Detecting LLM-Generated Content
This paper investigates how well LLMs can detect their own generated content in educational contexts, finding that detection accuracy varies by task type and is unreliable for short-answer questions.
When LLMs Over-Answer: Measuring and Mitigating Quality Issues in LLM-Based Hardware Description Language Question Answering
This paper investigates quality issues in LLM-generated answers for hardware description language questions, finding over-answering tendencies like redundancy (65.7%) and verbosity (69.1%), and proposes a multi-agent framework that reduces core answers by 37% and non-core content length by 31% while improving quality scores.