subjective-assessment

Tag

Cards List
#subjective-assessment

Does Reasoning Improve Psychological Depth in Large Language Models? It Depends on Who's Judging

arXiv cs.LG ↗ · 2026-09-15 Cached

The study investigates whether LLM-as-a-Judge evaluators reliably assess psychological depth in LLM-generated stories, revealing that human preferences are heterogeneous while judges exhibit bias towards reasoning outputs based on surface features.

0 favorites 0 likes
← Back to home

Submit Feedback