Tag
The study investigates whether LLM-as-a-Judge evaluators reliably assess psychological depth in LLM-generated stories, revealing that human preferences are heterogeneous while judges exhibit bias towards reasoning outputs based on surface features.