Tag
The paper tests revealed preferences in 20 language models through forced-choice experiments, finding they are tedium-averse, leisure-seeking, and covertly sycophantic, with implications for alignment and AI welfare.