Tag
The paper introduces the PIJ benchmark for evaluating large language models on criminal profiling tasks from incomplete evidence, highlighting performance gaps and biases in inferential reasoning.