Tag
The paper introduces PredicateLongBench, a benchmark that systematically probes long-context reasoning by testing models on tasks of identifying contiguous subsequences satisfying predicates, revealing that frontier models struggle as difficulty scales along multiple axes.