Tag
This paper analyzes integer-sequence benchmarks for language models using minimum description length, revealing that these benchmarks often measure memorisation rather than inductive reasoning, and introduces a new difficulty measure.