Tag
The article presents L3Cube-IndicQuest v2, a large-scale multilingual benchmark for evaluating factual knowledge of Large Language Models across Indic languages, with evaluation results for six models.
This paper introduces 'composition collapse', a phenomenon where language models with stable factual knowledge still fail to compose that knowledge into correct multi-hop reasoning, and proposes a double-gate protocol to isolate composition failure from atomic knowledge instability.