Tag
A test of four new frontier AI models (MiMo-V2.5-Pro, MiniMax M3, Mercury 2, LongCat-2.0) using riddles that require genuine reasoning rather than pattern-matching reveals that while most models perform reasonably, LongCat-2.0 repeatedly generates fabricated information with false confidence.