Tag
Researchers at MIT CSAIL and Harvard used a modified Battleship game to study and improve language models' question-asking abilities. By applying Monte Carlo inference strategies, they significantly boosted smaller models like Llama 4 Scout's win rate from 8% to 82% against humans, outperforming larger models at a fraction of the cost.