Tag
ChessLFM, an AI model, has been featured on the home page of Lichess, a popular chess platform, where the model's seed data was collected.
A systematic mapping study analyzing chess research across humans, engines, and language models to identify gaps and future directions in strategic reasoning.
A browser extension called ChessInsights AI has been developed for real-time chessboard detection and analysis using 100% client-side computer vision, ensuring privacy by processing all data locally without server uploads.
This paper introduces LLAMIA-Bench and a method called latent state internalization to enhance collaboration between language models and non-language agents, showing that internalizing continuous representations outperforms verbalization and matches frontier models like GPT-5.1.
The paper introduces LLAMIA-Bench, a benchmark for collaborative chess tasks between language models and non-language agents, and proposes latent state internalization to outperform text-based verbalization, with a 14B model matching or exceeding frontier models like GPT-5.1.
Chess.com has launched Gambit, a free online poker site emphasizing learning and social play without real-money gambling, and plans to expand into more classic games.
The article compares blindfold chess skills to effective AI-assisted programming, arguing that pattern recognition and attention control are key for both.
This paper leverages offline reinforcement learning to automatically generate high-quality chess puzzles, enhancing pedagogical value for beginners based on extensive user data from platforms like Chess.com and Lichess.
A demo of chessformer_lens shows that ablating a single attention head in a chess transformer causes it to stop recognizing Morphy's queen sacrifice, demonstrating the concentration of specific capabilities in individual heads.
Otter is a 15.3M-parameter human chess AI that extends Maia 2 by conditioning move predictions on game history and time pressure, achieving higher accuracy than Maia 2 with fewer parameters. Trained on 6.1 billion positions from Lichess games, it demonstrates that treating chess as a time-aware, sequential activity improves prediction of human play.
This paper introduces ACT-Eval, a tool-augmented evaluation framework for LLM chess commentary, and releases a benchmark of 325 position-move pairs. It finds that factual hallucinations remain pervasive in LLM chess commentary, and tool augmentation improves factual correctness but not expert-level strategic coverage.
A reflection on how Deep Blue's 1997 victory over Kasparov did not harm chess, but instead led to increased interest and more professional players.
This video builds a chess simulator and tests the reasoning abilities of multiple LLMs using puzzles and tournaments. Results show Gemini 3.1 wins, and there is a clear gap between open-source models and top closed-source models.
Researchers investigate how reinforcement learning (RL) interacts with pretraining quality, showing that stronger pretrained models benefit more from RL and that RL cannot fully compensate for weak pretraining. A joint scaling law is identified for pretraining and post-training compute allocation.
This paper investigates the relationship between pretraining and reinforcement learning (RL) post-training for large language models using chess as a controlled testbed. It establishes a scaling law connecting pretraining loss to post-RL performance and shows that RL amplifies correct moves on easy puzzles while surfacing new correct moves on hard ones.
This paper studies how pretraining choices (model size, data) affect returns from RL post-training on reasoning tasks, using chess as a controlled testbed. It finds that post-RL performance is well-predicted by pretraining loss and that RL amplifies correct moves on easy puzzles while surfacing correct moves on hard puzzles, with findings transferring to math domains.
Built 'Kibitz', a human move predictor for chess broadcasts, trained on RTX 5080, and automated its operation as a business using Hermes, Stripe, and NVIDIA AI Nemotron for a hackathon.
Proposes DD-Elo, a chess rating system that integrates move-level data via a drift-diffusion model to accelerate skill assessment and rating adaptation, while maintaining alignment with traditional Elo.
An informal experiment using a chessboard reveals that vision language models often fail at spatial reasoning and precise structured output, despite correctly recognizing pieces, highlighting a key gap in VLM evaluation.
Boxwood Chess is a chess pattern training tool without timers, streaks, or ratings.