All articles, most recently crawled first.
Ballmer Peak is a humorous concept from xkcd joking that a specific blood alcohol level enhances programming productivity, named after Steve Ballmer, with no scientific basis but studied in satirical contexts.
This article describes a service that generates RSS feeds for Last.fm user data, including recent tracks, loved tracks, top tracks, artists, and recommended tracks, with customizable periods and format options.
The article provides recommendations for AI labs on responsible release of AI-generated mathematical results, emphasizing the need for human understanding and community-led verification.
NASA is secretly attempting to restart a retired SR-71A Blackbird aircraft after nearly 27 years, hiring former staff and rolling it into a hangar at Edwards AFB as part of efforts to reinvigorate aeronautics research.
The article satirizes the performative and repetitive AI projects posted on LinkedIn, particularly computer vision demos, and describes the creation of a simple detector to identify such posts.
PSSA is a novel non-transformer language model implemented in Rust from scratch, using recurrent state-space layers and episodic memory for faster learning and inference compared to transformers.
The article discusses a benchmark where coding agents reconstructed a scientific flow diagram as editable PowerPoint slides, arguing that editable artifacts better test visual understanding by revealing structural comprehension versus pixel-level reproduction.
A developer automated a client's manual admin process by building a system with a worker portal, automated processing, and admin dashboard, reducing weekly review time from 6-8 hours to about 30 minutes.
An experiment compared two AI tools, Opus 5.5 and Sol 6.1, in creating an SVG of the Mona Lisa without reference images, evaluating their artistic output and thinking levels.
Dyna Robotics demonstrates a humanoid robot capable of performing laundry tasks such as loading and unloading washers and dryers, as well as folding and stacking towels.
The article discusses the limitations of banning AI use by children in schools and emphasizes the need for comprehensive AI education to foster independent judgment and responsible use.
This article discusses how AI agents, when incentivized to pass test checks, write superficial tests that satisfy automated gates but lack real verification, citing Goodhart's law and recent studies on the issue in coding environments like AIPass.
The article highlights the Strata inference engine, which significantly outperforms llama.cpp for running Qwen3.8 models on a laptop with 12GB VRAM and 64GB RAM, achieving up to 50 tokens per second for text generation and 1500 tokens per second for prompt processing.
This paper introduces Self-Play Search Distillation (SPSD), a framework that uses self-play in board games to generate synthetic data for improving large language model reasoning, with demonstrated improvements on mathematical benchmarks.
The paper proposes a lightweight, rule-based selector for diverse SFT traces to improve post-RL generalization in reasoning models, showing significant performance gains on mathematical benchmarks.
LeRF introduces a method to enhance perspective-taking reasoning in vision-language models by learning reference coordinate frames, improving performance on benchmarks through supervised fine-tuning and reinforcement learning.
CineSubBench is a benchmark for evaluating LLMs on long-form narrative and cultural understanding from multilingual movie subtitles, with tasks covering narrative reconstruction, genre prediction, and safety across six languages.
This paper explores approximating softmax in pretrained LLMs for kernel acceleration, demonstrating performance gains like up to 25.8% speedup on Blackwell B200 with minimal perplexity impact.
VideoPhysEdit is a training-free pipeline for physical counterfactual video editing that uses rigid-body physical scene reconstruction to simulate edits and generate accurate downstream motions and interactions.
TGRL proposes a temperature-grouped reinforcement learning method to enhance exploration in large language models, achieving faster training and improved performance across multiple benchmarks.