Tag
The article discusses how frontier LLM models are improving in reducing hallucination rates, and argues that humans also hallucinate frequently, suggesting we should trust advanced AI models more while maintaining critical thinking.
A tweet humorously notes that as AI becomes smarter, the author's ability to write sentences without spelling or grammar errors diminishes.
A benchmark release comparing 24 LLMs to human writers on 475 creative writing prompts, using a custom reward model to show that frontier models outperform amateurs but not professionals.
This article debates whether current top AI models are more intelligent than an average human, stressing the need to define intelligence first.
A paper compares 8 LLMs with over 18,000 human learners, finding that high accuracy in LLMs can hide disconnected foundational knowledge, and recommends evaluating with connected problem sets.
The paper presents a systematic cross-model evaluation of how large language models interpret verbal probability expressions, finding they track human benchmarks with fidelity but exhibit biases, particularly for negative expressions, with implications for human-AI uncertainty communication.
Discusses the advanced capabilities of frontier AI models like GPT-5.6 Sol, questioning whether they constitute AGI and reflecting on AI becoming a commodity for complex cognitive tasks.
This paper compares semantic search dynamics between humans and LLMs using verbal fluency data, finding that humans exhibit more variable and exploratory search patterns that current models fail to reproduce.