Tag
The paper 'Memory Reward Inflation in Self-Improving LLM Agents' shows that self-improving agents with frozen weights can still degrade by trusting flawed LLM-generated memory scores, with models endorsing 31–54% of their own wrong answers. This 'Echo Gap' persists across stronger LLMs.
A user reports that DeepSeek-V4-Flash-0731 is unreliable for non-coding office tasks like summarization and meeting notes, failing at concept extraction and speaker understanding despite strong benchmark scores, while Gemma-4-31B performs better.
monday.com's Director of Engineering Dor Cohen will speak at Interrupt London about building a dual-pipeline LLM eval system using LangSmith production traces for online and offline evaluation.
This paper presents HallDetect, a lightweight and reference-free framework for hallucination detection that decomposes generated content into atomic claims and verifies them with a compact entailment model. It outperforms comparably resourced baselines on multiple benchmarks and provides a claim-to-span audit trail.
A head-to-head evaluation of 32 local language models on a fact-extraction corpus finds that most models are statistically indistinguishable, with LFM2.5 models performing significantly worse despite larger sizes.
A side-by-side coding experiment comparing GPT-5.6 Luna and DeepSeek V4 Flash shows that DeepSeek's apparent 5x price advantage shrinks when retries are included. The article argues for more comprehensive benchmarks reporting cost per attempt and cost per verified success.
IslamicTurathBench (ISTB) is a new multi-task, multi-discipline benchmark for evaluating large language models on classical Islamic scholarship, containing 3,465 expert-reviewed questions across 35 works and seven fields.
This paper introduces ProverbIT, a novel Italian benchmark of 100 multiple-choice questions to test LLMs' ability to complete proverbs. Evaluating 13 models, it finds that performance drops significantly in multiple-choice formats without correct answers, suggesting reliance on memorized patterns rather than deep semantic understanding.
This arXiv paper evaluates theory of mind capabilities in reasoning LLMs, finding increased robustness to prompt variations and task perturbations. The authors interpret gains as evidence for a robustness-based account rather than a new ToM-specific ability.
Introduces CROWN-QA, a benchmark for completeness-sensitive negative reasoning in LLMs, showing models struggle to distinguish justified negative answers from insufficient evidence, often over-closing.
FinReportBench is an expert-grounded benchmark for measuring and improving institution-grade financial report generation, with 35 observable criteria across deliverability, report identity, and institutional completeness. It curates 244 bilingual tasks, evaluates nine model families, and uses benchmark-guided skill distillation to improve generation and self-review across five model families.
This paper shows that apparent LLM self-correction gains often stem from format repair rather than improved reasoning. Across multiple model scales, format effects dominate content effects, with content margins near zero on capable models, suggesting the field has misattributed a minority of measured self-correction to actual content improvement.
This paper investigates whether the native-vs-translate multilingual reasoning gap on MGSM is an artifact of output-token budgets. The authors show that the measured gap swings significantly across different hard caps and that length normalization can reverse strategy rankings, concluding that output caps should be treated as an independent variable in evaluations.
This paper applies Item Response Theory to eight safety benchmarks across 192 language models, identifying three latent factors, enabling 97-99% cost reduction via adaptive testing, and supporting sandbagging detection and model auditing.
This paper introduces MatrAIx, a population-scale simulated-user evaluation infrastructure using 8.3 billion persona records to test AI systems and digital products. It provides a quality-filtered coreset of ~1 million personas, multiple evaluation environments, and validation studies showing high persona adherence.
This paper introduces Alien Abduction, an interactive game to probe how LLMs acquire evidence, update hypotheses, and decide when to stop during abductive reasoning. It finds that models perform better with upfront evidence and with oracle-provided examples than with self-selected queries, revealing deficiencies in active information acquisition.
This paper tests whether commonsense benchmark scores predict real-world downstream task performance by evaluating 23 LLMs across four benchmarks and their reworked variants, finding that revisions preserve rankings but only offer task-dependent predictive validity.
This paper distinguishes between two tasks in LLM opinion simulation—emulation (generating individual responses) and estimation (directly predicting population distributions)—and finds that base models are better emulators while post-trained models are better estimators, using the Pew American Trends Panel.
This paper proposes a multidimensional evaluation framework for assessing statistical reasoning in large language models, combining response accuracy, response behavior, structural topic modeling, and lexical similarity analysis across 15 LLMs and 90 exam questions. It finds that accuracy alone is insufficient to characterize LLM statistical reasoning and that vendor-specific stylistic differences exist.
This paper measures the implicit assumptions language models make about 'a city' by scoring anonymized urban profiles across 40 indicators, finding a shared preference for larger, faster-growing, and more infrastructure-rich cities. It uses open-weight checkpoints and replication data to make the default portrait of cities in LLMs empirically traceable.