llm-evaluation

Tag

Cards List
#llm-evaluation

@rohanpaul_ai: A self-improving agent can keep its weights frozen and still get worse through the memories it learns to trust. These a…

X AI KOLs Following · 9h ago Cached

The paper 'Memory Reward Inflation in Self-Improving LLM Agents' shows that self-improving agents with frozen weights can still degrade by trusting flawed LLM-generated memory scores, with models endorsing 31–54% of their own wrong answers. This 'Echo Gap' persists across stronger LLMs.

0 favorites 0 likes
#llm-evaluation

Is anyone else finding DeepSeek-V4-Flash unreliable for non-coding tasks?

Reddit r/LocalLLaMA · yesterday

A user reports that DeepSeek-V4-Flash-0731 is unreliable for non-coding office tasks like summarization and meeting notes, failing at concept extraction and speaker understanding despite strong benchmark scores, while Gemma-4-31B performs better.

0 favorites 0 likes
#llm-evaluation

@LangChain: .@mondaydotcom is returning to Interrupt! Director of Engineering Dor Cohen will share how http://monday.com built a du…

X AI KOLs Timeline · 2d ago Cached

monday.com's Director of Engineering Dor Cohen will speak at Interrupt London about building a dual-pipeline LLM eval system using LangSmith production traces for online and offline evaluation.

0 favorites 0 likes
#llm-evaluation

Decomposed Entailment for Factuality Checking and Hallucination Detection

arXiv cs.CL · 2d ago Cached

This paper presents HallDetect, a lightweight and reference-free framework for hallucination detection that decomposes generated content into atomic claims and verifies them with a compact entailment model. It outperforms comparably resourced baselines on multiple benchmarks and provides a claim-to-span audit trail.

0 favorites 0 likes
#llm-evaluation

32 total local models tested head to head

Reddit r/LocalLLaMA · 3d ago

A head-to-head evaluation of 32 local language models on a fact-extraction corpus finds that most models are statistically indistinguishable, with LFM2.5 models performing significantly worse despite larger sizes.

0 favorites 0 likes
#llm-evaluation

A cheaper AI model is not necessarily cheaper once retries are counted

Reddit r/artificial · 3d ago

A side-by-side coding experiment comparing GPT-5.6 Luna and DeepSeek V4 Flash shows that DeepSeek's apparent 5x price advantage shrinks when retries are included. The article argues for more comprehensive benchmarks reporting cost per attempt and cost per verified success.

0 favorites 0 likes
#llm-evaluation

IslamicTurathBench: A Multi-Task, Multi-Discipline Benchmark for Evaluating Large Language Models on the Islamic Scholarly Tradition (turath)

arXiv cs.CL · 3d ago Cached

IslamicTurathBench (ISTB) is a new multi-task, multi-discipline benchmark for evaluating large language models on classical Islamic scholarship, containing 3,465 expert-reviewed questions across 35 works and seven fields.

0 favorites 0 likes
#llm-evaluation

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark

arXiv cs.CL · 3d ago Cached

This paper introduces ProverbIT, a novel Italian benchmark of 100 multiple-choice questions to test LLMs' ability to complete proverbs. Evaluating 13 models, it finds that performance drops significantly in multiple-choice formats without correct answers, suggesting reliance on memorized patterns rather than deep semantic understanding.

0 favorites 0 likes
#llm-evaluation

Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning

arXiv cs.CL · 3d ago Cached

This arXiv paper evaluates theory of mind capabilities in reasoning LLMs, finding increased robustness to prompt variations and task perturbations. The authors interpret gains as evidence for a robustness-based account rather than a new ToM-specific ability.

0 favorites 0 likes
#llm-evaluation

When Absence Is Evidence: Evaluating Completeness-Sensitive Negative Reasoning in Large Language Models

arXiv cs.CL · 3d ago Cached

Introduces CROWN-QA, a benchmark for completeness-sensitive negative reasoning in LLMs, showing models struggle to distinguish justified negative answers from insufficient evidence, often over-closing.

0 favorites 0 likes
#llm-evaluation

FinReportBench: Measuring and Improving Institution-Grade Financial Report Generation

arXiv cs.CL · 3d ago Cached

FinReportBench is an expert-grounded benchmark for measuring and improving institution-grade financial report generation, with 35 observable criteria across deliverability, report identity, and institutional completeness. It curates 244 bilingual tasks, evaluates nine model families, and uses benchmark-guided skill distillation to improve generation and self-review across five model families.

0 favorites 0 likes
#llm-evaluation

The Calibration Floor: Format Repair Can Masquerade as Self-Correction at Small-to-Mid Scale

arXiv cs.CL · 3d ago Cached

This paper shows that apparent LLM self-correction gains often stem from format repair rather than improved reasoning. Across multiple model scales, format effects dominate content effects, with content margins near zero on capable models, suggesting the field has misattributed a minority of measured self-correction to actual content improvement.

0 favorites 0 likes
#llm-evaluation

Mind the Cap: Output-Budget Regimes Change the Measured Multilingual Reasoning Gap

arXiv cs.CL · 3d ago Cached

This paper investigates whether the native-vs-translate multilingual reasoning gap on MGSM is an artifact of output-token budgets. The authors show that the measured gap swings significantly across different hard caps and that length normalization can reverse strategy rankings, concluding that output caps should be treated as an independent variable in evaluations.

0 favorites 0 likes
#llm-evaluation

Item Response Theory for AI Safety

arXiv cs.AI · 3d ago Cached

This paper applies Item Response Theory to eight safety benchmarks across 192 language models, identifying three latent factors, enabling 97-99% cost reduction via adaptive testing, and supporting sandbagging detection and model auditing.

0 favorites 0 likes
#llm-evaluation

MatrAIx: Simulating the World with 8.3 Billion Persona Agents

arXiv cs.AI · 3d ago Cached

This paper introduces MatrAIx, a population-scale simulated-user evaluation infrastructure using 8.3 billion persona records to test AI systems and digital products. It provides a quality-filtered coreset of ~1 million personas, multiple evaluation environments, and validation studies showing high persona adherence.

0 favorites 0 likes
#llm-evaluation

Don't Let Me Ask for It: LLMs Show Deficiencies in Active Multi-Turn Information Acquisition for Abductive Inference

arXiv cs.CL · 4d ago Cached

This paper introduces Alien Abduction, an interactive game to probe how LLMs acquire evidence, update hypotheses, and decide when to stop during abductive reasoning. It finds that models perform better with upfront evidence and with oracle-provided examples than with self-selected queries, revealing deficiencies in active information acquisition.

0 favorites 0 likes
#llm-evaluation

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks

arXiv cs.CL · 4d ago Cached

This paper tests whether commonsense benchmark scores predict real-world downstream task performance by evaluating 23 LLMs across four benchmarks and their reworked variants, finding that revisions preserve rankings but only offer task-dependent predictive validity.

0 favorites 0 likes
#llm-evaluation

Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation

arXiv cs.CL · 4d ago Cached

This paper distinguishes between two tasks in LLM opinion simulation—emulation (generating individual responses) and estimation (directly predicting population distributions)—and finds that base models are better emulators while post-trained models are better estimators, using the Pew American Trends Panel.

0 favorites 0 likes
#llm-evaluation

Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models

arXiv cs.CL · 4d ago Cached

This paper proposes a multidimensional evaluation framework for assessing statistical reasoning in large language models, combining response accuracy, response behavior, structural topic modeling, and lexical similarity analysis across 15 LLMs and 90 exam questions. It finds that accuracy alone is insufficient to characterize LLM statistical reasoning and that vendor-specific stylistic differences exist.

0 favorites 0 likes
#llm-evaluation

Mapping the City Through the Lens of Language Models

arXiv cs.CL · 4d ago Cached

This paper measures the implicit assumptions language models make about 'a city' by scoring anonymized urban profiles across 40 indicators, finding a shared preference for larger, faster-growing, and more infrastructure-rich cities. It uses open-weight checkpoints and replication data to make the default portrait of cities in LLMs empirically traceable.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback