validity

Tag

Cards List
#validity

Do Not Trust the Benchmark: Limitations of General LLM Rankings and a Case for Task-Specific Evaluation

arXiv cs.AI ↗ · 5d ago Cached

This paper examines the limitations of general LLM rankings, including benchmark saturation and commercial incentives, and advocates for task-specific evaluation methods to improve reliability.

0 favorites 0 likes
#validity

Append-only memory is exactly wrong when an agent needs to change its mind

Reddit r/AI_Agents ↗ · 2026-08-10

A new preprint called TEPA treats memory validity as a first-class state, revoking outdated precedents when new evidence conflicts while keeping audit trails. It outperforms append-only and last-write-wins in a complete-reversal experiment, though results are not yet independently reproduced.

0 favorites 0 likes
#validity

Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks

arXiv cs.AI ↗ · 2026-08-03 Cached

This paper audits four agent-safety benchmarks (R-Judge, InjecAgent, AgentHarm, AgentDojo) across many models, showing that their scores are confounded by capability, metrics have artifacts, and rankings disagree across benchmarks, undermining interchangeable safety claims.

0 favorites 0 likes
#validity

Hidden Consensus:Preference-Validity Compression in Human Feedback

arXiv cs.CL ↗ · 2026-06-10 Cached

This paper argues that standard RLHF's scalarization of human preferences collapses multiple valid interpretations into a single target, mis-measuring alignment in culturally plural societies. Analyzing a Malaysian dataset, they find 79% of prompts have multiple majority-supported responses that single-winner aggregation discards.

0 favorites 0 likes
#validity

Generative-Evaluative Agreement: A Necessary Validity Criterion for LLM-Enabled Adaptive Assessment

arXiv cs.AI ↗ · 2026-05-20 Cached

Introduces Generative-Evaluative Agreement (GEA), a validity criterion for LLM-enabled adaptive assessments, and measures it on a two-stage adaptive test, finding that the model recovers about half the intended variance with systematic bias.

0 favorites 0 likes
← Back to home

Submit Feedback