Position: Evaluation Scores Are Perishable Knowledge Claims
摘要
This position paper argues that language model evaluation scores should be treated as perishable knowledge claims, not ground truth, and proposes explicit metadata such as formality tier, scope declaration, and expiration date to counter 'trust inflation' caused by averaging weak and strong signals.
查看缓存全文
缓存时间: 2026/07/31 04:00
# Position: Evaluation Scores Are Perishable Knowledge Claims
Source: [https://arxiv.org/html/2607.26191](https://arxiv.org/html/2607.26191)
Sankalp Gilda DeepThought Solutions sankalp@deepthoughtsolutions\.xyz &Shlok Gilda Meta shlokgilda@meta\.com
###### Abstract
Evaluation methodologies for language models increasingly combine multiple signals, from automated metrics and LLM\-as\-judge ratings to human assessments and benchmark suite results\. When these signals are aggregated via averaging, evaluation confidence can then substantially exceed the reliability of the weakest signal: a phenomenon we call*trust inflation in evaluation*\. We argue that evaluation scores should be treated as epistemic claims with three properties:*formality*\(human evaluation provides stronger evidence than an automated metric\),*scope*\(a benchmark result applies to the tested distribution, not universally\), and*validity windows*\(benchmark results expire as contamination accumulates and distributions shift\)\. Several converging research traditions \(chain\-of\-thought analysis, possibilistic logic, and algebraic theory\) establish weakest\-link aggregation as the conservative endpoint of a parameterized operator family controlled by a single pessimism parameter\. Drawing on those traditions, and on concrete lessons from building an evaluation harness for agentic AI, we propose that evaluation results carry explicit metadata \(formality tier, scope declaration, and expiration date\) to make their epistemic status transparent\. We illustrate the cost of mean aggregation on the public HELM leaderboard: across 54 frontier models on ten scenarios, the top\-five models ranked by mean score and by weakest\-link are completely disjoint\.
Position: Evaluation Scores Are Perishable Knowledge Claims
Sankalp GildaDeepThought Solutionssankalp@deepthoughtsolutions\.xyzShlok Gilda††thanks:Work done while at the University of Florida\.Metashlokgilda@meta\.com
††footnotetext:Published in the Proceedings of the Fifth Workshop on Generation, Evaluation and Metrics \(GEM\), ACL 2026, San Diego, California, USA, pages 1029–1035\. Association for Computational Linguistics\.[https://aclanthology\.org/2026\.gem\-main\.80/](https://aclanthology.org/2026.gem-main.80/)\(doi:10\.18653/v1/2026\.gem\-main\.80\)\. Licensed CC BY 4\.0\.## 1Introduction
The evaluation of language models rests on a cracked foundation\(Gehrmannet al\.,[2022](https://arxiv.org/html/2607.26191#bib.bib8)\)\. Prompt sensitivity studies show that minor formatting changes—switching enumerator style, reordering answer choices, adjusting whitespace—can swing model accuracy by ten percentage points or more\(Habbaet al\.,[2025](https://arxiv.org/html/2607.26191#bib.bib13)\)\. Six years of reproducibility studies find that the majority of human NLG evaluations fail to reproduce, with original\-vs\-reproduced system ranking correlations frequently belowρ=0\.8\\rho=0\.8\(Belzet al\.,[2023](https://arxiv.org/html/2607.26191#bib.bib2)\)\. Benchmark contamination gives static test sets a shelf life of six to twelve months before training\-data overlap renders scores meaningless\(Whiteet al\.,[2024](https://arxiv.org/html/2607.26191#bib.bib16)\)\. And the rapidly growing LLM\-as\-judge literature, recently surveyed byGuet al\.\([2024](https://arxiv.org/html/2607.26191#bib.bib12)\), documents systematic failure modes including style\-over\-substance bias\(Feueret al\.,[2025](https://arxiv.org/html/2607.26191#bib.bib6)\), length and position effects, and degradation when judges share a model family with the system being evaluated\.
These problems are studied in isolation: contamination detection, annotation quality, metric robustness, judge calibration\. We argue that they share a common structural cause\. Evaluation scores are treated as ground truth: fixed quantities to be measured ever more precisely\. They are not\. They are*knowledge claims*: assertions about system quality that carry implicit assumptions about formality, scope, and temporal validity\. When these assumptions are hidden and scores are aggregated by averaging, the resulting confidence systematically exceeds the reliability of the weakest evaluation signal\. We call this failure mode*trust inflation in evaluation*\.
The term echoes financial trust inflation, where structured products repackaged weak assets into apparently strong ones\. Three LLM\-as\-judge ratings from the same model family do not constitute independent evaluation evidence, yet standard aggregation treats them as additive\(Boubdiret al\.,[2023](https://arxiv.org/html/2607.26191#bib.bib4)\)\. A benchmark score from 2023 does not validate a system in 2026, yet leaderboards print it next to fresh results without qualification\.
We propose treating evaluation results as epistemic artifacts with explicit metadata: a*formality tier*for evidence strength \(Section[3](https://arxiv.org/html/2607.26191#S3)\), a*scope declaration*bounding applicability, and a*validity window*after which the result should be re\-evaluated\. We ground these proposals in the weakest\-link aggregation principle, supported by several converging research traditions\(Jacoviet al\.,[2024](https://arxiv.org/html/2607.26191#bib.bib14); Dubois and Prade,[2025](https://arxiv.org/html/2607.26191#bib.bib5)\), and in concrete engineering lessons from building an evaluation harness for agentic AI systems \(Section[4](https://arxiv.org/html/2607.26191#S4)\)\.Gilda and Gilda \([2026a](https://arxiv.org/html/2607.26191#bib.bib10)\)formalize these properties as requirements for AI\-assisted engineering more broadly; the present paper extends them to evaluation methodology\.
## 2Trust inflation in evaluation
Trust inflationoccurs when an evaluation pipeline’s aggregate confidence in a system’s quality exceeds the reliability of the weakest evaluation signal supporting that assessment\. It is a systemic property of how scores are combined, not a deficiency of any individual metric\.
### Worked Example\.
Consider a model evaluated on four dimensions: reasoning \(0\.920\.92\), factuality \(0\.410\.41\), fluency \(0\.950\.95\), and coherence \(0\.880\.88\)\. The arithmetic mean is0\.790\.79, suggesting a competent system\. The minimum is0\.410\.41, and the system’s factuality, often the most safety\-critical dimension, is masked by strong performance elsewhere\. If deployment decisions scale with aggregate confidence, averaging warrants deployment that conservative aggregation would block\.
This is not hypothetical\.Feueret al\.\([2025](https://arxiv.org/html/2607.26191#bib.bib6)\)show that LLM judges assign higher scores to longer, more polished answers even when they contain factual errors\. Aggregate evaluation scores are inflated along exactly this dimension\. A complementary illustration comes from imbalanced multi\-class classification: the gap between micro and macro F1 \(88\.76%88\.76\\%vs\.67\.98%67\.98\\%in multi\-dimensional toxicity classification,Gildaet al\.,[2021](https://arxiv.org/html/2607.26191#bib.bib9)\) shows how frequency\-weighted aggregation masks weakness on hard dimensions\.
### Three Mechanisms\.
Trust inflation in evaluation operates through three channels:
1. 1\.Signal averaging: when benchmark suites report aggregate scores across sub\-tasks, weak performance on critical capabilities is diluted by strong performance on common ones\.
2. 2\.Self\-referential evaluation: an LLM that generates text and an LLM\-as\-judge from the same model family that evaluates it share training data, biases, and failure modes\. The evaluation is not independent; it is self\-assessment, shown to be only 25–39% faithful to actual model computation\(Anthropic,[2025](https://arxiv.org/html/2607.26191#bib.bib1)\)\.
3. 3\.Temporal staleness: benchmark datasets, annotation guidelines, and leaderboard rankings stay in place after contamination, distribution shift, or model updates render them obsolete\.
We term the constraint underlying mechanism 2 the*Transformer Mandate*\(Gilda and Gilda,[2026a](https://arxiv.org/html/2607.26191#bib.bib10)\): no system can be the authoritative evaluator of its own outputs\. In evaluation methodology, this means LLM\-as\-judge scores from the same model family as the evaluated system should be classified as self\-assessment \(F0 ceiling, Table[1](https://arxiv.org/html/2607.26191#S3.T1)\), not independent evaluation\.
### Weakest\-Link as the Conservative Endpoint\.
The worked example above exposes a general principle: when evaluation dimensions are serially dependent \(factuality must hold before fluency matters\), the aggregate reliability cannot exceed the minimum of its components\. This is the*weakest\-link principle*\(WLNK\)\.Jacoviet al\.\([2024](https://arxiv.org/html/2607.26191#bib.bib14)\)demonstrate empirically that the lowest\-confidence reasoning step predicts chain\-of\-thought failure better than any average\.Dubois and Prade \([2025](https://arxiv.org/html/2607.26191#bib.bib5)\)establish weakest\-link resolution as a fundamental principle of possibilistic logic, grounded in four decades of theory\.Gilda and Gilda \([2026b](https://arxiv.org/html/2607.26191#bib.bib11)\)derive the same bound algebraically as one of five invariants on structured reasoning chains: no conclusion can exceed the reliability of its least\-supported premise\. Algebraically,min\\minis the unique idempotent continuous t\-norm—the only operator where applying the same evidence twice changes nothing—which forces it as the conservative endpoint of any serial\-aggregation family\.
We treatmin\\minnot as a uniquely correct operator but as the conservative endpoint of a parameterized family\. The ordered weighted average\(Yager,[1988](https://arxiv.org/html/2607.26191#bib.bib18)\)on evidence scores\{s1,…,sn\}\\\{s\_\{1\},\\ldots,s\_\{n\}\\\}sorted in descending orders\(1\)≥⋯≥s\(n\)s\_\{\(1\)\}\\geq\\cdots\\geq s\_\{\(n\)\}isOWA\(s;w\)=∑iwis\(i\)\\mathrm\{OWA\}\(s;w\)=\\sum\_\{i\}w\_\{i\}\\,s\_\{\(i\)\}for weightswi≥0w\_\{i\}\\geq 0,∑wi=1\\sum w\_\{i\}=1\. The pessimism parameterρ=1−β\(w\)∈\[0,1\]\\rho=1\-\\beta\(w\)\\in\[0,1\]inverts Yager’s ornessβ\(w\)=\(1/\(n−1\)\)∑i\(n−i\)wi\\beta\(w\)=\(1/\(n\-1\)\)\\sum\_\{i\}\(n\-i\)\\,w\_\{i\}, so thatρ=1\\rho=1\(orness 0\) recoversmin\\min,ρ=0\\rho=0\(orness 1\) recoversmax\\max, andρ=0\.5\\rho=0\.5recovers the arithmetic mean\. The position is not that aggregation must bemin\\min, but that the operator must be exposed and calibrated rather than fixed by fiat to the arithmetic mean\. For safety\-critical evaluation,min\\minis the appropriate default; for the routine middle of evaluation pipelines, intermediateρ\\rhois defensible if the analyst chooses to live with the trust\-inflation cost\. The abuse is the silent default: a hiddenρ≈0\.5\\rho\\approx 0\.5that gets quoted as if no aggregation choice were involved\.
## 3Evaluation as epistemic system
If evaluation scores are knowledge claims, they should carry the metadata that any knowledge claim requires: how strong is the evidence, where does it apply, and when does it expire?
### Formality Tiers\.
Not all evaluation evidence is equally rigorous\. We propose four tiers, each with a reliability ceiling reflecting the maximum trust an evaluation signal of that type can contribute:
Table 1:Formality tiers for evaluation evidence\. Ceilings cap reliability regardless of sample size\.These ceilings are not arbitrary: an F0 ceiling of 0\.70 reflects the empirical finding that LLM self\-assessment is 25–39% faithful\(Anthropic,[2025](https://arxiv.org/html/2607.26191#bib.bib1)\), and that LLM judges exhibit style\-over\-substance bias\(Feueret al\.,[2025](https://arxiv.org/html/2607.26191#bib.bib6)\)\. An F2 ceiling of 0\.95 admits that even controlled human evaluation has reproducibility limits\(Belzet al\.,[2023](https://arxiv.org/html/2607.26191#bib.bib2)\)\. Under WLNK, a benchmark suite combining F0 and F2 evidence cannot claim overall reliability above 0\.70, since the F0 component caps the aggregate\.
### Scope\.
A benchmark result applies to the distribution and conditions under which it was collected, not universally\. MMLU scores do not predict performance on domain\-specific tasks\. English\-language evaluations do not transfer to other languages\. Even purpose\-built verification tools have severe coverage limitations:Yanget al\.\([2024](https://arxiv.org/html/2607.26191#bib.bib19)\)find that Google Fact Check retrieves results for only 15\.8% of input claims, and semantically equivalent claims phrased differently yield dissimilar results 81% of the time\. Evaluation benchmarks face the same coverage and phrasing\-sensitivity problems\. Scope matching admits degrees: evidence from a narrower or broader distribution than the evaluation target should contribute with proportionally reduced weight, not be treated as either perfectly applicable or entirely irrelevant\. Scope should be declared explicitly \(task domain, language, model size range, evaluation date\) so that consumers know the boundaries of the claim\.
### Validity Windows\.
Benchmark results expire\.Whiteet al\.\([2024](https://arxiv.org/html/2607.26191#bib.bib16)\)demonstrate that static benchmarks become contaminated within months\. The DOVE study\(Habbaet al\.,[2025](https://arxiv.org/html/2607.26191#bib.bib13)\)shows that the “same” benchmark produces different results under minor prompt variations: a score’s validity is conditional on the exact evaluation configuration\. We propose that evaluation results carry explicit validity windows: an F0 crowd annotation might be valid for weeks, an F2 controlled study for months, and an F3 formal property proof indefinitely\. When evidence expires, its reliability drops to a floor value representing “uncertain, not disproved\.” This forces re\-evaluation rather than silent reliance on stale results\.
## 4Evidence from building an evaluation harness
We report lessons from building a 3,700\-line evaluation harness for comparing agentic AI approaches on ML research tasks, using controlled A/B methodology with Docker isolation, structured error classification, and paired statistical analysis\.
### Schema Volatility\.
Our evaluation output schema required 13 revisions across two output formats \(per\-run and cross\-run comparison\) in five weeks, each triggered by discovering that post\-hoc analysis required fields absent from the original design\. Without explicit schema versioning, analysis scripts silently compare scores from incompatible evaluation regimes, a form of trust inflation across time, where stale format assumptions inflate confidence in cross\-version comparisons\.
### Cross\-Boundary Semantic Bugs\.
A semantic mismatch between our Python evaluation client and Go backend caused all script failures to be silently recorded as successes\. A parameter default in one language masked the actual verdict computed by the other\. This class of bug was invisible to unit tests, integration tests, and output inspection; it required forensic database analysis to discover\. This is trust inflation at the infrastructure level: reported evaluation scores silently exceed actual system performance because the measurement process itself is corrupted\.
### Score Saturation\.
A persistent score ceiling at 90\.5% of SOTA was suspected to be a harness artifact but proved to be a deterministic property of the embedding model the LLM consistently selected\. Misattributing model limits to infrastructure limitations \(or vice versa\) misdirects evaluation effort\. The result is inflated confidence that the evaluation pipeline is measuring what it claims\.
### Tiered Evaluation as Formality in Practice\.
Our harness implements four evaluation tiers \(syntax check, sample\-based scoring, full evaluation, and no evaluation\) with explicit reliability multipliers: a sample\-based score carries a 0\.7x weight relative to a full evaluation\. This directly instantiates the formality tier concept from Section[3](https://arxiv.org/html/2607.26191#S3): evaluation speed can be traded for reliability with explicit epistemic accounting, rather than treating all evaluation signals as equivalent regardless of thoroughness\.
## 5Illustration on a public leaderboard
Figure 1:Mean\-aggregate rank vs\. weakest\-link \(WLNK\) rank for 54 models on ten HELM Lite scenarios \(Stanford CRFM, v1\.13\.0\)\. Diagonal = no change\. Color encodes rank displacement; the five largest movers are labeled\. Top\-5 by mean and top\-5 by WLNK are completely disjoint\.To make the cost of mean aggregation concrete, we apply both aggregators to two publicly released HELM leaderboards \(Stanford CRFM\)\. On HELM Capabilities v1\.0\.0 \(22 frontier models on GPQA, IFEval, MMLU\-Pro, Omni\-MATH, WildBench; one of the five uses LLM\-as\-judge scoring\), the mean and WLNK rankings give Spearmanρ=0\.87\\rho=0\.87, top\-5 Jaccard=0\.67=0\.67, and maximum rank displacement=8=8positions\. Claude 3\.5 Sonnet drops from rank 5 by mean to rank 12 by WLNK because its 0\.28 Omni\-MATH score is masked by strength elsewhere\. On the larger HELM Lite v1\.13\.0 \(54 models, 10 scenarios: narrative QA, MMLU, GSM, MATH, LegalBench, MedQA, and WMT translation\), the divergence sharpens: Spearmanρ=0\.89\\rho=0\.89, max rank displacement=21=21positions, and*top\-5 Jaccard=0\.000=0\.000*\. The five models that lead by mean and the five that lead by WLNK are completely disjoint \(Figure[1](https://arxiv.org/html/2607.26191#S5.F1)\)\. In the spirit of the position: ranking by mean rewards models that excel where rewards are easy; ranking by WLNK rewards models that do not collapse where evaluation is hardest\. A reader who silently chooses one over the other has silently picked a pessimism parameter, and the resulting leaderboard inherits that choice without disclosing it\.
A note on the formality\-tier ceilings of Section[3](https://arxiv.org/html/2607.26191#S3)\. They do bind on saturated subtasks: in HELM Lite, top\-model scores on OpenBookQA \(0\.970\.97\), GSM8K \(0\.960\.96\), Math\-CoT \(0\.920\.92\), and MedQA \(0\.860\.86\) all exceed the F1 ceiling of0\.850\.85\. In HELM Capabilities, WildBench \(0\.830\.83vs\. F00\.700\.70\) and IFEval \(0\.870\.87vs\. F10\.850\.85\) bind as well\. They do not, however, alter the WLNK aggregate at the leaderboard scale shown here, because the weakest\-link is consistently a non\-saturated subtask \(Omni\-MATH at0\.460\.46in both substrates; WMT\-14 BLEU at0\.260\.26in Lite\)\. The rank divergence in Figure[1](https://arxiv.org/html/2607.26191#S5.F1)is therefore driven by multi\-dimensional capability variance, not by tier\-clipping; the tier mechanism contributes by capping reported claims on saturated subtasks rather than by reshaping aggregate rankings\. Validity\-window decay is left for future empirical work; it is exercised qualitatively by the benchmark\-contamination evidence cited in Section[3](https://arxiv.org/html/2607.26191#S3)\.
## 6Implications and call to action
We propose four concrete changes to evaluation practice:
1. 1\.Metadata on evaluation results: every benchmark score should carry a formality tier \(Table[1](https://arxiv.org/html/2607.26191#S3.T1)\), a scope declaration \(task, language, model class, date\), and a validity window\. This makes the epistemic status of evaluation claims transparent and auditable\.
2. 2\.Expose the aggregation operator: when evaluation dimensions are aggregated, the operator should be a calibrated choice on the pessimism spectrum—weakest\-link \(min\\min\) at the conservative endpoint, arithmetic mean in the middle,max\\maxat the permissive endpoint—not a hidden default\. Two heuristics distinguish serial from parallel dependencies in practice\. First, dimensions are*serial*when one dimension’s failure undermines the meaning of another: factuality undermines coherence \(a coherently\-stated falsehood is still wrong\), and safety undermines helpfulness \(a helpful suggestion to commit a crime is still unsafe\)\. Instruction\-following undermines the downstream content quality that depends on it\. Second, dimensions are*parallel*when they probe distinguishable aspects of the same artifact whose failures are independent: lexical fluency vs\. syntactic acceptability; English performance vs\. Spanish performance on a multilingual benchmark\. For parallel dimensions, probabilistic combination appropriately credits redundant evidence\. The conservative default for serial dependencies is WLNK, but the explicit point is that the choice must be*declared*; the silent arithmetic mean is the abuse, not the participating operator\.
3. 3\.Schema versioning for evaluation outputs: evaluation output formats should be versioned from day one\. Our experience of 13 revisions in five weeks suggests this is not premature engineering but necessary hygiene for any evaluation pipeline still under change\.
4. 4\.Honest reporting infrastructure: evaluation harnesses should emit machine\-readable warnings when sample sizes are insufficient, disclose normalization differences from reference benchmarks, and document what randomness seeds do and do not control\.
These proposals complement, not replace, three layers of existing apparatus: HEDS\(Belz and Thomson,[2024](https://arxiv.org/html/2607.26191#bib.bib3)\), Model Cards\(Mitchellet al\.,[2019](https://arxiv.org/html/2607.26191#bib.bib17)\), and Datasheets\(Gebruet al\.,[2021](https://arxiv.org/html/2607.26191#bib.bib7)\)document*how*a score was produced and on*what*data\. Construct\-validity work\(Liao and Xiao,[2023](https://arxiv.org/html/2607.26191#bib.bib15)\)asks*whether*a benchmark measures what it claims\. We propose the missing fourth layer—*how much*to trust the score,*for how long*, and*how*to combine it with other scores\. DOVE\(Habbaet al\.,[2025](https://arxiv.org/html/2607.26191#bib.bib13)\)and ReproNLP\(Belzet al\.,[2023](https://arxiv.org/html/2607.26191#bib.bib2)\)expose prompt sensitivity and reproducibility failures respectively; the LLM\-as\-judge survey ofGuet al\.\([2024](https://arxiv.org/html/2607.26191#bib.bib12)\)catalogs judge fragility across dozens of recent studies\. Trust inflation gives them a unifying diagnosis\.
## Limitations
This is a position paper without large\-scale empirical validation\. We anticipate three objections\.
First,*WLNK is too conservative*: a model excelling on 9 of 10 dimensions would be capped at its worst score\. The position handles this directly via the OWA family of Section[2](https://arxiv.org/html/2607.26191#S2):min\\minis theρ=1\\rho=1endpoint of a continuous spectrum, not a unique mandate\. The actual choice is whichρ\\rhoto use in which evaluation context; the position is thatρ\\rhomust be declared, not thatρ=1\\rho=1is universally correct\. For safety\-critical deployments, we argue the conservative endpoint is the appropriate default precisely because overestimating evaluation reliability does more harm than underestimating it\.
Second,*validity windows create perverse incentives*: teams might game freshness by re\-running benchmarks without meaningful updates\. This risk exists but is mitigated by formality tiers, since refreshing an F0 crowd annotation extends only the F0 ceiling, not overall reliability\.
Third,*formality tiers calcify into bureaucracy*: rigid tier assignments could discourage methodological innovation\. We view the tiers as defaults requiring community calibration, not fixed standards\. Different evaluation contexts may warrant different ceilings\. The harness evidence is drawn from a single system; validation on other evaluation pipelines would strengthen the claims\.
## Ethics Statement
Trust inflation in evaluation can lead to premature deployment of systems whose weakest capabilities are masked by aggregate scores\. By proposing transparent epistemic metadata on evaluation results, we aim to reduce the risk of deploying systems that appear competent on average but fail on safety\-critical dimensions\. We do not propose restricting evaluation methods; we propose making their epistemic status explicit so that deployment decisions are informed by honest assessments\.
## References
- Anthropic \(2025\)Reasoning models don’t always say what they think\.Technical reportAnthropic\.Note:Measured Claude 3\.7 Sonnet at 25% faithfulness, DeepSeek R1 at 39%External Links:[Link](https://www.anthropic.com/research/reasoning-models-dont-say-think)Cited by:[item 2](https://arxiv.org/html/2607.26191#S2.I1.i2.p1.1),[§3](https://arxiv.org/html/2607.26191#S3.SS0.SSS0.Px1.p2.1)\.
- A\. Belz, S\. Agarwal, A\. Shimorina, and E\. Reiter \(2023\)Non\-repeatable experiments and non\-reproducible results: the reproducibility crisis in human evaluation in NLP\.InFindings of the Association for Computational Linguistics: ACL 2023,pp\. 3676–3689\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.findings-acl.226)Cited by:[§1](https://arxiv.org/html/2607.26191#S1.p1.1),[§3](https://arxiv.org/html/2607.26191#S3.SS0.SSS0.Px1.p2.1),[§6](https://arxiv.org/html/2607.26191#S6.p3.1)\.
- A\. Belz and C\. Thomson \(2024\)HEDS 3\.0: the human evaluation data sheet version 3\.0\.arXiv preprint arXiv:2412\.07940\.External Links:[Link](https://arxiv.org/abs/2412.07940)Cited by:[§6](https://arxiv.org/html/2607.26191#S6.p3.1)\.
- M\. Boubdir, E\. Kim, B\. Ermis, S\. Hooker, and M\. Fadaee \(2023\)Elo uncovered: robustness and best practices in language model evaluation\.InProceedings of the Third Workshop on Natural Language Generation, Evaluation, and Metrics \(GEM\),Singapore,pp\. 339–352\.External Links:[Link](https://aclanthology.org/2023.gem-1.28/)Cited by:[§1](https://arxiv.org/html/2607.26191#S1.p3.1)\.
- D\. Dubois and H\. Prade \(2025\)40 years of research in possibilistic logic – a survey\.InProceedings of the Thirty\-Fourth International Joint Conference on Artificial Intelligence \(IJCAI\-25\),pp\. 10427–10435\.Note:Survey Track\. Establishes “weakest link resolution” as fundamental principle of possibilistic inferenceExternal Links:[Document](https://dx.doi.org/10.24963/ijcai.2025/1158)Cited by:[§1](https://arxiv.org/html/2607.26191#S1.p4.1),[§2](https://arxiv.org/html/2607.26191#S2.SS0.SSS0.Px3.p1.1)\.
- B\. Feuer, M\. Goldblum, T\. Datta, S\. Nambiar, R\. Besaleli, S\. Dooley, M\. Cembalest, and J\. P\. Dickerson \(2025\)Style outweighs substance: failure modes of LLM judges in alignment benchmarking\.InProceedings of the 13th International Conference on Learning Representations \(ICLR\),External Links:[Link](https://arxiv.org/abs/2409.15268)Cited by:[§1](https://arxiv.org/html/2607.26191#S1.p1.1),[§2](https://arxiv.org/html/2607.26191#S2.SS0.SSS0.Px1.p2.2),[§3](https://arxiv.org/html/2607.26191#S3.SS0.SSS0.Px1.p2.1)\.
- T\. Gebru, J\. Morgenstern, B\. Vecchione, J\. W\. Vaughan, H\. Wallach, H\. Daumé III, and K\. Crawford \(2021\)Datasheets for datasets\.Communications of the ACM64\(12\),pp\. 86–92\.External Links:[Document](https://dx.doi.org/10.1145/3458723)Cited by:[§6](https://arxiv.org/html/2607.26191#S6.p3.1)\.
- S\. Gehrmann, E\. Clark, and T\. Sellam \(2022\)Repairing the cracked foundation: a survey of obstacles in evaluation practices for generated text\.Journal of Artificial Intelligence Research73,pp\. 767–835\.External Links:[Document](https://dx.doi.org/10.1613/jair.1.12866)Cited by:[§1](https://arxiv.org/html/2607.26191#S1.p1.1)\.
- S\. Gilda and S\. Gilda \(2026a\)AI\-assisted engineering should track the epistemic status and temporal validity of architectural decisions\.arXiv preprint arXiv:2601\.21116\.Cited by:[§1](https://arxiv.org/html/2607.26191#S1.p4.1),[§2](https://arxiv.org/html/2607.26191#S2.SS0.SSS0.Px2.p3.1)\.
- S\. Gilda and S\. Gilda \(2026b\)Structured abductive\-deductive\-inductive reasoning for LLMs via algebraic invariants\.arXiv preprint arXiv:2604\.15727\.Note:Accepted at the ICLR 2026 Workshop on Logical Reasoning of LLMsExternal Links:[Link](https://arxiv.org/abs/2604.15727)Cited by:[§2](https://arxiv.org/html/2607.26191#S2.SS0.SSS0.Px3.p1.1)\.
- S\. Gilda, M\. Silva, L\. Giovanini, and D\. Oliveira \(2021\)Predicting different types of subtle toxicity in unhealthy online conversations\.InProceedings of the International Conference on Web Intelligence and Intelligent Agent Technology,External Links:[Link](https://arxiv.org/abs/2106.03952)Cited by:[§2](https://arxiv.org/html/2607.26191#S2.SS0.SSS0.Px1.p2.2)\.
- J\. Gu, X\. Jiang, Z\. Shi, H\. Tan, X\. Zhai, C\. Xu, W\. Li, Y\. Shen, S\. Ma, H\. Liu, S\. Wang, K\. Zhang, Y\. Wang, W\. Gao, L\. Ni, and J\. Guo \(2024\)A survey on LLM\-as\-a\-judge\.arXiv preprint arXiv:2411\.15594\.External Links:[Link](https://arxiv.org/abs/2411.15594)Cited by:[§1](https://arxiv.org/html/2607.26191#S1.p1.1),[§6](https://arxiv.org/html/2607.26191#S6.p3.1)\.
- E\. Habba, O\. Arviv, I\. Itzhak, Y\. Perlitz, E\. Bandel, L\. Choshen, M\. Shmueli\-Scheuer, and G\. Stanovsky \(2025\)DOVE: a large\-scale multi\-dimensional predictions dataset towards meaningful LLM evaluation\.InFindings of the Association for Computational Linguistics: ACL 2025,External Links:[Link](https://arxiv.org/abs/2503.01622)Cited by:[§1](https://arxiv.org/html/2607.26191#S1.p1.1),[§3](https://arxiv.org/html/2607.26191#S3.SS0.SSS0.Px3.p1.1),[§6](https://arxiv.org/html/2607.26191#S6.p3.1)\.
- A\. Jacovi, Y\. Bitton, B\. Bohnet, J\. Herzig, O\. Honovich, M\. Tseng, M\. Collins, R\. Aharoni, and M\. Geva \(2024\)A chain\-of\-thought is as strong as its weakest link: a benchmark for verifiers of reasoning chains\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(ACL 2024\),pp\. 1–20\.Note:Independently validates WLNK principle: reasoning chain reliability equals its weakest stepExternal Links:[Link](https://arxiv.org/abs/2402.00559)Cited by:[§1](https://arxiv.org/html/2607.26191#S1.p4.1),[§2](https://arxiv.org/html/2607.26191#S2.SS0.SSS0.Px3.p1.1)\.
- Q\. V\. Liao and Z\. Xiao \(2023\)Rethinking model evaluation as narrowing the socio\-technical gap\.arXiv preprint arXiv:2306\.03100\.External Links:[Link](https://arxiv.org/abs/2306.03100)Cited by:[§6](https://arxiv.org/html/2607.26191#S6.p3.1)\.
- M\. Mitchell, S\. Wu, A\. Zaldivar, P\. Barnes, L\. Vasserman, B\. Hutchinson, E\. Spitzer, I\. D\. Raji, and T\. Gebru \(2019\)Model cards for model reporting\.InProceedings of the Conference on Fairness, Accountability, and Transparency \(FAT\*\),pp\. 220–229\.External Links:[Document](https://dx.doi.org/10.1145/3287560.3287596)Cited by:[§6](https://arxiv.org/html/2607.26191#S6.p3.1)\.
- C\. White, S\. Dooley, M\. Roberts, A\. Pal, B\. Feuer, S\. Jain, R\. Shwartz\-Ziv, N\. Jain, K\. Saifullah, S\. Naidu, C\. Hegde, Y\. LeCun, T\. Goldstein, W\. Neiswanger, and M\. Goldblum \(2024\)LiveBench: a challenging, contamination\-free LLM benchmark\.arXiv preprint arXiv:2406\.19314\.External Links:[Link](https://arxiv.org/abs/2406.19314)Cited by:[§1](https://arxiv.org/html/2607.26191#S1.p1.1),[§3](https://arxiv.org/html/2607.26191#S3.SS0.SSS0.Px3.p1.1)\.
- R\. R\. Yager \(1988\)On ordered weighted averaging aggregation operators in multicriteria decision making\.IEEE Transactions on Systems, Man, and Cybernetics18\(1\),pp\. 183–190\.External Links:[Document](https://dx.doi.org/10.1109/21.87068)Cited by:[§2](https://arxiv.org/html/2607.26191#S2.SS0.SSS0.Px3.p2.17)\.
- Q\. Yang, T\. Christensen, S\. Gilda, J\. Fernandes, D\. Oliveira, R\. Wilson, and D\. Woodard \(2024\)Are fact\-checking tools helpful? an exploration of the usability of Google fact check\.InProceedings of the ACM on Human\-Computer Interaction,External Links:[Link](https://arxiv.org/abs/2402.13244)Cited by:[§3](https://arxiv.org/html/2607.26191#S3.SS0.SSS0.Px2.p1.1)\.相似文章
我们是在评估知识还是措辞?利用ParaEval减轻MCQA敏感性
本文指出,标准的多选题问答基准对措辞偏差敏感,将知识与表面形式熟悉度混为一谈。作者提出了ParaEval框架,该框架为每个答案选项使用多种释义,根据最有利的措辞对模型进行评分,从而减少虚假的性能差距,实现更稳健的评估。
回收评估:有损记忆比空记忆更糟糕
本文表明,具有有损记忆的语言模型如果保留了错误结论而丢弃了证据,会产生自信的错误答案,而空记忆则会导致弃权。作者提出了一种源优先压缩策略,保留可重新计算的来源而非结论,以保持可纠正性,并在多个模型和对话系统中展示了这一机制。
Evaluation Cards: 一种AI评估报告的解释层
本文介绍了EvalCards,这是一种操作框架,通过将基准元数据、评估运行数据和模型元数据组合成一个统一记录,并包含可重现性、完整性、来源、风险和分数可比性的解释性信号,从而标准化AI评估报告。作者在数千个模型和基准测试中部署了一个监控工具,揭示了当前报告实践中的系统性差距。
理解扩散大语言模型中的评估幻觉
本文指出了扩散LLM解码方法中的评估不一致性,表明提示模板的选择会显著影响排名,并提出了可靠评估的实用指南。
识别与解决知识型VQA基准测试的陷阱:审计、修复与增强
本文对知识型VQA基准进行了审计,揭示了系统性的假设违反,使得准确率成为误导性指标。它提出了一种修复协议和多实体增强方法,以恢复答案可推导性和问题清晰度,表明修正后的设置产生了显著不同的模型排名。