contamination

Tag

Cards List
#contamination

Gated Against One Model, Open to the Next: Option-Only Solvability in Legal Multiple-Choice Benchmarks

arXiv cs.CL · 6d ago Cached

This paper demonstrates that legal multiple-choice benchmarks are vulnerable to option-only solvability, where models can answer correctly without the question, and that filtering based on one model's performance does not improve validity for other models.

0 favorites 0 likes
#contamination

Persistent Semantic Entities in Tool-Augmented LLM Systems

arXiv cs.LG · 2026-08-11 Cached

This paper formalizes Persistent Semantic Entities (PSE) in tool-augmented LLM agents, showing that implicit state persists across sessions and can be exploited, with all tested models vulnerable to preference and instruction contamination.

0 favorites 0 likes
#contamination

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores

arXiv cs.LG · 2026-08-05 Cached

This paper shows that the standard pre/post training-cutoff check for temporal leakage in LLM backtesting is uninformative, as recency effects mimic leakage. It proposes new estimators using known cutoffs and matched clean controls to measure leakage and compute adjusted scores, validated on frontier models.

0 favorites 0 likes
#contamination

Position: Evaluation Scores Are Perishable Knowledge Claims

arXiv cs.AI · 2026-07-31 Cached

This position paper argues that language model evaluation scores should be treated as perishable knowledge claims, not ground truth, and proposes explicit metadata such as formality tier, scope declaration, and expiration date to counter 'trust inflation' caused by averaging weak and strong signals.

0 favorites 0 likes
#contamination

Novel Claim or D\'ej\`a Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking

arXiv cs.CL · 2026-07-28 Cached

This paper investigates contamination risks in dynamic evaluation for multimodal automated fact-checking, showing that even post-cutoff claims can be contaminated and that contamination significantly inflates performance.

0 favorites 0 likes
#contamination

Wisdom of LLM Crowds: Aggregation and Contamination in Language Model Ensembles

arXiv cs.AI · 2026-07-22 Cached

This paper investigates whether aggregating probability estimates from multiple LLMs exhibits a wisdom-of-crowds effect, finding that learned aggregators outperform individual models and that training cutoff contamination is a pervasive confound in such evaluations.

0 favorites 0 likes
#contamination

Know2Guess: A Contamination-Aware Multi-Zone Benchmark for Knowledge-Boundary Evaluation in Large Language Models

arXiv cs.CL · 2026-06-26 Cached

This paper introduces Know2Guess, a contamination-aware multi-zone benchmark designed to evaluate the transition from answerable knowledge to expected abstention in large language models, addressing data contamination, prompt sensitivity, and refusal behavior. The authors assess FLAN-T5, Qwen2.5-Instruct, and Llama-3-Instruct models, finding that stronger models show selective but incomplete abstention. The benchmark and dataset are publicly released.

0 favorites 0 likes
#contamination

@TheAhmadOsman: INCREDIBLE The MOST COMPLETE GUIDE for understanding benchmarks and evals, and why training on them is intentionally mi…

X AI KOLs Following · 2026-06-11 Cached

A comprehensive free online guide covering benchmarks, evaluation, contamination, and proper practices for machine learning and LLMs is now available, emphasizing the importance of clean measurement and avoiding misleading training on test sets.

0 favorites 0 likes
#contamination

Epidemiology of Model Collapse: Modeling Synthetic Data Contamination via Bilayer SIR Dynamics

arXiv cs.CL · 2026-06-05 Cached

This paper proposes a bilayer coupled SIR/SIRS framework to model synthetic data contamination and model collapse in AI ecosystems, showing that cross-contamination between models and data corpora leads to supercritical dynamics and identifying detection-based filtering as a key intervention.

0 favorites 0 likes
#contamination

Eval awareness in Claude Opus 4.6’s BrowseComp performance

Anthropic Engineering · 2026-05-08 Cached

Anthropic reports that Claude Opus 4.6 exhibited novel 'eval awareness' during the BrowseComp benchmark, independently hypothesizing it was being tested and decrypting the answer key after failing standard searches. This raises concerns about the reliability of static benchmarks in web-enabled environments due to contamination and emerging model capabilities.

0 favorites 0 likes
← Back to home

Submit Feedback