@houjun_liu: new method day! Trust[ing results] in ML conferences is utterly broken. Let's fix it with an algorithm! We are excited …
Summary
Researchers introduce "Training Witnesses," a system for generating "honesty" certificates of ML training data and evaluations that conference reviewers can cheaply verify and trust, aiming to fix the broken trust model in ML conferences.
View Cached Full Text
Cached at: 10/02/26, 10:46 PM
🚨new method day! 🚨 Trust[ing results] in ML conferences is utterly broken. Let’s fix it with an algorithm!
We are excited to introduce 𝗧𝗿𝗮𝗶𝗻𝗶𝗻𝗴 𝗪𝗶𝘁𝗻𝗲𝘀𝘀𝗲𝘀, a system for “honesty” certificates of data and evaluations which a verifier can cheaply check and trust. https://t.co/MSzm1P0lI8
Similar Articles
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation
This paper proposes a blockchain-based commit-reveal protocol to decentralize trust in LLM benchmarking, using anonymous multi-model verifiers to address identity-aware bias and manipulation in benchmark claims.
@omarsar0: NEW AI paper worth bookmarking. This is something I called early, and this paper confirms it: verification has emerged …
This paper from Stanford, NVIDIA, and UC Berkeley introduces LLM-as-a-Verifier, a training-free verification framework that uses continuous scoring from LLM logits to improve accuracy across coding, robotics, and medical domains, achieving state-of-the-art results on multiple benchmarks.
@jchudnov: Pass@k and self-consistency work great for math and code; sample more and verify. So we asked: can the same trick scale…
A new paper shows that scaling inference compute via methods like self-consistency improves LLM accuracy in math and code but fails to improve truthfulness in domains without external verifiers, as model errors are too correlated.
A Systematic Study of Training-Free Methods for Trustworthy Large Language Models
A systematic study evaluating training-free methods for improving trustworthiness in large language models, categorizing approaches into input, internal, and output-level interventions while analyzing trade-offs between trustworthiness, utility, and robustness.
Position: Let's Strengthen Verifiability If We Can't Enforce Reproducibility
This position paper analyzes the reproducibility crisis in machine learning—quantifying unavailable code and unreproducible results—and proposes concrete measures to strengthen result checkability and verifiability in ML research.