@thsottiaux: Do you still trust benchmarks or do you just listen to your friends? What makes you try a new model?
Summary
A tweet questioning the trustworthiness of benchmarks and asking what drives users to try new AI models.
Similar Articles
@svpino: I don't trust benchmarks. We've all seen this movie: New model beats everyone else on a benchmark. People hype it. Then…
The tweet critiques traditional AI benchmarks and introduces TRACES, a new benchmark that evaluates AI's discovery process by focusing on how models reach answers, including tool usage, error correction, and evidence tracing.
@vasuman: The only benchmark that matters is how AI power users on Twitter feel about your model
A tweet by @vasuman suggests that the sentiment of AI power users on Twitter is the most crucial benchmark for evaluating AI models.
I stopped trusting model benchmarks and started running my own eval set, here is what changed[D]
The author describes losing faith in public AI model benchmarks due to vendor-created metrics, self-reported parameters, and lack of independent verification, and advocates for building custom evaluation sets from real production traffic to make more relevant model comparisons.
@charliermarsh: I fear the gap between benchmarks and capabilities is getting wider
A Twitter discussion highlights growing concerns that AI benchmarks fail to reflect real-world model capabilities, as they test isolated tasks rather than complex, long-running prompts encountered by users.
@Alex_ybuild: Hehe, let me solemnly introduce to everyone the human-facing benchmark—humanbench https://humanbench.ybuild.ai Come on,…
The tweet introduces humanbench, a benchmark tool for evaluating AI models to determine their size and personality characteristics.