Reduced my Jev judge’s calibration error [D]

Reddit r/MachineLearning Tools

Summary

The author benchmarked the Jev judge on a dataset and reduced its calibration error by 68.1% through learning from human-labelled examples, improving confidence alignment for production use without significantly changing classification accuracy.

I benchmarked Jev on TRIVIA+ dataset using an untouched 645-example test set. Before calibration: ECE: 0.0982 After learning from human-labelled examples: ECE: 0.0313 That’s a 68.1% reduction in calibration error. But hallucination-detection F1 only moved: 0.5833 → 0.5877 So what improved? Not the judge’s ability to classify examples. Its confidence became much more aligned with reality. And that distinction matters. If a judge score is only being displayed on a dashboard, maybe not much. But if confidence > 0.8 means “automatically approve this answer,” or < 0.4 means “escalate to a human,” then badly calibrated scores can quietly become a production problem. That’s one of the reasons I’ve been building calibration into Typed Evals instead of treating raw judge confidence as trustworthy by default. Benchmark/methodology report: https://github.com/TrustifAI/typed_evals/blob/main/docs/BENCHMARK.md Would be interested to know how people here pick production thresholds for LLM/Jev judges.
Original Article

Similar Articles

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

Hugging Face Daily Papers

This paper introduces JEV-as-a-Judge, a cost-effective evaluation method for LLMs that uses a decision-only judge with confidence thresholds to accept certain verdicts and escalate uncertain ones, achieving comparable accuracy to state-of-the-art models at significantly lower cost.