typed-evals

Tag

Cards List
#typed-evals

Reduced my Jev judge’s calibration error [D]

Reddit r/MachineLearning ↗ · yesterday

The author benchmarked the Jev judge on a dataset and reduced its calibration error by 68.1% through learning from human-labelled examples, improving confidence alignment for production use without significantly changing classification accuracy.

0 favorites 0 likes
← Back to home

Submit Feedback