ai-judges

Tag

Cards List
#ai-judges

Judging LLM-as-a-Judge: Concerning Rubric Artifacts in LLM-based Automated Text Generation Evaluation

arXiv cs.CL · 4d ago Cached

This paper demonstrates that rubric text alone can predict LLM judge outputs, challenging the assumption of rubric-based evaluation and raising concerns about its reliability in automated text generation assessment.

0 favorites 0 likes
#ai-judges

LLM-as-judge anchored on one confidence value in 10 of 16 evals. Asking for a label fixed it.

Reddit r/AI_Agents · 2026-08-28

An LLM judge consistently returned a fixed confidence score of 0.72 in evaluations, but switching to categorical labels improved score distribution, showing that models are better at classification than numerical estimation for assessments.

0 favorites 0 likes
#ai-judges

Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response Evaluation

arXiv cs.CL · 2026-08-10 Cached

Introduces counterfactual audits to test whether audio language model judges actually use paralinguistic evidence when evaluating speech-to-speech responses, finding that contrastive success often overstates native reliability and similar accuracies can hide different failure modes across Gemini, GPT, and open models.

0 favorites 0 likes
#ai-judges

BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges

arXiv cs.LG · 2026-07-21 Cached

BACON proposes a four-stage pipeline that combines budgeted human labels with multiple AI judge outputs to produce calibrated item-level surrogate predictions, supporting both population-level estimation and individual-level scoring with improved accuracy and reduced bias.

0 favorites 0 likes
#ai-judges

@omarsar0: LLM-as-a-Judge explained in ~10 mins. Knowing how to build AI verifiers and judges is one of the most important emergin…

X AI KOLs Following · 2026-06-29 Cached

A quick introduction to the LLM-as-a-Judge concept, explaining how to build AI verifiers and judges, and pointing to resources to learn more.

0 favorites 0 likes
← Back to home

Submit Feedback