rubric-scoring

Tag

Cards List
#rubric-scoring

K-Bench: measuring model performance on real scientific agent requests

arXiv cs.AI · 16h ago Cached

The paper introduces K-Bench 01, a benchmark for evaluating AI agents on real scientific requests, revealing that no model consistently meets the threshold for acceptable performance, with overclaiming as a common failure.

0 favorites 0 likes
#rubric-scoring

EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking

Hugging Face Daily Papers · 4d ago Cached

EviRank reformulates multimodal image re-ranking as semantic constraint satisfaction using structured evidence packages, achieving state-of-the-art performance without training.

0 favorites 0 likes
#rubric-scoring

From Scoring to Explanations: Evaluating SHAP and LLM Rationales for Rubric-based Teaching Quality Assessment

arXiv cs.CL · 2026-06-05 Cached

This paper proposes a framework for sentence-level interpretability of rubric-based scoring, comparing SHAP and LLM-generated rationales. It finds that fine-tuned pretrained language models outperform LLMs in prediction accuracy, and SHAP provides more faithful and transferable explanations.

0 favorites 0 likes
← Back to home

Submit Feedback