Tag
The paper introduces K-Bench 01, a benchmark for evaluating AI agents on real scientific requests, revealing that no model consistently meets the threshold for acceptable performance, with overclaiming as a common failure.
EviRank reformulates multimodal image re-ranking as semantic constraint satisfaction using structured evidence packages, achieving state-of-the-art performance without training.
This paper proposes a framework for sentence-level interpretability of rubric-based scoring, comparing SHAP and LLM-generated rationales. It finds that fine-tuned pretrained language models outperform LLMs in prediction accuracy, and SHAP provides more faithful and transferable explanations.