Tag
This paper proposes a conditional generalizability framework to evaluate nonuniform dependability across response conditions in automated essay scoring.
Systematic study shows LLM-based dense retrievers outperform BERT baselines on typos and poisoning but remain vulnerable to semantic perturbations, with embedding geometry predicting robustness.