Tag
Small language models can grade open-ended exam answers as reliably as more expensive models when using an explicit rubric, reducing the need for advanced AI intelligence in grading tasks.
WuYuEval is a multi-level benchmark for evaluating large language models in solid waste management, covering foundational knowledge, domain reasoning, and expert decision-making. It includes 4,590 multiple-choice and 247 scenario-based open-ended questions, and reports performance across 33 LLMs.