open-ended-questions

Tag

Cards List
#open-ended-questions

Grading Needs a Rubric, Not Intelligence

arXiv cs.CL · 6d ago Cached

Small language models can grade open-ended exam answers as reliably as more expensive models when using an explicit rubric, reducing the need for advanced AI intelligence in grading tasks.

0 favorites 0 likes
#open-ended-questions

WuYuEval: A Multi-Level Benchmark for Large Language Models in Solid Waste Management

arXiv cs.CL · 2026-08-11 Cached

WuYuEval is a multi-level benchmark for evaluating large language models in solid waste management, covering foundational knowledge, domain reasoning, and expert decision-making. It includes 4,590 multiple-choice and 247 scenario-based open-ended questions, and reports performance across 33 LLMs.

0 favorites 0 likes
← Back to home

Submit Feedback