Tag
This paper introduces a Bloom-aligned framework for measuring educational control in LLMs—the ability to adjust cognitive demand while preserving instructional intent—and evaluates it on programming tasks using Qwen3 models, revealing that models excel at increasing difficulty but struggle to lower it.
This paper evaluates cross-dataset generalization of supervised ML/DL models and prompted LLMs for automatic Bloom's taxonomy classification of assessment questions, finding that LLMs are more robust across diverse educational contexts.
BloomBench is a cognitively grounded bilingual (English-Arabic) multimodal benchmark for Vision-Language Models, systematically evaluating six cognitive levels based on Bloom's Taxonomy. Experiments reveal significant cognitive asymmetries and cross-lingual performance gaps in current models.