Tag
This paper presents a modular pipeline for educational analogy generation, decomposing the task into four stages and evaluating 12 LLMs and 7 embedding models. Results show that sub-concept grounding improves explanation quality and retrieval precision, with a novel LLM-as-a-judge evaluation validated against human annotations.