Tag
This paper introduces DiagnosticIQ, a benchmark for evaluating LLMs in translating industrial symbolic maintenance rules into actionable steps. It highlights that while frontier models perform well on standard tasks, they exhibit brittleness and pattern-matching behaviors under structural perturbations.