Tag
This paper introduces DiagFlowBench, a benchmark dataset of 1,676 multi-turn diagnostic conversations derived from industrial flowcharts, designed to evaluate how well language models handle off-procedure inputs and abstain from giving inappropriate advice.
This paper introduces DiagnosticIQ, a benchmark for evaluating LLMs in translating industrial symbolic maintenance rules into actionable steps. It highlights that while frontier models perform well on standard tasks, they exhibit brittleness and pattern-matching behaviors under structural perturbations.