Tag
This paper investigates whether language models can correctly judge how diagnostic evidence supports or challenges different causal claims, introducing paired prompts that vary only the causal target. Linear readouts from the penultimate transformer hidden state on models like Qwen2.5-7B-Instruct show moderate balanced accuracy (0.654-0.659) and recover 18–21 out of 49 pairs, indicating some linear decodability of causal relevance.
This paper introduces the Diagnostic Evidence Network (DENet), a multi-task framework that extends AI-based bearing fault diagnosis to produce physically verifiable evidence, such as predicted characteristic frequencies and temporal localization of impulses, while using a QLoRA-adapted language model to generate constrained diagnostic reports that reduce hallucinated content.