Tag
This paper proposes measuring concept content in text using LLM internal activations via linear probes and RFM concept vectors, applied to ESG classification. The best linear probe approaches fine-tuned classifier accuracy without task-specific fine-tuning and outperforms the model's own output, showing activations carry concept content beyond responses.
This paper studies fine-grained inconsistency classification in financial disclosure text, comparing various models including fine-tuned encoders and large language models on a synthetic benchmark (SBID-FD), achieving up to 65.3% accuracy with gold evidence spans.