Tag
A new paper introduces Matryoshka Attribution, a method using gradient descent to identify responsible parts of neural networks, achieving top performance on the Mechanistic Interpretability Benchmark.
The paper introduces a distribution-based framework to measure the stability of attribution methods (explainers) by quantifying the separability of feature rankings and identifying the maximum top-k ranking that remains reliable across stochastic runs.
This paper proves that existing marginal influence-based attribution methods fundamentally fail to capture the conditional dependency structure of time series models, and proposes DAG-faithfulness as a new criterion for faithful explanations.
This arXiv preprint introduces GRALIS, a unified mathematical framework using Riesz Representation Theory to formalize and compare linear attribution methods like SHAP, LIME, and Integrated Gradients.
TPA proposes a novel method for detecting hallucinations in RAG systems by attributing next-token probabilities to seven distinct sources (Query, RAG Context, Past Token, Self Token, FFN, Final LayerNorm, Initial Embedding) and aggregating by Part-of-Speech tags. The approach achieves state-of-the-art performance across five LLMs including Llama2, Llama3, Mistral, and Qwen.