feature-attribution

Tag

Cards List
#feature-attribution

NObSP: Functional Decomposition of Neural Networks via Oblique Subspace Projections

arXiv cs.LG · yesterday Cached

This paper introduces NObSP, a framework for decomposing neural network predictions into per-feature contribution functions and interaction residuals using oblique subspace projections, improving interpretability and reducing attribution errors compared to existing methods.

0 favorites 0 likes
#feature-attribution

RelShap: Relationally Consistent Shapley Explanations

arXiv cs.LG · 2026-08-13 Cached

This paper proposes RelShap, a framework that incorporates relational constraints and data provenance into Shapley value computation, making explanations more faithful to the data-generating process. It is estimator-agnostic and composes with existing SHAP estimators while exploiting functional dependencies to reduce runtime.

0 favorites 0 likes
#feature-attribution

Measuring Explainer Stability via Attribution Separability

arXiv cs.LG · 2026-08-05 Cached

The paper introduces a distribution-based framework to measure the stability of attribution methods (explainers) by quantifying the separability of feature rankings and identifying the maximum top-k ranking that remains reliable across stochastic runs.

0 favorites 0 likes
#feature-attribution

Shapley-Value-Based Feature Attribution for Data Masking

arXiv cs.LG · 2026-08-03 Cached

Proposes a Shapley-value-based feature attribution framework for data masking that balances disclosure risk and data utility at the feature level, agnostic to specific masking methods.

0 favorites 0 likes
#feature-attribution

FADEx: Feature Attribution and Distortion-based Explanation of Dimensionality Reduction

arXiv cs.LG · 2026-07-31 Cached

FADEx introduces a novel local per-instance feature attribution method for explaining dimensionality reduction techniques, using Taylor expansion and Singular Value Decomposition to provide model-agnostic explanations and distortion analysis.

0 favorites 0 likes
#feature-attribution

Feature Attribution in Directed Acyclic Graphs Using Edge Intervention

arXiv cs.AI · 2026-06-16 Cached

Proposes DAG-SHAP, a novel feature attribution method based on edge intervention for directed acyclic graphs, addressing limitations of existing Shapley value methods in capturing feature interactions and causal relationships.

0 favorites 0 likes
#feature-attribution

The Attribution Contract: Feature Attribution for Generative Language Models

arXiv cs.LG · 2026-05-25 Cached

This paper introduces the Attribution Contract, a specification for feature-attribution claims in generative language models, addressing ambiguities in what constitutes a feature and how attribution methods should be evaluated. It uses autoregressive and diffusion models as case studies to show when attribution is informative or misleading.

0 favorites 0 likes
#feature-attribution

The Attribution Impossibility: No Feature Ranking Is Faithful, Stable, and Complete Under Collinearity

arXiv cs.LG · 2026-05-22 Cached

This paper proves that no feature ranking can be simultaneously faithful, stable, and complete under collinearity, characterizing the full attribution design space and providing a formally verified impossibility theorem in explainable AI.

0 favorites 0 likes
#feature-attribution

model-agnostic sensitivity approximator [P]

Reddit r/MachineLearning · 2026-05-18

A 16-year-old developer created sage-explainer, a Python package that approximates prediction sensitivity to features for black-box models like random forests and XGBoost, offering more stable results than centered finite differences.

0 favorites 0 likes
#feature-attribution

From Weight Perturbation to Feature Attribution for Explaining Fully Connected Neural Networks

arXiv cs.LG · 2026-05-18 Cached

Introduces a weight perturbation-based feature attribution method (XWP and XWPc) for fully connected neural networks, achieving competitive performance on standard baseline metrics.

0 favorites 0 likes
#feature-attribution

Prune, Interpret, Evaluate: A Cross-Layer Transcoder-Native Framework for Efficient Circuit Discovery via Feature Attribution

arXiv cs.CL · 2026-04-21 Cached

Researchers introduce PIE, a CLT-native framework for efficient circuit discovery via feature attribution-based pruning, achieving ~40× compression in feature selection while maintaining behavioral fidelity on IOI and Doc-String tasks.

0 favorites 0 likes
← Back to home

Submit Feedback