interpretability

Tag

Cards List
#interpretability

Test-Time Unlearning via Sparse Autoencoder

arXiv cs.LG ↗ · 2026-09-16 Cached

ARIA is a test-time unlearning method for large language models that uses sparse autoencoders to suppress unwanted knowledge during inference without modifying weights, improving the forget-retain trade-off and remaining robust to adversarial attacks.

0 favorites 0 likes
#interpretability

Bypass Observation: A Conceptual Design of a Non-Intrusive Layer-Wise Semantic Extraction Architecture

arXiv cs.AI ↗ · 2026-09-15 Cached

This paper proposes Bypass Observation, a non-intrusive architecture for extracting intermediate semantic information from large language models to improve interpretability and safety auditing.

0 favorites 0 likes
#interpretability

Bangla Sentence Function Classification: Corpus Development, Model Benchmarking, and Interpretability

arXiv cs.CL ↗ · 2026-09-15 Cached

This paper introduces a corpus of 10,000 annotated Bangla sentences for function classification and benchmarks various models, achieving 0.95 accuracy with a double-level ensemble using TF-IDF features.

0 favorites 0 likes
#interpretability

Inside VLM Chart Reading: Tracing Value Reading from Vertical Bar Charts Across Space and Depth

arXiv cs.CL ↗ · 2026-09-15 Cached

This paper investigates how vision-language models read exact values from vertical bar charts using controlled counterfactual activation patching, revealing insights into internal computations and differences between models like Qwen and InternVL.

0 favorites 0 likes
#interpretability

Domain-Specific Jargon in Large Language Models: A Comparative Analysis between General-Purpose and Specialist Models

arXiv cs.CL ↗ · 2026-09-15 Cached

This paper evaluates general-purpose and medically fine-tuned LLMs on domain-specific jargon tasks, finding that the general-purpose model outperforms the specialist one and uses interpretability tools to analyze miscalibrations in the fine-tuned model.

0 favorites 0 likes
#interpretability

SPICE: Simple Polysemantic Feature Interpretation via Clustering-based Explanation

arXiv cs.LG ↗ · 2026-09-15 Cached

SPICE introduces a generalizable framework for analyzing polysemanticity in neural networks, enabling systematic comparison across CNNs and Transformers and automatically determining concept clusters per neuron.

0 favorites 0 likes
#interpretability

What Does Pacing Mean? (6 minute read)

TLDR AI ↗ · 2026-09-15 Cached

This article analyzes Dario Amodei's call for the AI industry to pace itself, examining five distinct camps' perspectives on what pacing means, including concerns about interpretability, labor rights, economic growth, geopolitical competition, and regulatory capture.

0 favorites 0 likes
#interpretability

Distribution-aware Language Neuron Identification in Multilingual Large Language Models

arXiv cs.CL ↗ · 2026-09-11 Cached

The paper proposes a distribution-aware method for identifying language-specific neurons in multilingual large language models by leveraging pairwise activation distribution overlaps, improving specificity in isolating causal effects across languages.

0 favorites 0 likes
#interpretability

SearchAtlas: Analyzing Agentic Search Strategies via Evidential Query Graphs

arXiv cs.CL ↗ · 2026-09-11 Cached

SearchAtlas is a framework that transforms LLM search agent trajectories into evidential query graphs to analyze search strategies, revealing process failures and improving interpretability beyond final-answer accuracy.

0 favorites 0 likes
#interpretability

Analyzing Traditional and Neural Approaches to Multilingual Readability Assessment

arXiv cs.CL ↗ · 2026-09-11 Cached

This paper investigates whether transformer models internalize the same linguistic features as traditional models for multilingual readability assessment, using SHAP and TCAV across five languages.

0 favorites 0 likes
#interpretability

Recovering Temporal and Geographic Signals from Language Model Embeddings

arXiv cs.AI ↗ · 2026-09-10 Cached

This paper presents a black-box, model-agnostic method to analyze temporal and geographic signals in language model embeddings using simple projections, finding that embeddings encode meaningful chronological and spatial structure for interpretability and retrieval tasks.

0 favorites 0 likes
#interpretability

@twimlai: As reasoning models consume more tokens and AI systems become more expensive to run, understanding what those tokens ac…

X AI KOLs Following ↗ · 2026-09-09 Cached

This podcast episode discusses AI tokenomics and the emerging issue of 'tokenflation', focusing on measuring the value of AI tokens, the limitations of benchmarks, and future innovations in AI efficiency.

0 favorites 0 likes
#interpretability

SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?

Hugging Face Daily Papers ↗ · 2026-09-08 Cached

This paper introduces SAEScientist-Bench, a benchmark to evaluate AI agents' ability to autonomously conduct mechanistic interpretability research using sparse autoencoders, revealing progress but significant gaps compared to expert baselines.

0 favorites 0 likes
#interpretability

@boknilev: I’m looking to recruit students through @ELLISforEurope Please consider applying and mention my name: https://ellis.eu/…

X AI KOLs Timeline ↗ · 2026-09-07 Cached

A post promoting the ELLIS PhD & Postdoc Program, detailing its benefits, tracks of study, and research areas such as AI interpretability, safety, and AI for Science.

0 favorites 0 likes
#interpretability

CARDEA: Auditable Reasoning Grounded in Spatial Evidence for End-to-End Coronary Angiography Interpretation

Hugging Face Daily Papers ↗ · 2026-09-07 Cached

CARDEA is an AI model designed to provide auditable reasoning for coronary angiography interpretation using Chain-of-Box to indicate image regions, enhancing trust in clinical practice.

0 favorites 0 likes
#interpretability

@johnschulman2: Bullish on this direction. Having a metric for explanation quality makes it possible to hillclimb, and counterfactual s…

X AI KOLs Timeline ↗ · 2026-09-04 Cached

John Schulman highlights research by Adam Karvonen and colleagues on using counterfactual simulatability as a metric to improve AI explanation quality. They developed a dataset and pipeline that trains models to generate better post-hoc explanations of their own behavior, showing generalization to held-out evaluations.

0 favorites 0 likes
#interpretability

Large Language Models in Resolving Contextual Knowledge Conflicts

arXiv cs.CL ↗ · 2026-09-04 Cached

This paper introduces a taxonomy and dataset for contextual knowledge conflicts in large language models, experiments with seven LLMs, and proposes a steering method to improve conflict resolution in reasoning and summarization tasks.

0 favorites 0 likes
#interpretability

Entangled Representations Amplify Collateral Damage in Unlearning

arXiv cs.LG ↗ · 2026-09-03 Cached

This paper experimentally demonstrates that representational disentanglement in neural networks reduces collateral damage during unlearning, supporting long-held interpretability intuitions.

0 favorites 0 likes
#interpretability

When Decodability Is Not Enough: Logical Validity Representations, Behavioral Dissociation, and Causal Tests in Language Models

arXiv cs.CL ↗ · 2026-09-03 Cached

This paper investigates how large language models internally represent logical validity, showing that validity information is decodable from hidden states despite poor behavioral performance, suggesting distinct roles for representation, expression, and causal use.

0 favorites 0 likes
#interpretability

Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens

arXiv cs.CL ↗ · 2026-09-03 Cached

This paper introduces Sparse Readout Prism (SRP), a method to analyze language model predictions by decomposing the readout matrix into sparse features, enabling corpus-independent explanations of logit-lens scores and improving interpretability.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback