Tag
ARIA is a test-time unlearning method for large language models that uses sparse autoencoders to suppress unwanted knowledge during inference without modifying weights, improving the forget-retain trade-off and remaining robust to adversarial attacks.
This paper proposes Bypass Observation, a non-intrusive architecture for extracting intermediate semantic information from large language models to improve interpretability and safety auditing.
This paper introduces a corpus of 10,000 annotated Bangla sentences for function classification and benchmarks various models, achieving 0.95 accuracy with a double-level ensemble using TF-IDF features.
This paper investigates how vision-language models read exact values from vertical bar charts using controlled counterfactual activation patching, revealing insights into internal computations and differences between models like Qwen and InternVL.
This paper evaluates general-purpose and medically fine-tuned LLMs on domain-specific jargon tasks, finding that the general-purpose model outperforms the specialist one and uses interpretability tools to analyze miscalibrations in the fine-tuned model.
SPICE introduces a generalizable framework for analyzing polysemanticity in neural networks, enabling systematic comparison across CNNs and Transformers and automatically determining concept clusters per neuron.
This article analyzes Dario Amodei's call for the AI industry to pace itself, examining five distinct camps' perspectives on what pacing means, including concerns about interpretability, labor rights, economic growth, geopolitical competition, and regulatory capture.
The paper proposes a distribution-aware method for identifying language-specific neurons in multilingual large language models by leveraging pairwise activation distribution overlaps, improving specificity in isolating causal effects across languages.
SearchAtlas is a framework that transforms LLM search agent trajectories into evidential query graphs to analyze search strategies, revealing process failures and improving interpretability beyond final-answer accuracy.
This paper investigates whether transformer models internalize the same linguistic features as traditional models for multilingual readability assessment, using SHAP and TCAV across five languages.
This paper presents a black-box, model-agnostic method to analyze temporal and geographic signals in language model embeddings using simple projections, finding that embeddings encode meaningful chronological and spatial structure for interpretability and retrieval tasks.
This podcast episode discusses AI tokenomics and the emerging issue of 'tokenflation', focusing on measuring the value of AI tokens, the limitations of benchmarks, and future innovations in AI efficiency.
This paper introduces SAEScientist-Bench, a benchmark to evaluate AI agents' ability to autonomously conduct mechanistic interpretability research using sparse autoencoders, revealing progress but significant gaps compared to expert baselines.
A post promoting the ELLIS PhD & Postdoc Program, detailing its benefits, tracks of study, and research areas such as AI interpretability, safety, and AI for Science.
CARDEA is an AI model designed to provide auditable reasoning for coronary angiography interpretation using Chain-of-Box to indicate image regions, enhancing trust in clinical practice.
John Schulman highlights research by Adam Karvonen and colleagues on using counterfactual simulatability as a metric to improve AI explanation quality. They developed a dataset and pipeline that trains models to generate better post-hoc explanations of their own behavior, showing generalization to held-out evaluations.
This paper introduces a taxonomy and dataset for contextual knowledge conflicts in large language models, experiments with seven LLMs, and proposes a steering method to improve conflict resolution in reasoning and summarization tasks.
This paper experimentally demonstrates that representational disentanglement in neural networks reduces collateral damage during unlearning, supporting long-held interpretability intuitions.
This paper investigates how large language models internally represent logical validity, showing that validity information is decodable from hidden states despite poor behavioral performance, suggesting distinct roles for representation, expression, and causal use.
This paper introduces Sparse Readout Prism (SRP), a method to analyze language model predictions by decomposing the readout matrix into sparse features, enabling corpus-independent explanations of logit-lens scores and improving interpretability.