Multimodal Model Diffing for Feature Discovery and Control
Summary
MMDiff uses multimodal sparse autoencoders to isolate, detect, and control features in multimodal language models, improving interpretability and targeted steering of visual and safety behaviors.
View Cached Full Text
Cached at: 08/17/26, 07:46 AM
Paper page - Multimodal Model Diffing for Feature Discovery and Control
Source: https://huggingface.co/papers/2608.09928
Abstract
MMDiff uses multimodal sparse autoencoders to isolate, detect, and control specific features in multimodal language models, improving interpretability and targeted steering of visual and safety behaviors.
Multimodal Large Language Models(MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control. While applicable to post-hoc inspection, hidden states that are decomposed into interpretable feature directions usingsparse autoencoders(SAEs) neither readily isolate which features are changed by multimodal training, nor are they directly useful for targeted control. We introduceMMDiff, a multimodal model-diffing framework that trainsmultimodal SAEsand turns them into feature-level interfaces for discovering and controlling multimodal behavior.MMDiffsupports three uses: (i)feature isolation, by diffing a base-LM SAE against its multimodal-adapted counterpart to identify features altered by multimodal training; (ii) task-specific feature detection, via per-tokencontrastive firing analysisthat isolates causal features; and (iii)feature-level control, by causally removing orsteeringthe discovered feature directions. We trainmultimodal SAEsfor three MLLM families, LLaVA-MORE, PaliGemma 2, and InternVL3.5, and evaluate onvisual-spatial understanding,multimodal safety, and OCR.MMDiffdiscovers sparse, causally specific features whose removal selectively degrades target behaviors by an average of 12% on spatial tasks and 17% on OCR, and reduces attack success rate by 24% onmultimodal safetyattacks, with no impact on VQA performance.Steeringthese features improves spatial and OCR accuracy by +3.6% and +1.8% on average over a standard single-layersteeringbaseline. These results show thatmultimodal SAEscan serve not only as interpretability tools, but as mechanisms for auditing,steering, and controlling MLLMs behavior toward safer and more capable generations.
View arXiv pageView PDFProject pageGitHub0Add to collection
Get this paper in your agent:
hf papers read 2608\.09928
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.09928 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.09928 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.09928 in a Space README.md to link it from this page.
Collections including this paper2
Similar Articles
MMDiff: Extending Diffusion Transformers for Multi-Modal Generation
MMDiff extends frozen diffusion transformers into multi-modal generative systems using lightweight decoders, achieving significant improvements in semantic segmentation and other perceptual tasks through multi-timestep feature fusion.
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models
PerceptionDLM introduces a multimodal diffusion language model that enables parallel region perception via structured attention masking and efficient prompting, achieving faster inference without sacrificing caption quality. Experiments show competitive performance with substantial speed improvements for multi-region perception tasks.
CL-DMDF:Dynamic Multimodal Data Fusion Model Based on Contrastive Learning
This paper proposes CL-DMDF, a dynamic multimodal data fusion model that uses contrastive learning and a dual-dimensional attention mechanism to handle missing modalities and improve discriminative learning.
MBDiff: Multi-view Behavior-aware Diffusion Model for Probabilistic Utility Data Imputation
Presents MBDiff, a multi-view behavior-aware diffusion model for probabilistic utility data imputation that learns user behavior from global, local, and instance-level views and uses a conditional attentional denoising network. Evaluated on real utility data from Florida, it outperforms state-of-the-art baselines.
Faster but Different: Diagnosing and Controlling Content Drift in Accelerated Multimodal Diffusion Language Models
This paper investigates how training-free acceleration can silently change generated content in diffusion-based multimodal large language models, and proposes paired diagnostics and consistency-control methods to mitigate content drift.