model-editing

Tag

Cards List
#model-editing

Anthropic published research on GRAM: a technique to surgically remove dangerous knowledge from AI models at the weight level

Reddit r/artificial · 2026-07-09

Anthropic published research on GRAM, a technique for surgically removing dangerous knowledge from AI models at the weight level, advancing AI safety.

0 favorites 0 likes
#model-editing

Revocable Learned State via Process Sidecars

arXiv cs.LG · 2026-07-01 Cached

This paper introduces process sidecars, a two-coefficient edit family for revoking learned state from language models after safety training, achieving second-order accuracy and outperforming naive task arithmetic in experiments across multiple models.

0 favorites 0 likes
#model-editing

BrainSurgery: Reproducible and Reliable Declarative Weight Manipulations for Model Editing and Upcycling

Hugging Face Daily Papers · 2026-06-08 Cached

BrainSurgery is a tool for reproducible and declarative weight manipulations on neural network checkpoints, enabling model editing and upcycling through YAML plans with built-in validation.

0 favorites 0 likes
#model-editing

@jiqizhixin: New from NVIDIA! You can edit a model’s compressed memory without scrambling what it already knows! Enter Gated DeltaNe…

X AI KOLs Timeline · 2026-05-22 Cached

NVIDIA introduces Gated DeltaNet-2, a method for editing compressed model memory without catastrophic forgetting, using independent gates for erase and write operations. It outperforms existing models like Mamba-2 and Mamba-3 on language modeling and long-context tasks.

0 favorites 0 likes
#model-editing

Modality-Decoupled Online Recursive Editing

arXiv cs.LG · 2026-05-21 Cached

Proposes M-ORE, a modality-decoupled online recursive editor for lifelong adaptation of multimodal large language models, addressing cross-modal conflict and inter-edit interference with constant per-edit overhead.

0 favorites 0 likes
#model-editing

HoReN: Normalized Hopfield Retrieval for Large-Scale Sequential Model Editing

arXiv cs.LG · 2026-05-12 Cached

This paper introduces HoReN, a parameter-preserving model editing method that uses normalized Hopfield retrieval to handle large-scale sequential updates to large language models. It addresses issues of knowledge accumulation and routing challenges, demonstrating stable performance on 50K sequential edits where prior methods degrade.

0 favorites 0 likes
← Back to home

Submit Feedback