ablation

Tag

Cards List
#ablation

chessformer_lens demo: ablating 1 of a chess transformer's 128 attention heads makes the model stop finding Morphy's queen sacrifice [P]

Reddit r/MachineLearning · 2026-08-13

A demo of chessformer_lens shows that ablating a single attention head in a chess transformer causes it to stop recognizing Morphy's queen sacrifice, demonstrating the concentration of specific capabilities in individual heads.

0 favorites 0 likes
#ablation

Ablation, Statistical Inference, and Validation for KV-Cache Compression

arXiv cs.LG · 2026-07-14 Cached

This paper presents a systematic comparative study of KV-cache compression schemes (TurboQuant and SpectralQuant), introduces a statistical validation methodology, and offers regime-specific recommendations for efficient transformer inference.

0 favorites 0 likes
#ablation

Nex-N2-Mini-Ultra-Uncensored-Heretic Is Out Now, an Agentic Model With Agentic Thinking Now Uncensored With 5/100 Refusals and 0.0020 KLD, Available in Safetensors and GGUF Formats!

Reddit r/LocalLLaMA · 2026-06-24 Cached

A new uncensored version of the Nex-N2-mini model, called Nex-N2-mini-ultra-uncensored-heretic, has been released. It achieves 93% fewer refusals while preserving quality with low KL divergence, and is available in safetensors and GGUF formats.

0 favorites 0 likes
#ablation

New ablation operator. (apostate)

Reddit r/LocalLLaMA · 2026-06-22

A new contrastive ablation operator called apostate is introduced that reduces model refusal from 96% to 5% while preserving harmless behavior with only 0.081 KL divergence, tested on Granite 3.3-8B.

0 favorites 0 likes
#ablation

Contrastive targeted SFT as a mechinterp method - has anyone mapped causal dependency interactions this way? [D]

Reddit r/MachineLearning · 2026-06-17

A researcher shares an experimental plan for identifying causal dependencies between capability dimensions in a 31B model using contrastive targeted SFT and circuit tracing, seeking feedback on methodology and related work.

0 favorites 0 likes
#ablation

Tower-Plus-72B-Ultra-Uncensored-Heretic, a Model That Support 22 Languages Making it Great for Multilingual Tasks and is Especially Strong on Translation Related Workflows Where No Censorship Is Essential, Now Ultra Uncensored With 5/100 Refusals!

Reddit r/LocalLLaMA · 2026-06-15 Cached

Tower-Plus-72B-Ultra-Uncensored-Heretic is a decensored version of Unbabel/Tower-Plus-72B, supporting 22 languages and excelling in translation tasks with minimal refusals.

0 favorites 0 likes
#ablation

Ablation-Reversible Heads Don't Transfer: A Stress Test for Mechanistic Role Claims in Transformers

arXiv cs.AI · 2026-06-09 Cached

This paper shows that attention heads meeting common criteria for mechanistic role claims (necessity, linear decodability, ablation reversibility) routinely fail to transfer computations across prompts, and introduces the KID (Knowing/Intent/Doing) framework and a three-stage pipeline for more rigorous role assignment.

0 favorites 0 likes
#ablation

Recursive Self-Improvement for Skills (Skill RSI)

Reddit r/AI_Agents · 2026-06-03

Skill RSI is a free tool that recursively evaluates and improves AI skills via procedural evaluations and a research agent, supporting standalone or Codex plugin usage.

0 favorites 0 likes
#ablation

Why our #1 LightGBM feature by importance made predictions worse [D]

Reddit r/MachineLearning · 2026-06-01

A blog post from Flyback demonstrates how a LightGBM feature that ranked #1 in importance actually worsened predictions due to target encoding leakage, highlighting the danger of relying solely on feature importance metrics.

0 favorites 0 likes
#ablation

Measuring, Localizing, and Ablating Alignment Signatures in LLMs

arXiv cs.LG · 2026-06-01 Cached

This paper investigates how post-training of LLMs introduces AI-like stylistic regularities and proposes PASTA, a training-free method to localize and ablate these alignment signatures, reducing AI detection rates while maintaining coherence across 11 models and 6 detectors.

0 favorites 0 likes
#ablation

@NousResearch: To check that CNA isolates only the intended behavior, we evaluate steered models on MMLU across a range of steering st…

X AI KOLs Following · 2026-05-19 Cached

Nous Research released Contrastive Neuron Attribution (CNA), a method to steer LLM behavior by identifying and ablating sparse circuits in MLP neurons without training sparse autoencoders or degrading general benchmarks, validated on multiple large language models.

0 favorites 0 likes
← Back to home

Submit Feedback