jacobian-lens

Tag

Cards List
#jacobian-lens

Survival of the Fitted: Qwen3.6-27B’s Jacobian lens reads and steers Qwen3.8-27B with zero refitting [R]

Reddit r/MachineLearning · 2026-08-15

This article tests the transferability of a Jacobian interpretability lens from Qwen3.6-27B to Qwen3.8-27B, finding that it can read and steer the newer model with zero refitting for specific tasks.

0 favorites 0 likes
#jacobian-lens

Verbalizable Representations Form a Global Workspace in Language Models

arXiv cs.AI · 2026-07-20 Cached

This paper presents evidence that language models maintain a privileged set of internal representations (J-space) that are verbally reportable and function like a global workspace, introducing the Jacobian lens interpretability technique. The findings offer a window into model cognition and have implications for alignment and safety.

0 favorites 0 likes
#jacobian-lens

J-space comparisons across open models

Hacker News Top · 2026-07-15 Cached

This blog post extends Anthropic's verbalizable workspace paper by measuring how far steering directions from middle layers reach, when the structure forms during training, whether it transfers between models, and how it scales, all on open models.

0 favorites 0 likes
#jacobian-lens

J-Wash: A novel way to brainwash and customize large language models based on Anthropic's Jacobian-Lens!

Reddit r/LocalLLaMA · 2026-07-13 Cached

J-Wash is a tool that uses Anthropic's Jacobian lens to let users edit the behavior and identity of large language models by modifying token directions, then export the edited model as a standalone checkpoint without training or fine-tuning.

0 favorites 0 likes
#jacobian-lens

Evaluating J-space entropy as an error predictor across 7 datasets on Qwen3-4B [R]

Reddit r/MachineLearning · 2026-07-13

This study evaluates whether J-space entropy (inspired by Anthropic's Jacobian Lens) can serve as an error predictor across seven datasets on Qwen3-4B. Results show it can complement output confidence for factual retrieval but is not a general hallucination detector, with strong task dependence.

0 favorites 0 likes
#jacobian-lens

Interactive Jacobian-Lens visualizer and live steerer for GGUF models on llama.cpp

Reddit r/LocalLLaMA · 2026-07-12

An interactive Jacobian-lens visualizer and live steerer for GGUF models running on llama.cpp, enabling real-time model interpretability and control.

0 favorites 0 likes
#jacobian-lens

I created a super harmful model ! :D (by tweaking it's J-Space!!!)

Reddit r/LocalLLaMA · 2026-07-11

The author created a tool based on Anthropic's Jacobian-Lens to manually tweak a model's Jacobian Space, producing an uncensored model called Nikusui-v1, released with GGUF quantizations.

0 favorites 0 likes
#jacobian-lens

Anthropic found a hidden space where Claude puzzles over concepts

MIT Technology Review · 2026-07-09 Cached

Anthropic developed the Jacobian lens (J-lens) to reveal a hidden 'J-space' inside Claude Opus 4.6, offering unprecedented insight into an LLM's internal reasoning process before it outputs tokens. The technique allows monitoring and control of model behavior by surfacing the words the model is about to produce.

0 favorites 0 likes
#jacobian-lens

No Space Like J-Space (54 minute read)

TLDR AI · 2026-07-08 Cached

An article discussing a new interpretability technique called the Jacobian Lens and the discovery of J-space, a region in LLMs where verbalizable representations form a global workspace, marking a significant advance in understanding LLM reasoning.

0 favorites 0 likes
#jacobian-lens

@arafatkatze: Inspired by Anthropic's JSpace paper, we @cline used @modal to host a public demo so anyone can watch a Jacobian-lens v…

X AI KOLs Following · 2026-07-07 Cached

A public demo of a Jacobian-lens view of language model representations, inspired by Anthropic's JSpace paper, allowing users to explore model internals across layers and tokens.

0 favorites 0 likes
#jacobian-lens

I tested Anthropic’s new Jacobian Lens on open models, then it turned into a local-model hallucination router

Reddit r/LocalLLaMA · 2026-07-07

The author tested Anthropic's Jacobian Lens on open models, then it evolved into a local-model hallucination router for detecting AI hallucinations.

0 favorites 0 likes
#jacobian-lens

you can just watch a language model think now. i built a way to visualize the words AI doesn’t say

Reddit r/artificial · 2026-07-06

A developer built Subtext, a tool that visualizes the internal 'silent words' of language models using Anthropic's Jacobian lens, allowing real-time observation of the model's reasoning before it outputs tokens. The tool runs on a single 12GB GPU and streams at full generation speed.

0 favorites 0 likes
← Back to home

Submit Feedback