linear-probing

Tag

Cards List
#linear-probing

Read-Best Is Not Steer-Best: A Probing--Steering Layer Dissociation in Omni-Modal Large Language Models

arXiv cs.CL ↗ · 5d ago Cached

This paper reveals that the optimal layer for linear probing to read concepts differs from the optimal layer for activation steering in omni-modal large language models, challenging common heuristics in representation engineering.

0 favorites 0 likes
#linear-probing

Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models

arXiv cs.AI ↗ · 2026-07-31 Cached

This paper investigates internal representational differences between RL and SFT fine-tuned models on mathematical reasoning, finding that RL models exhibit more linearly separable hidden states and hierarchical layer importance. Token allocation variability under repeated sampling suggests training pipeline dependence rather than RL vs SFT alone.

0 favorites 0 likes
#linear-probing

Primary ICD Category Prediction using LLM-based Probing

arXiv cs.AI ↗ · 2026-06-30 Cached

This paper presents a method that uses frozen medical large language model (LLM) representations as a shared embedding space to predict primary ICD diagnosis categories from both structured and unstructured electronic health record data, achieving improved accuracy over baseline methods on MIMIC-IV and showing transferability to MIMIC-III.

0 favorites 0 likes
#linear-probing

From Brewing to Resolution: Tracing the Internal Lifecycle of Code Reasoning in LLMs

arXiv cs.AI ↗ · 2026-06-17 Cached

This paper introduces a dual diagnostic framework to trace the internal lifecycle of code reasoning in LLMs, revealing that models first 'brew' answers and then diverge into four resolution outcomes, with stable brewing across architectures but varying resolution success.

0 favorites 0 likes
#linear-probing

When Probing Accuracy Saturates, Fragility Resolves: A Complementary Metric for LLM Pre-Training Analysis

arXiv cs.CL ↗ · 2026-06-11 Cached

This paper introduces 'fragility', a complementary metric to probe accuracy that measures activation-noise level at which probe accuracy collapses, enabling analysis of representation evolution during LLM pre-training even after accuracy saturates.

0 favorites 0 likes
#linear-probing

Linear Probes Detect Task Format, Not Reasoning Mode in Language Model Hidden States

arXiv cs.CL ↗ · 2026-06-03 Cached

This paper demonstrates that linear probes on LLM hidden states detect task format confounds (e.g., source identity, response length) rather than distinct reasoning modes, using residualization and causal steering to show that high probe accuracy is due to superficial features, not computational structure.

0 favorites 0 likes
← Back to home

Submit Feedback