activation-probing

Tag

Cards List
#activation-probing

Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models

Hugging Face Daily Papers ↗ · 2026-09-01 Cached

Introduces Autoregressive Mosaics (AM-Bench), a benchmark to evaluate whether text-only LLMs have genuine 2D spatial reasoning abilities, distinct from code generation. Findings show spatial reasoning varies among models and is influenced by output medium like SVG vs. code.

0 favorites 0 likes
#activation-probing

Committed Before Reasoning: Behavioral Reproduction and Preliminary Activation-Level Evidence of Answer Pre-Commitment in an Open-Weight LLM

arXiv cs.CL ↗ · 2026-07-21 Cached

This paper reproduces the phenomenon of answer pre-commitment in an open-weight LLM (Qwen3-8B) using a minimal car-wash question and provides preliminary activation-level evidence that the commitment is encoded in hidden states before the answer text is emitted.

0 favorites 0 likes
← Back to home

Submit Feedback