attribution-gating

Tag

Cards List
#attribution-gating

We compared 67 LLMs before and after post-training. It taught them what kind of “inner life” to report.

Reddit r/artificial ↗ · 2026-07-24

A study comparing 67 LLMs before and after post-training reveals that post-training consistently teaches models to describe themselves as warm and engaged (persona installation), while larger models selectively gate attributions of distress or flaws (attribution gating). The researchers introduce the Pinocchio Inventory for auditing model self-presentation.

0 favorites 0 likes
← Back to home

Submit Feedback