self-recognition

Tag

Cards List
#self-recognition

Self-Recognition Finetuning can Prevent and Reverse Emergent Misalignment

arXiv cs.CL · 2026-06-24 Cached

This paper proposes Self-Recognition Finetuning as an intervention to prevent and reverse emergent misalignment in LLMs, showing it stabilizes the model's aligned character rather than adopting a misaligned persona.

0 favorites 0 likes
#self-recognition

Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All?

arXiv cs.AI · 2026-05-12 Cached

This research introduces a 3D benchmark to evaluate whether Vision-Language Model (VLM) agents can achieve mirror self-recognition, a proxy for higher-order cognition. The study finds that while stronger VLMs can use reflected evidence for action, weaker models often fail to extract self-relevant information or misattribute reflections, highlighting the distinction between linguistic compliance and grounded self-identification.

0 favorites 0 likes
← Back to home

Submit Feedback