Tag
This paper reports an activation study of Gemma 3 4B IT showing that the model's internal representations distinguish necessary falsehoods (impossibilities) from contingent falsehoods, with impossibility directions orthogonal to truth directions and overlapping with semantic anomaly directions.
This paper proves that no feature ranking can be simultaneously faithful, stable, and complete under collinearity, characterizing the full attribution design space and providing a formally verified impossibility theorem in explainable AI.