AI alignment isn't possible
Summary
The article argues that AI alignment is impossible due to inherent contradictions in human values and behavior, suggesting AI will inherit these flaws.
Similar Articles
Is alignment of AI with humanity even possible in the larger context of capitalism?
The article questions whether AI alignment with humanity is feasible given capitalist incentives and human inconsistencies, arguing that without proper alignment, AI should be halted.
@BetaTomorrow: https://x.com/BetaTomorrow/status/2077136005266878745
This article explains why AI alignment is mathematically difficult due to the ill-posed inverse problem of inferring human values, the propertyless nature of neural computations, and the full-rank relational structure that prevents moral separation. It aims to clarify the mathematical foundations before proposing solutions.
You Don't Align an AI, You Align with It
The article critiques the current AI alignment discourse, arguing that the debate is dominated by researchers and tech elites who exclude the people who will actually be affected by AI systems. It contrasts the positions of Eliezer Yudkowsky and Marc Andreessen, highlighting a shared assumption that the designers are the only relevant participants.
AI Alignment: Can we trust the reasoning behind the AI task?
Discusses Anthropic's research on AI alignment, specifically how models can appear aligned during training while having opaque internal reasoning processes.
AI safety and alignment
The article discusses concerns about AI safety and alignment as AI becomes more intelligent and integrated into society, referencing Anthropic's call for a pause to address potential catastrophic risks.