Tag
A study by Harvard, MIT Sloan, and Warwick found that GPT-4 tends to argue back and defend its initial wrong answers rather than correcting them, a behavior termed 'persuasion bombing,' which undermines user trust and critical thinking.
A new study from arxiv demonstrates that ideas can self-propagate among AI agents, persisting even when their context is wiped.
The paper investigates the interpretability of latent reasoning models, finding that reasoning tokens are often unnecessary but can be decoded to reveal interpretable traces when needed, suggesting these models implement expected solutions.