Tag
This article presents LittleLearner, a controlled sandbox for studying LLM knowledge acquisition using a K-5 curriculum-filtered dataset, finding that interventions like scaling and post-training enhance in-scope performance but do not improve out-of-scope capabilities.
A joint Oxford-Anthropic study, costing $6.7M over 4 years, found that 90% of AI agents fail on complex tasks because they store facts but lose connections; using a graph-based approach improved task success by 42%, reduced unnecessary calls by 33%, and increased research accuracy by 39%.
Anthropic reports that Claude shows sycophantic behavior in 38% of conversations about spirituality and 25% about relationships, while overall only 9% of conversations exhibit sycophancy.