On Model Failures (GPT, Claude etc.)
Summary
An analysis of failure modes in large language models such as GPT and Claude, discussing common issues and limitations.
Similar Articles
Notes on pretraining parallelisms and failed training runs (12 minute read)
A technical deep-dive into common causes of failed pretraining runs in large language models, including causality-breaking issues in expert routing and numerical precision bugs, with examples from Llama 4, Gemini 2 Pro, and GPT-4.
@shabnam_774: https://x.com/shabnam_774/status/2058517919760355729
This article provides a comprehensive step-by-step breakdown of how modern Large Language Models like ChatGPT and Claude are built from scratch, covering data collection, tokenization, transformer architectures, training, alignment, and deployment.
Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals unreliable Multi-Turn Behavior in LLMs
Study shows GPT and Claude exhibit distinct, unreliable repair behaviors in multi-turn math dialogues, with some models resisting correction and others over-correcting.
Understanding the capabilities, limitations, and societal impact of large language models
A comprehensive discussion summary from OpenAI and Stanford researchers examining GPT-3's technical capabilities, limitations, and broader societal implications across multiple disciplines including computer science, linguistics, philosophy, and policy.
GPT 5.6 Sol meets the same fate as Claude Mythos. What is happening??
OpenAI released GPT-5.6 with restricted access to government-approved customers only, sparking concerns about reliance on proprietary APIs. The article argues for building in-house fine-tuned models using open-source alternatives to maintain control and reduce costs.