@HowToAI_: Microsoft Research + Salesforce has published a paper that should scare every single AI builder right now. It’s called …

X AI KOLs Timeline Papers

Summary

A new paper by Microsoft Research and Salesforce reveals that LLM performance drops significantly in multi-turn conversations due to a 'Lost in Conversation' phenomenon, challenging the reliability of current single-turn benchmarks.

Microsoft Research + Salesforce has published a paper that should scare every single AI builder right now. It’s called “LLMs Get Lost in Multi-Turn Conversation.” And it reveals a massive, hidden performance cliff that every major model is currently falling off. We’ve been measuring AI intelligence all wrong. We test models on single, perfectly crafted prompts. But in the real world, we use AI in conversations. We go back and forth. We clarify. We refine. The researchers just proved that even the "smartest" models like GPT, Claude, and Gemini, are actually terrible at this. The finding is brutal: When an AI moves from a single prompt to a multi-turn conversation, its performance doesn't just dip. It craters by an average of 35% to 39%. The researchers tested 15 different models across 200,000 simulated conversations. The results were universal. As the conversation continues, the AI doesn't get smarter. It gets lost. The paper identifies the "Lost in Conversation" (LiC) phenomenon, and it happens for four specific reasons: 1. Premature Solutions: The AI tries to solve the problem before it has all the info. 2. False Assumptions: It fills in the gaps of what you didn't say with total hallucinations. 3. Over-reliance on Errors: Once it makes a mistake in turn 2, it clings to that mistake for the rest of the chat. 4. Context Bloat: It gets distracted by its own previous wordy responses. In simpler terms: when an LLM takes a wrong turn, it doesn't recover. It just keeps driving deeper into the woods. It’s a crisis for anyone building autonomous agents. If your agent needs to ask "What did you mean by that?" or "Which file should I use?", its chances of successfully finishing the task drop from 90% to as low as 35%. The more the AI talks, the less reliable it becomes. We have spent three years building "conversational AI," only to find out that the conversation itself is what's breaking the AI. The takeaway is a wake-up call for the entire industry: Stop trusting single-turn benchmarks. They are measuring a version of the AI that doesn't exist in production.
Original Article

Similar Articles

Microsoft is reportedly training salespeople to talk down OpenAI and Anthropic

TechCrunch AI

Microsoft is reportedly training its sales team to negatively compare rival AI products from OpenAI, Anthropic, and Google against its own, emphasizing the full end-to-end system. The move comes after Microsoft dropped exclusivity with OpenAI and is swapping external models for its own in flagship apps.