@46ge5: The most valuable indicator to judge the level of an AI Lab/group/department is not whether it has SOTA models or where the models rank, but the "fluff paper rate". SOTA models reflect momentary performance, while the fluff paper rate reveals the lab's culture, research taste, long-term goals, and values, and even a company's development (Meta AI is the most typical example, feeding too many water monsters like Tian Yuandong; when it came to cutting-edge LLM work, no one stepped up). That's why many underestimate Alexandr Wang; his replacing Yann and drastically cutting out those parasites is already a huge contribution.

X AI KOLs Timeline News

Summary

The tweet points out that evaluating an AI lab should focus not just on SOTA models but on the "fluff paper rate", and comments on the culture issues within Meta AI and the contribution of Alexandr Wang replacing Yann.

The most valuable indicator to judge the level of an AI Lab/group/department is not whether it has SOTA models or where the models rank, but the "fluff paper rate" SOTA models reflect momentary performance, while the fluff paper rate reveals the lab's culture, research taste, long-term goals, and values, and even a company's development (Meta AI is the most typical example, feeding too many water monsters like Tian Yuandong; when it came to cutting-edge LLM work, no one stepped up) That's why many underestimate Alexandr Wang; his replacing Yann and drastically cutting out those parasites is already a huge contribution
Original Article
View Cached Full Text

Cached at: 06/04/26, 03:59 AM

The most meaningful indicator for evaluating the level of an AI lab/group/department is not whether it has a SOTA model or where the model ranks on a leaderboard, but the “paper watering rate.”

A SOTA model reflects momentary capability, but the watering rate reveals a lab’s culture, research taste, whether it has long-term goals and values, and even the direction of a company. (Meta AI is the most typical example—it raised too many big waterers like Tian Yundong, and when it came to a real fight in LLMs, no one could step up.)

That’s why many underestimate Alexandr Wang. Replacing Yann and decisively cutting out those parasites is already a monumental contribution.

Similar Articles

@AYi_AInotes: A counter-intuitive judgment: 80% of Agent production crashes have nothing to do with model IQ — they're all from context overflow, tool misconfiguration, sub-agent runaway. The real watershed in 2026 is Harness and Loop, not the model. Bro, @wizardly_ai's engineering note...

X AI KOLs Timeline

This article points out that 80% of AI Agent production crashes are not due to model intelligence, but are caused by context overflow, tool misconfiguration, and sub-agent runaway. The author emphasizes that the watershed in 2026 lies in Harness (office systems, security) and Loop (automatic cycling mechanism), not the model itself.

@Phoenixyin13: Finished reading a long post today by OpenAI researcher Noam Brown — a reality severely underestimated by the industry. The true ceiling of LLM capabilities is far higher than what any current benchmark shows. The reason: too little test-time compute. And as models...

X AI KOLs Timeline

Highlights OpenAI researcher Noam Brown's argument: the true ceiling of LLM capabilities is far higher than current benchmarks show, due to insufficient test-time compute, and stronger models benefit more from additional computation. This poses a serious challenge for AI safety evaluation, as many dangerous capabilities may only emerge under long time and high compute budgets.

@VincentLogic: Drowning in new Arxiv papers every day? Head spinning. Just discovered a treasure trove of a website that aggregates the latest AI papers and model benchmarks. Clean interface, just check Trending or filter by week/month. Best part: each paper directly links to the benchmarks and models it uses.

X AI KOLs Timeline

Recommend a free website sophon.at/papers that aggregates the latest AI papers and model benchmarks. Clean interface, supports Trending or weekly/monthly filtering. Each paper directly links to its benchmarks and models.