@rohanpaul_ai: "Long-horizon tasks are still a joke. They do not work, and I do not care what anybody says. Do not show me a stupid ev…
Summary
Chamath criticizes long-horizon tasks in AI as ineffective, predicts a hype cycle leading to disillusionment, and suggests using symbolic spaces to guide embedded spaces for better AI performance.
View Cached Full Text
Cached at: 08/29/26, 06:03 AM
“Long-horizon tasks are still a joke. They do not work, and I do not care what anybody says. Do not show me a stupid evaluation. Do not tell me about some dumb script you ran for 48 hours. Long-horizon tasks are not handled well. They simply do not work.”
- Chamath at Stanford AI Club
“2nd, complex problems also do not work. They are neither addressed nor handled well.
Why is this important? If AI develops like any other technology, we are going to experience an initial rise—the hype cycle. Then, we will see a natural contraction because, somehow and somewhere, something is going to fail. We are all going to see this, and then we will enter what is called the “trough of disillusionment.” I think the business and MBA folks will confirm whether that is true.
Afterward, you typically see the slow and gradual adoption of the real, final solution. This happened with the internet, and it has happened in many other cases.
The problem is that we are spending hundreds of billions, potentially trillions, of dollars trying to figure out how to cross this chasm. So, what do we do?
If we do not figure this out, people will reach the trough of disillusionment and say that AI was a joke. I think we need to be able to bring AI into highly complicated environments and make it work.
What is my solution? At a very basic level, you need a symbolic space that guides the embedded space.“
From “techniahqrobot” YouTube channel, (full video link in comment)
Similar Articles
@rohanpaul_ai: This paper is a brutal reality check for long-horizon AI. Give an agent a year of interconnected decisions, delayed fee…
A paper evaluates eight leading AI models on long-horizon tasks, finding that even the best-performing model achieves only 27.3% of human performance, highlighting significant limitations for dependable long-horizon AI execution.
@rohanpaul_ai: Long-horizon agent reliability has not arrived yet with better models. On WeaveBench's 114 hybrid GUI-CLI tasks, the be…
This paper argues that long-horizon AI agent reliability is lacking despite better models, as seen on WeaveBench with only 41.2% pass rate, and proposes the LongHorizon-Harness to manage task state for improved performance.
@jietang: Recent thoughts: The Shift to Long-Horizon Tasks The most likely breakthrough this year will be in long-horizon tasks. …
The article discusses the anticipated breakthrough in long-horizon AI tasks and autonomous agents, suggesting a shift from 'one-person' to 'none-person' companies. It highlights technical pillars like memory, continual learning, and self-judging as key to realizing fully self-evolving AI systems that could redefine AGI and operating systems.
@rohanpaul_ai: Chamath on how AI agents are making the "10x engineer" distinction disappear because the most efficient "code paths" ar…
Chamath Palihapitiya argues that AI agents are erasing the '10x engineer' distinction by making the most efficient code paths obvious to everyone, comparing it to how AI removed the mystery from optimal chess moves.
Dario and Dwarkesh: hard to watch as Dwarkesh seems so wrong it makes me cringe
Dario Amodei argues in an interview that AI's exponential growth phase is nearing its end, predicting high-confidence outcomes for verifiable tasks within years while addressing skepticism about scaling laws and RL.