Chamath Palihapitiya told Stanford AI Club that long-horizon tasks 'simply do not work' and dismissed benchmark evaluations as 'stupid.' He warned that without progress, AI faces a trough of disillusionment with hundreds of billions at stake.
Key facts
- Chamath Palihapitiya: long-horizon tasks 'simply do not work'.
- Warns of 'trough of disillusionment' for AI industry.
- Spending 'hundreds of billions, potentially trillions' on AI.
- Proposes symbolic space guiding embedded space.
- Talk at Stanford AI Club, video on techniahqrobot channel.
Chamath Palihapitiya, the venture capitalist and former Facebook executive, delivered a blunt assessment of AI's limitations at a Stanford AI Club event. According to @rohanpaul_ai, he said: "Long-horizon tasks are still a joke. They do not work, and I do not care what anybody says. Do not show me a stupid evaluation. Do not tell me about some dumb script you ran for 48 hours. Long-horizon tasks are not handled well. They simply do not work."
Palihapitiya's critique extends beyond long-horizon tasks to complex problems generally. "Second, complex problems also do not work. They are neither addressed nor handled well," he said. This is not a minor technical quibble; it's a structural warning about the industry's trajectory. He invoked the standard technology adoption curve: initial hype, then a natural contraction leading to a "trough of disillusionment," followed by gradual adoption of the real solution. The internet followed this path, he noted.
The stakes are enormous. "We are spending hundreds of billions, potentially trillions, of dollars trying to figure out how to cross this chasm," Palihapitiya said. If the industry fails, he warned, "people will reach the trough of disillusionment and say that AI was a joke."
His proposed solution is notable for its contrarian bent: "At a very basic level, you need a symbolic space that guides the embedded space." This is a direct challenge to the scaling paradigm that dominates current AI research, where more data and compute are assumed to yield general intelligence. Palihapitiya is suggesting that pure statistical learning may be insufficient for tasks requiring long-range planning and reasoning.
The timing of these remarks matters. In 2026, the industry has poured unprecedented capital into AI infrastructure, with frontier labs like OpenAI, Anthropic, and Google DeepMind competing for dominance. Yet independent evaluations—such as the 2025-2026 swe-bench results and agentic benchmarks—have shown persistent failures on tasks requiring multi-step execution over extended horizons. For instance, many agent frameworks still struggle with tasks exceeding a few hours of simulated work, despite claims of "agentic" capabilities.
Palihapitiya's "symbolic space" idea echoes older AI paradigms—GOFAI (Good Old-Fashioned AI) and neuro-symbolic approaches—that were largely abandoned in the deep learning era. His suggestion is not new, but it carries weight coming from a prominent investor who has backed AI companies through his venture firm, Social Capital. The question is whether the industry will heed the warning or dismiss it as the musings of a contrarian billionaire.
The video, posted on the "techniahqrobot" YouTube channel, captures the full talk. Palihapitiya's candor is a refreshing counterpoint to the relentless optimism of vendor press releases, but his critique is not without precedent. Researchers have long noted that LLMs struggle with planning and long-horizon tasks—see the 2024 paper by Valmeekam et al., which showed GPT-4's poor performance on planning benchmarks. Palihapitiya is amplifying a known problem to a broader audience.
What is missing is data. The source does not include specific benchmark numbers or examples of failures, only Palihapitiya's assertions. That leaves room for skepticism: is he speaking from direct experience with his portfolio companies, or is this a general impression? The answer is unclear from the source.
What to watch

Watch for Palihapitiya's portfolio companies' AI products in 2026: if he backs a neuro-symbolic startup, that signals real conviction. Also track next-generation agent benchmarks (e.g., new swe-bench releases) to see if long-horizon scores improve materially within 12 months.









