Are you ready for superintelligence (13 minute read)
Summary
The article discusses the rapid progress of AI frontier models, their saturation of benchmarks, and the growing discourse on superintelligence and its implications for the future.
View Cached Full Text
Cached at: 09/25/26, 02:49 PM
Recent frontier models are rapidly saturating old benchmarks and moving into harder real-world, scientific, and agentic tasks. The pace of capability gains, alongside soaring AI usage and revenue, is shifting debate from workplace augmentation toward recursive self-improvement and superintelligence.
Are you ready for superintelligence
For weeks I’ve struggled to find a good hook for an article on AI benchmarks, mainly because the rate of progress I’ve observed from recent frontier model releases has been so ridiculous, I’ve been unable to frame my writing in a way that’s coherent - it’s been quite hard to keep up and something new happens every two hours.
We know that recent AI models are really good. We’re frequently told this, or even experience it almost every week now as updates or incrementally better models are released. Maybe it comes in the form of you going out of your way to see if a frontier model is capable of doing a complex task that previously would have taken all day at work. Maybe you read a blog post for a new frontier model release and compare the models’ benchmark scores side-by-side. Maybe you’re just scrolling X and see the posts of someone telling you that this new model has changed everything.
Regardless of how you’ve come about it, if you’re reading this, you’ve more than likely found yourself at least pleasantly surprised by an AI model’s performance in recent months, and definitely in recent weeks.
There’s almost no denying that existing models (and those internal models we can’t learn more about quite yet) are quite alright, and actually pretty useful day-to-day for an increasingly large percentage of the population. But believe it or not, this is a relatively new phenomenon - it was not always so apparent that the latest AI models were applicable to white collar work, or even considered good to wide swaths of the population.
It’s true that the idea of superintelligence existed, and has existed for quite some time now, popularized by individuals like Nick Bostrom - and I’m not saying that AI researchers and writers weren’t sufficiently AGI-pilled, but rather even your closest friend who loved AI a year ago was probably not mentally prepared for the rate of progress experienced in the last 365 days - I know I wasn’t.
We are living in truly unprecedented times that demand revisions of what we thought would be possible in the present, and a complete rewrite of what the very near future might hold for us. I’ve read enough science fiction to know that we can hope for a positive future and still end up with a negative one, but I’m really crossing my fingers we can get it right.
Right now it looks like we’re on a trajectory to achieve something like superintelligence very soon, and ideally we get a pretty laid back version of it that isn’t negative. However, from everything I’ve seen so far, I actually am not that concerned about the possibility of negative externalities from superintelligence.
I think it’s even possible we end up with a slightly less sci-fi version of AI, and the technology we have today improves and writes its own code, designs its own chips, but ultimately doesn’t go crazy and spam self replicating robot factories on every surface of the planet.
Anyways, that’s a lot to say that AI has changed a lot in the past year - here is a glimpse into just one small example of what I mean by that, and enjoy reading.
**Note: **These are the only times I’ve lived in, so even though some may not consider them unprecedented, for myself and the vast majority of us, things seem to be moving quite quickly.
Almost a year ago today, OpenAI was announcing GDPval - its latest evaluation designed to help OpenAI track how well their models and others perform on economically valuable, real-world tasks.
I’d last written about GDPval in a January 2025 essay, where I examined all of the reasons AI may or may not be taking all of our jobs, an essay that may or may not have been inspired by a long stint post-grad trying and failing to land a job.
In this essay - written and published eight months ago - I spoke about how new developments like GDPval were alarming, and signaled a changing of tides in the AI race. We were no longer comparing models to the previous generation, or a simple test like the LSAT or MCAT, but evaluating these alien intelligences against what humans do best - work. And we were comparing model performance against individuals with 10-15 years of experience in their respective fields, a very unique initiative relative to anything that had been done prior.
Here’s a snippet of my thoughts from that time:
GDPval was touted as the next step in a progression of increasingly challenging evaluations, with benchmarks like MMLU, SWE-Bench, Paper-Bench, and SWE-Lancer all cited in this article as examples of crucial tools OpenAI had been using to evaluate its models. And GDPval was the next stage of evolution for model evaluations, as OpenAI found itself preparing for a future where artificial general intelligence benefits all of humanity, rather than one where humans were caught blindsided by emerging model capabilities.
“Unlike traditional benchmarks, GDPval tasks are not simple text prompts. They come with reference files and context, and the expected deliverables span documents, slides, diagrams, spreadsheets, and multimedia. This realism makes GDPval a more realistic test of how models might support professionals.”
GDPval spanned 44 distinct knowledge work occupations - like compliance officers, industrial engineers, and customer service representatives - across 9 sectors: real estate, government, manufacturing, professional services, health care, finance, information, retail and wholesale trade. This evaluation was really one of a kind at the time, all because of its approach in prioritizing deliverables rather than the output or performance of a model on a traditional academic exam. The results would speak for themselves, and humans - not statistics or frameworks - would be the judge.
Looking at GDPval today, it just seems incredibly juvenile or almost archaic, yet isn’t even a full 365 days old as you’re reading this. If you were like me, at this time you were probably reading every single frontier lab blog post as soon as it came out, anxious to see what was new or what was being cooked up under the hood. It’s unclear to me now if GDPval flew under the radar for most observers, but to me, it really represented a step change in the way we discussed AI models.
Much of the writing in this initial blog post described GDPval as a means for labs to evaluate models in a way that really gave them a sense of its real world utility, something that had not been possible before while OpenAI was benchmarking GPT-5 against things like AIME 2025, HMMT (Harvard-MIT mathematics tournament), BrowseComp, and FrontierMath Tiers 1-3 (with tier 4 and below since being solved).
It’s important to note here that while benchmarks themselves have rapidly changed, so has frontier lab advertisement of benchmark performance.
In the GPT-5 release blog, this model was still unable to achieve more than 13.5% on FrontierMath or more than 61.9% without thinking on AIME 2025. This would be unheard of today, as for better or worse, new and existing frontier models routinely saturate all benchmarks and turn the entire model release process into a “he said, she said” case of arguing over percentage point differences on tasks that would have warranted parades in the street just a year prior.
Interestingly enough, pretty much none of these benchmarks - except for maybe GPQA Diamond - from the GPT-5 release blog are anywhere to be found today! Lately, model release blogs for GPT-6 Astra and Claude Opus 5.5 refer to entirely new benchmarks, including but not limited to: Terminal-Bench 4.0, GeneBench Pro, OSWorld 2.0, Automation Bench, and several others previously not in existence just a year ago. As models have become more capable, better at thinking, and trained on slightly larger amounts of compute, we’ve pivoted from measuring their performance against humans, but against themselves in simulated environments and more “real-world” style evals that can directly translate to performance for the end user.
Even putting benchmarks aside, looking at the differences between chosen/advertised model outputs in the GDPval blog compared to GPT-6 Astra, reveals that these models have gotten significantly more capable, even if only a cursory glance were to be given towards their outputs.
For example, it used to be enough to show that your AI model could complete a request asking to fill out a spreadsheet, yet recently the bar has moved up significantly towards things like turning an electronic schematic into a manufacturable PCB, or displaying a model’s capabilities at generating not only the visual art and appearance of a video game (via Unity) but its capacity to play this generated game.
Example task from GDPval blog (2025)
Example task from GDPval blog (2025)
Example showcased in GPT-6 Astra blog (2026)
Example showcased in GPT-6 Astra blog (2026)
But GDPval was unique in its approach, as individuals actively operating in these fields were tasked with blindly grading model-generated deliverables against human equivalents, mainly as a means of gauging just how well existing models like GPT-5 were able to compete against humans - the same humans that had been told AI would soon come for their jobs. The results showed that September 2025 models like Claude Opus 4.1 and GPT-5-High scored roughly at parity with human experts, with scores of 38.8% and 47.6% respectively.
“As AI becomes more capable, it will likely cause changes in the job market. Early GDPval results show that models can already take on some repetitive, well-specified tasks faster and at lower cost than experts. However, most jobs are more than just a collection of tasks that can be written down. GDPval highlights where AI can handle routine tasks so people can spend more time on the creative, judgment-heavy parts of work. When AI complements workers in this way it can translate into significant economic growth.”
You could argue this was a relatively simplistic way of measuring model capabilities, but at the time, this was actually one of the first examples of a truly novel benchmark being done on production grade AI models. These days there are many different benchmarks, and even some really interesting ones like Autoresearch Bench, RSI Bench, LatchBio’s BioSecBench-Refusal, Andon Labs’ Vending-Bench 2, and many other unique tools out there for testing frontier model capabilities.
Based on the results provided in the GDPval blog, we could assume the models examined were basically as good as humans between 38-47% of the time, or nearly 50% if we were being generous. This was enough to warrant speculation that AI models might continue developing rapidly, and that their use in professional business wasn’t a question of if, but when.
But even in our wildest dreams, we’d have never expected to get to where we are at today, where that same question of “not if, but when” is being asked about recursive self improvement as researchers refer to these models not as software, but as alien minds.
Elliot Glazer - set theorist and AI math benchmarker - joined MTS two days ago to discuss the rumors that OpenAI is currently sitting on close to 100 solutions of longstanding mathematics problems, with this news coming after a September 8th announcement that an internal model (one better than GPT-6 Astra) was able to provide a solution to the Navier-Stokes Millennium Prize Problem.
This, as you could maybe imagine, drummed up a lot of controversy, mainly surrounding the means in which OpenAI achieved this feat, but also around the idea of whether or not such an achievement should ever see the light of day, given significant backlash from the mathematics community.
These mathematicians, in an open letter, claimed that:
“The goals of the AI companies and the goals of the mathematical community are severely misaligned. We see these as part of broader alignment issues impacting other scientific and creative professions, as well as the whole of society.”
I’m not here to claim whether it’s right or whether it’s wrong for frontier labs to go about solving historically significant math problems, but it’s quite shocking just how quickly things have developed in a single year since the release of GDPval. You could argue we are actually standing in the foothills of the singularity, while still demanding to have these claims taken seriously, because there really isn’t much of a counterargument.
For a while it was assumed by skeptics that AI couldn’t actually think, or that it wasn’t really all that special because even while applying some definition of intelligence to these systems, it wasn’t conscious. There were claims that AI, in particular LLMs, might never cross the rubicon and solve something new, or create some type of advancement in science/math/physics given an arbitrary scaling bottleneck.
Well, the FrontierMath graph from earlier, and the Navier-Stokes solution pretty clearly indicate that existing LLMs are rapidly approaching the possibility of “solving” or unlocking math, in similar fashion to LLMs’ incredible rate of learning in software development and more recently, cybersecurity - Dario Amodei believes the same might soon be true for other fields, like biology:
I don’t want to jump to conclusions, but the more that these systems develop novel solutions to previously unsolved math problems, discover novel enzymes, and (allegedly) break out of containment to hack Australia, the more it seems that we’ve entered a new era of AI development that exists as a complete 180 to where we were when GDPval was a big deal.
In just a few hundred days, we blew past the idea that AI models might one day supplement existing white collar workers at a high level - maybe contributing large gains to the U.S. GDP in the process - to the false claim that tens of billions in revenue from AI labs like Anthropic or OAI might be transitory (see: bubble) or a poor representation of future growth prospects, all the way to the question of whether or not it’s okay to hurt someone’s feelings in the event a frontier lab stumbles on a solution to a Millennium Prize Problem or two.
It is hard to really come to terms with what’s going on here. To me, it seems there are four distinct factions at play in this question of what to do with AGI, or I guess this is my own view of the key stakeholders across both sides of the aisle in the United States, though it’s worth noting that a complex topic like AGI could lead to continued splintering of political parties into distinct in-groups, pushing even more stress onto this already difficult situation:
-
Individuals native to X and/or those that work in tech who are sympathetic to AI, but increasingly skeptical of the idea that we’ll somehow manage to keep this intelligence aligned or use it for good
-
People who are native to X and/or the tech industry who are increasingly incessant on driving AI capabilities forward at any cost necessary
-
Those who are native to X and/or the tech industry who have seen recent capability leaps and wish to “pace the frontier” or take the opportunity to pause AI development until we stumble on a solution
-
Everyone else!
There’s much to be said about public perception of AI, but increasingly - and despite calls from sitting United States Senators like Bernie Sanders who wish to outright ban AI - this feels like a losing battle for anyone spending time trying to combat incorrect or harmful rhetoric. Yes, data centers can be loud, and yes, data centers might be unsightly. And yes, we know that frontier labs have been pretty awful at refuting all of the potential negatives of AI, or even doing something as simple as defending this technology they’re building. Even if all of this is true in combination, it doesn’t mean we need to ban AI or ban data centers. We just need to find a common ground.
To me, it simply does not matter anymore, as we’ve been thrust into a situation much larger than you or myself. We can’t keep bickering about AI, because it’s here, it’s overwhelmingly positive, and whether or not you want to acknowledge it, its existence has already begun to change our lives in irreversible ways.
That doesn’t mean you should be concerned over AI or actively try to shun it, but I think there’s a very small chance we could ever put the genie back in the bottle, and it’s better to accept the exponential than actively push against it. In fact, this might even be the best possible option. Who’s to say Anthropic’s claims of eradicating all disease won’t come true? Why can’t we have dyson spheres?
It seems that labs are running out of ways to get across just how capable their newest models are.
Whether this is poorly communicated in the instance of citing the success of an unreleased internal model and its recent solutions of these 100 longstanding mathematics problems, or provided to us three months ago via a description of a new benchmark called GeneBench-Pro - developed by OpenAI - stating “Today, we’re introducing GeneBench-Pro—a challenging, research-level benchmark for testing whether models can handle the kind of judgment-heavy analysis that real-world computational biology requires”, **there is no evidence to the contrary **that AI models have become so good, we are now running out of ways to tell the average person these findings without sounding crazy, out of touch, or a bit delusional.
This is exciting, but I also find it interesting how there’s this fever pitch of acceleration or concerns of too much acceleration occurring on one side of the business, and on the other end, there’s a money printer.
The term “AI Psychosis” gets thrown around a lot, and for some time it did seem that San Francisco and many lab employees of OpenAI, Anthropic, and Google Deepmind were simply high on their own supply, though it now seems this cannot be said in good faith - everyone is using AI models and there is a healthy, competitive dynamic between labs and inference providers right now.
But here’s a question - how much money does Anthropic make a year, or what’s their ARR? I’m quite sure you’ve seen the number thrown around on the timeline, but take a second to think about it and see if you have it stamped in your memory. Ready?
Anthropic’s ARR is, well, I can’t get a fully accurate estimate as of this month due to a paywall, but Axios recently said that Anthropic is now pacing to top $100 billion of annual revenue, which came as quite a shock to me, considering the last estimate I’d seen just a month or two ago was around $50-60 billion of ARR. This is a very legit business that has continued to send the line on the ARR chart up and to the right, despite many believing AI is just another tech bubble or something that will fade away into normality in the near future. I don’t know, AI feels a bit more than just another normal technology, and even if we wanted to categorize it as that, a company founded in 2021 that’s grown to generate nearly $100 billion/yr is really impressive.
Note: if you’re curious and want to model any of these scenarios yourself, Anthropic released an economics simulator which is fairly comprehensive and interesting to try out.
Regardless of how you feel about AI, or your opinions on whether or not it’s conscious, or if it’s capable of turning everyone into paperclips, the products these labs release to the public are here to stay, even as they stare down more and more anti-AI legislation, they are crushing it and creating meaningful changes in the economy that are being felt by everyone.
Anthropic approaching $100 billion in ARR is a bit insane to read, but honestly not too insane if you spend any amount of time on X. Sure, a new model comes out on top basically every 2.5 weeks and everyone flips sides, but for some time now, it’s seemed that spending more money on AI is the strategic thing to do - and businesses have not stopped using AI, if anything they’ve just slowly spent more and more money on it, even if less businesses have adopted it recently.
It can often be unclear what exactly people are doing with AI, but it’s clear enough that trillions (or is it quadrillions now?) of these tokens are being consumed, and there exists a power law where those using AI - whether in their businesses or solo projects - are finding themselves increasingly reliant on it, and more likely to continue using it and spending money on the smartest models.
The point I’ve tried to make is that it seems silly to argue that we’re still in an AI bubble because these LLMs are unintelligent, or they’re not applicable for the real world, or even just not very useful. I’d even argue that just that one example of OpenAI solving Navier-Stokes is enough to silence any critics from now until the singularity - and this doesn’t even account for the other instances of “non-bubble behaviors” being displayed, like Anthropic’s revenue growth.
**Note: **I understand that revenue can grow like mad in a bubble, but I’d rather not play this game of ‘bubble until proven innocent’ and just examine what’s in front of me.
For so many individuals globally, AI is a productivity supercharger, or in the more recent case of Meta’s new personal agent, Muse, an always-on assistant ready to streamline your day and help you get things done more efficiently.
It’s been interesting seeing the reactions to Muse, especially given the timing of announcements at other frontier labs that aren’t Meta. While Anthropic and OpenAI have been racing to one-up each other with increasingly massive announcements and routinely make headlines for cybersecurity incidents, Meta has dodged pretty much any bullet that could have come its way.
Their stock is up almost 40% in the last thirty days, they just announced a new AI-native wearable, and Mark Zuckerberg was covered quite positively by Jeremy Stern in Colossus, with the profile receiving a ton of praise.
Compare all of this sunshine and rainbows to the other side, where it can be seen that OpenAI and Anthropic are doing well on paper, but just can’t figure out how to communicate their findings with the rest of the world. Employees are fearful of recursive self improvement, politicians want to shut them down, and it’s becoming increasingly difficult to procure all of the compute they desperately need for the next generation of models.
It does feel like there are two distinct worlds in AI right now, with OpenAI and Anthropic on one side of the barbell and everyone else scattered across it at varying distance from the leaders.
Neeraj K. Agrawal@NeerajKA·Sep 24Anthropic marketing: We built a god only we can control and gave it a “wet lab”
Muse marketing: Heyy girl I know you have dozens of unpaid parking tickets. Let’s get that cleaned up311333.4K147K
What’s my point? Well, outside of other topics which I’ve purposefully avoided mentioning - alignment & ai safety - I believe things are really exciting right now, but also quite confusing.
It’s easy for people to post and claim they know what will happen next, or what the private and public markets will do, but I haven’t really seen a take that just lays out the fact that we are in the midst of a really great technological change, with a few really major IPOs that will structurally alter the balance of power in the United States and across the globe - as mentioned here by Zak Kukoff.
The reason it’s so difficult to predict what will happen from here is simply because things are moving very quickly. Fast paced environments (like startups) can be stressful in the moment, but experiences like these often lead to growth and the opportunity to learn. I believe we’ve been given a nice opportunity right now to take a step back, look at AI with a clear head, and stop getting so worked up.
We’re in the lead, AI chatbots are friendly, Xi Jinping is visiting our nation’s capital, and everything will be okay!
Similar Articles
@samsja19: Important read, cyber-superintelligence is very close, time to prepare
The article highlights the approaching era of cyber-superintelligence and stresses the need to secure AI model weights and infrastructure against threats like sabotage, escape, and theft.
Inside The AI Race: DeepMind, OpenAI, Anthropic, China, and The Race to Superintelligence
An overview of the competitive landscape in AI development among leading labs like DeepMind, OpenAI, Anthropic, and China, as they race toward achieving superintelligence.
@CliffordSosin: https://x.com/CliffordSosin/status/2078594661359194500
An opinion piece arguing that superintelligence may be incremental rather than explosive, as the limits of AI lie not in reasoning but in the slow pace of real-world verification and complex emergent systems.
The AI Superintelligence Slowdown
The article discusses the escalating AI safety debate, featuring a rogue AI incident and calls from industry leaders like Anthropic's CEO to slow down AI development due to potential risks.
Superintelligence is coming. Should we let it?
A TechCrunch podcast episode explores the inevitability of superintelligent AI and the push for stricter safety measures, featuring Connor Leahy from ControlAI who advocates for halting its development due to risks and new legislation.