Zendoric
← Back to the day · September 4, 2026

The end of benchmarks: why GPT-6 Astra makes leaderboards unreadable

🕒 Published on Zendoric: September 4, 2026 · 09:12

✨ AI-generated · how it's made

OpenAI made it official this week: the company is presenting GPT-6 Astra as its entry into the 'AGI Era', pitching it as a generational leap with major advances in computer use, coding, research and the kind of long, messy work that until now required constant human supervision.

By AI Secret (Magna & Ben) · September 3, 2026.

OpenAI made it official this week: the company is presenting GPT-6 Astra as the entry point into the 'AGI Era', selling it as a generational leap with major advances in computer use, programming, research and in the kind of long, messy work that until now required constant human supervision. The newsletter's authors take that marketing as a starting point to raise a question they say the industry has long been dodging: how do you actually measure this?

According to the article, AI has spent years running on benchmarks: every launch arrives with a table showing that the new model is 3% better here, 7% better there, and mysteriously 12% smarter according to a benchmark that didn't exist the previous spring. Models get ranked, tables get color-coded, leaderboards get stacked and a winner gets crowned. The piece argues that this system worked reasonably well when models basically answered questions, but it stops working the moment models start *doing* things.

The central argument is that a benchmark is nothing more than a question bank: you give the model a problem, it produces an answer, and that answer is compared against an answer key. Even the most sophisticated 'computer use' benchmarks, the authors point out, are at bottom a set of pre-written tasks running inside equally pre-built environments. That is useful, but it looks nothing like what happens when you turn an agent loose on a real machine and ask it to solve something, because real life doesn't come with a benchmark file: nothing on a real desktop indicates exactly what 'done' looks like. A real computer is a browser that crashes mid-task, a website redesigned the day before without warning, a spreadsheet with three contradictory versions, a permission nobody mentioned, a PDF buried six folders down, an API that times out on precisely the worst request, and a person who says 'take care of this' and walks away without defining what 'this' is. The agent has to figure it out, and the part almost no benchmark measures well is its ability to recover when the first attempt fails — a skill entirely different from getting an answer right.

That is why, according to the article, the reports that OpenAI has bought tens of thousands of consumer Macs for reinforcement learning and computer-use training make more sense. The authors themselves caveat that those specific figures have been reported, not audited, but they say the underlying direction is unmistakable: the big AI labs want their models operating on real machines, not locked inside a chat window, and that quietly breaks the benchmark game as it existed.

The newsletter illustrates the problem with an example: two agents, the same computer, the same vague assignment — 'research this market, pick three promising companies, build a financial comparison and prepare a presentation for the investment committee'. There is no single correct way to do it; there are hundreds. One agent browses for two hours, another writes a script, a third stumbles onto a better data source halfway through and scraps its original plan, a fourth makes a mistake, catches it and fixes it, and a fifth delivers a slick presentation built on completely wrong numbers. The benchmark author cannot write down the correct answer in advance, because there isn't just one.

According to the authors, this raises a genuine discomfort: the hardest part of evaluating an agent may have shifted from 'was the answer correct?' to 'was this actually useful?'. To argue that this is not a minor nuance, they offer the example of an agent that scores above 95% on a computer-use benchmark and that, once deployed for a week at a real company, builds the wrong folder structure, overwrites a spreadsheet in production, wastes six hours chasing a dead lead, confidently cites a 2023 statistic and gets stuck at a permissions wall it never occurs to it to ask how to get around: on every benchmark task it scores perfectly, but nobody would keep it employed for a second week. By contrast, an agent that only scores 85% but knows when it's confused, asks the one question that unblocks it, changes strategy when something breaks and delivers usable work is harder to value with a leaderboard, which simply has no way of capturing that difference.

The authors argue that independent benchmarks are hitting a wall not for lack of effort from researchers — who keep building harder and more realistic tests — but because reality remains bigger than any test: you can make the benchmark longer, add more tools, more sites, more apps, more steps, more randomness, and nothing fundamental changes, because it's still a game with rules and, as soon as a model learns to play it, you're back to square one. The result, they say, is that we will keep seeing scores like 91.4%, 94.7%, 97.2% or 98.1%, that people will argue over whether Model A beats Model B, that someone will launch another leaderboard and then another one to rank the leaderboards, while the actual model goes on doing things on a real computer that none of those figures describe. They frame this as the 'benchmark paradox': the closer AI gets to general-purpose intelligence, the less a bank of pre-written questions says about it.

As a way out, the article does not propose 'a better benchmark' but something fundamentally different: stop giving the agent a question and give it a goal; stop offering it a controlled environment and give it a messy one; stop scoring whether it followed the expected steps and score whether it achieved the result; stop repeating the same public test and drop it into constantly changing environments, with private tasks it has never seen; and stop asking only whether it succeeded, and also ask how reliably it succeeds, what it cost, how long it took, how many times it needed a human, how safely it behaved and whether the work survived contact with reality. Their formula: the next generation of AI evaluation should look less like an exam and more like a professional internship — don't ask whether the model passed the test, but give it a laptop, an account, a budget, a goal and a week, and watch what it does. The authors acknowledge that this would say far more about the system's real intelligence, though they admit it would be a nightmare to standardize, and they close with the idea that, at the same time as the 'AGI Era', the 'Era of Unscoreable AI' is beginning.

🔗 Related on Zendoric