Zendoric
← Back to the day · September 3, 2026

These riddles sink AI models: could you solve them better?

🕒 Published on Zendoric: September 3, 2026 · 10:20

✨ AI-generated · how it's made

Games and puzzles have accompanied artificial intelligence since its origins. The term "machine learning" itself was popularized in 1959 by a paper from IBM scientist Arthur Samuel on an algorithm capable of learning to play checkers, and chess and Go later became…

Games and puzzles have been part of artificial intelligence since its very beginnings. The term "machine learning" itself was popularized in 1959 by a paper from IBM scientist Arthur Samuel on an algorithm that could learn to play checkers, and chess and Go later became famous testbeds. A feature by Grace Huckins published in MIT Technology Review, in its September/October 2026 issue, revisits that tradition to offer readers seven puzzle-style tests for checking where AI still lags behind human reasoning.

The starting point is that, measured solely by its skill at solving puzzles, AI is advancing fast. A team at Columbia University showed that in late 2024 the best models got only about 18% of the New York Times' famous "Connections" puzzles right; by early 2025, some models were solving them almost perfectly. But that aggregate progress conceals very specific and revealing failures, which the article organizes into several thematic blocks.

The first is spatial reasoning, illustrated by the classic "mental rotation" problems typical of IQ tests, in which you have to determine whether different images represent the same object seen from different angles. Although today's models already process images, they still fail strikingly at this kind of task: despite all the talk of "world models" capable of understanding physical environments, LLMs cannot manipulate three-dimensional objects the way trained spatial thinkers—architects or mechanical engineers—do. The text cites as a reference the study "Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models" (2025).

The second block addresses memory and adaptability through "knights and knaves" puzzles, in which some characters always tell the truth and others always lie. A 2024 study by researchers at Google and the University of Illinois Urbana-Champaign trained and evaluated models on slight variants of this classic puzzle and found that when a question closely resembles another one seen during training, the model tends to overlook the key differences and answers with what it memorized, instead of reasoning about the specific case. The article notes that the same mechanism could explain how models behave on SimpleBench, a set of questions designed to resemble more complex problems the models probably saw during training: humans spot the trick easily, but even state-of-the-art models get them wrong.

The third block deals with abstract and visual reasoning in two dimensions, with the ARC-AGI benchmark as its central reference. These problems require inferring an abstract, general rule from just a few examples, and models perform better when they receive each grid not as an image but as a string of numbers encoding the color of every cell. According to the research cited, even when models get these puzzles right they usually do so by applying convoluted rules that generalize poorly, whereas humans fall back on simple visual concepts. Even so, the article acknowledges that models have improved a great deal on ARC-AGI over the past year, although certain specific puzzles—such as the one included in the feature itself—still leave them stuck.

A fourth section turns the premise around: it is not only AI that falls into cognitive traps; humans have our own, and we do not share all of them with the models. The feature describes a set of "lightning round" questions devised by psychologists to invert the SimpleBench phenomenon: here it is people who tend to give impulsive answers—because of intuitive arithmetic errors, or because of questions phrased to suggest an obvious answer that falls apart as soon as you read carefully—while the models respond more deliberatively. The academic reference is the study by Hagendorff, Fabi and Kosinski published in Nature Computational Science in 2023 on human-like reasoning biases that emerged in large language models.

The fifth and final substantive block is growing complexity. Research by Apple showed that LLMs have no trouble solving simple versions of the Tower of Hanoi (moving a stack of disks without ever placing a larger one on top of a smaller one) and of the classic river-crossing problems, but start to fail beyond about six disks or six people. That Apple study went viral, although—the article recalls—it also sparked debate over whether it really reveals a limitation inherent to LLM reasoning or simply reflects the fact that it is normal to make more mistakes as a problem gets harder. Along the same lines, researchers at the University of Washington, Stanford and the Allen Institute for AI (in the work known as ZebraLogic) observed that models stumble in a similar way on "logic grid" puzzles, which require deducing several people's attributes from a list of clues.

Beyond the playful format, the feature leaves an underlying message: the contrast between near-perfect scores on heavily worked benchmarks and embarrassing stumbles on subtle variations of classic puzzles suggests that much of the models' performance depends on recognizing patterns memorized during training, rather than on genuinely generalizable reasoning. At the same time, the fact that humans retain intuitive biases that AI does not share complicates any simplistic narrative that "AI already thinks better than we do at everything": each puzzle, the author notes, illuminates a different way in which human and machine cognition diverge, and anyone who manages to solve all the examples offered can say, at least for now, that they beat AI on its own turf.

🔗 Related on Zendoric

Sources & references