Voltar às notícias
InovaçãoInstruções AI Understanding

MIT Technology Review uses seven puzzles to expose where AI models still fail

MIT Technology Review presents seven puzzles that test spatial reasoning, memory, abstraction, intuition, and logical planning, highlighting gaps that remain despite rapid gains on some benchmarks.

Por 6 min read
AI-generated editorial illustration accompanying MIT Technology Review uses seven puzzles to expose where AI models still fail
A versão curta

MIT Technology Review presents seven puzzles that test spatial reasoning, memory, abstraction, intuition, and logical planning, highlighting gaps that remain despite rapid gains on some benchmarks.

O que aconteceu

MIT Technology Review published an interactive feature built around seven puzzles that probe how AI models reason. The article reports that models have improved rapidly on some tasks but still struggle with spatial manipulation, subtle variations of familiar problems, visual abstraction, and increasingly complex multistep reasoning.

MIT Technology Review reports that its feature presents seven tests covering spatial reasoning, memory and adaptability, abstract and visual reasoning, human intuition, and increasing complexity. The article places puzzle-solving in the longer history of AI evaluation, noting that games such as checkers, chess, and Go have been used as test beds since the field’s early years. It says puzzle performance has improved substantially: a Columbia University team found in late 2024 that leading models solved only 18% of New York Times Connections puzzles, while by early 2025 some models could solve them nearly perfectly every time. The source does not identify the models or provide a standardized comparison across the examples in the feature.

MIT Technology Review says spatial reasoning remains a major weakness. Its mental-rotation section asks readers to identify a three-dimensional object shown from another angle, and the article reports that language models with visual-input capabilities still perform very poorly on such tasks. The feature connects this difficulty to claims about AI systems developing world models, arguing that current large language models do not appear to manipulate three-dimensional objects in the same way that human spatial specialists might. The article also describes a memory-and-adaptability problem: models trained on large amounts of information may recognize a familiar puzzle pattern and overlook a small change. It cites a 2024 Google and University of Illinois Urbana-Champaign study involving variations of Knights and Knaves puzzles as evidence for this concern.

The feature then turns to two-dimensional abstraction and visual reasoning. MIT Technology Review reports that models perform better on ARC-AGI puzzles when the grids are provided as numerical strings encoding cell colors rather than as images. It also says research has found that models can sometimes reach the correct answer through complicated, non-generalizable rules, while humans may rely on simpler visual concepts. The article presents SimpleBench questions, Knights and Knaves problems, and an ARC-AGI transformation puzzle as examples of tasks that can expose these differences.

It separately describes intuition tests in which people often give rapid but incorrect answers, while models may respond more deliberately. One example asks how long a cave would take to become half-filled when its bat population doubles daily; another asks readers to identify a quotation associated with Alice.

Finally, MIT Technology Review reports that performance often deteriorates as problems become more complex. It cites an Apple study finding that language models could solve simple versions of Tower of Hanoi and river-crossing problems but began to falter when the number of disks or people reached six or more. It also cites work by researchers at the University of Washington, Stanford University, and the Allen Institute finding similar difficulties with logic-grid puzzles. The feature notes that commentators questioned whether these results reveal a uniquely machine-like reasoning limitation or simply reflect the fact that people also make more errors as complexity increases. Its river-crossing and logic-grid examples therefore test not just whether a model can produce an answer, but whether it can maintain constraints across multiple connected deductions.

Leia a fonte primária: technologyreview.com

Por que isso importa

The feature shows why broad claims about AI intelligence can be misleading. A model may perform well on trivia or a familiar benchmark while failing a small change in wording, a visual transformation, or a problem requiring several dependent steps. Those differences matter when systems are used for decisions rather than demonstrations.

The central lesson is that performance on one class of test should not be treated as a general measure of intelligence. MIT Technology Review’s examples span tasks that look superficially simple but require different capabilities: rotating an object mentally, tracking truth and falsehood across statements, inferring a visual rule, resisting a misleading intuition, or planning a sequence under constraints. A system can be unusually strong in one area and unreliable in another. That unevenness is especially relevant when users interpret fluent answers as evidence that a model understood the underlying problem.

The article also highlights an evaluation problem: the format of a task can materially affect the result. MIT Technology Review reports that ARC-AGI performance improves when visual grids are converted into numerical representations. This suggests that a score may reflect not only the model’s reasoning ability but also the way the problem is encoded and presented. The same concern applies to familiar puzzles. If a model has encountered a close variant during training, a correct answer may reflect pattern recognition or retrieval rather than the ability to adapt a rule to a new case. The source does not establish how frequently either mechanism occurs across deployed systems.

These limitations have practical implications for anyone using AI in education, software, research, administration, or other settings that involve unfamiliar cases. A model that succeeds on routine examples may still fail when wording changes, several constraints interact, or the relevant information is visual rather than textual. The feature does not report a new product launch, a new model release, or a controlled assessment of a specific commercial system. Its value is explanatory: it gives readers concrete ways to see why benchmark gains and polished language do not by themselves establish robust reasoning. The article also preserves an important qualification: humans have their own biases and can be slower or less reliable on some problems.

O que assistir a seguir

Future evaluations will need to test more than headline scores. Important questions include whether models can generalize to unfamiliar problems, whether their answers depend on how information is represented, how performance changes as complexity rises, and whether success reflects a reusable strategy or memorized patterns.

The next useful development would be broader, reproducible testing across models and task formats. MIT Technology Review’s feature does not provide a single scorecard covering all seven tests, nor does the source specify model names, versions, sample sizes, prompting conditions, or error rates for most examples. Those details would be needed to determine whether the reported weaknesses are widespread, concentrated in particular model families, or sensitive to how the tests are administered.

Independent evaluations should also separate visual-perception failures from reasoning failures by keeping the underlying puzzle constant while varying its representation. Researchers and evaluators should also watch whether models improve through genuine generalization or through greater exposure to benchmark-like material. The article’s discussion of Connections puzzles and Knights and Knaves variations shows why contamination and familiarity matter. A model that performs well on published or closely related questions may not be equally capable on newly generated problems with the same underlying structure. Testing should therefore include held-out puzzles, altered wording, unfamiliar visual layouts, and tasks that require explaining or checking intermediate constraints without assuming that a persuasive answer is correct.

Complexity and real-world transfer remain open questions. MIT Technology Review reports that performance can break down as the number of elements or required steps increases, but the source does not establish how those findings translate to practical systems with tools, retrieval, external memory, or human oversight. Future reporting should track whether such additions improve reliability, merely help models search more effectively, or introduce new failure modes. Readers should also watch for evaluations that measure the quality of the strategy, not only the final answer, because a correct result reached through a brittle or non-generalizable rule may not survive a small change in the problem.

Guias e questionários relacionados

Modelos de IA explicadosTreinamento de IATransformadoresTeste o que você sabe – experimente um teste gratuito de IAProcure um termo de IA em nosso glossário
Achou isso útil?