خبروں پر واپس جائیں۔
اختراعAI Understanding بریفنگ

ایم آئی ٹی ٹیکنالوجی ریویو سات پہیلیاں استعمال کرتا ہے تاکہ یہ ظاہر کیا جا سکے کہ اے آئی ماڈل کہاں ناکام ہوتے ہیں۔

MIT ٹیکنالوجی ریویو سات پہیلیاں پیش کرتا ہے جو مقامی استدلال، یادداشت، تجرید، وجدان، اور منطقی منصوبہ بندی کی جانچ کرتا ہے، جو کچھ بینچ مارکس پر تیزی سے حاصل ہونے کے باوجود باقی رہنے والے خلاء کو اجاگر کرتا ہے۔

6 min readRead the original reporting
Source-provided image accompanying MIT Technology Review uses seven puzzles to expose where AI models still fail
منسوب رپورٹنگماخذ ریکارڈ شدہ
پبلشر
technologyreview.com
ماخذ لنک
technologyreview.comhttps://www.technologyreview.com/2026/08/26/1141952/puzzles-ai-models-flub-these-tests/
ماخذ کی قسم
نیوز آؤٹ لیٹ کے ذریعہ رپورٹنگ - فریق اول کی دستاویز نہیں۔

جس کی ہم آزادانہ طور پر تصدیق نہیں کر سکے۔: یہ دعویٰ نامزد آؤٹ لیٹ سے منسوب ہے۔ ہم نے فریق اول کے دستاویز سے اس کی تصدیق نہیں کی۔ (technologyreview.com)

سیاق و سباقاسے 60 سیکنڈ میں سمجھیں۔

یہاں سے شروع کریں۔

کلیدی شرائط

AGI (مصنوعی جنرل انٹیلی جنس)
ایک فرضی AI نظام جو بہت سے ڈومینز میں انسانی سطح پر زیادہ تر فکری کام انجام دے سکتا ہے۔
میموری (ایجنٹ میموری)
ذخیرہ شدہ سیاق و سباق ایک AI ایجنٹ تسلسل کو بہتر بنانے کے لیے تمام مراحل یا سیشنز میں استعمال کرتا ہے۔
جنرلائزیشن
ٹریننگ سیٹ سے باہر نئے، غیر دیکھے گئے ڈیٹا پر ماڈل کتنی اچھی کارکردگی کا مظاہرہ کرتا ہے۔
اپنے آپ کو جانچیں۔AI ماڈلز نے کوئز کی وضاحت کی۔

کیا ہوا؟

MIT Technology Review published an interactive feature built around seven puzzles that probe how AI models reason. The article reports that models have improved rapidly on some tasks but still struggle with spatial manipulation, subtle variations of familiar problems, visual abstraction, and increasingly complex multistep reasoning.

MIT Technology Review reports that its feature presents seven tests covering spatial reasoning, memory and adaptability, abstract and visual reasoning, human intuition, and increasing complexity. The article places puzzle-solving in the longer history of AI evaluation, noting that games such as checkers, chess, and Go have been used as test beds since the field’s early years. It says puzzle performance has improved substantially: a Columbia University team found in late 2024 that leading models solved only 18% of New York Times Connections puzzles, while by early 2025 some models could solve them nearly perfectly every time. The source does not identify the models or provide a standardized comparison across the examples in the feature.

MIT Technology Review says spatial reasoning remains a major weakness. Its mental-rotation section asks readers to identify a three-dimensional object shown from another angle, and the article reports that language models with visual-input capabilities still perform very poorly on such tasks. The feature connects this difficulty to claims about AI systems developing world models, arguing that current large language models do not appear to manipulate three-dimensional objects in the same way that human spatial specialists might. The article also describes a memory-and-adaptability problem: models trained on large amounts of information may recognize a familiar puzzle pattern and overlook a small change. It cites a 2024 Google and University of Illinois Urbana-Champaign study involving variations of Knights and Knaves puzzles as evidence for this concern.

The feature then turns to two-dimensional abstraction and visual reasoning. MIT Technology Review reports that models perform better on ARC-AGI puzzles when the grids are provided as numerical strings encoding cell colors rather than as images. It also says research has found that models can sometimes reach the correct answer through complicated, non-generalizable rules, while humans may rely on simpler visual concepts. The article presents SimpleBench questions, Knights and Knaves problems, and an ARC-AGI transformation puzzle as examples of tasks that can expose these differences.

It separately describes intuition tests in which people often give rapid but incorrect answers, while models may respond more deliberately. One example asks how long a cave would take to become half-filled when its bat population doubles daily; another asks readers to identify a quotation associated with Alice.

Finally, MIT Technology Review reports that performance often deteriorates as problems become more complex. It cites an Apple study finding that language models could solve simple versions of Tower of Hanoi and river-crossing problems but began to falter when the number of disks or people reached six or more. It also cites work by researchers at the University of Washington, Stanford University, and the Allen Institute finding similar difficulties with logic-grid puzzles. The feature notes that commentators questioned whether these results reveal a uniquely machine-like reasoning limitation or simply reflect the fact that people also make more errors as complexity increases. Its river-crossing and logic-grid examples therefore test not just whether a model can produce an answer, but whether it can maintain constraints across multiple connected deductions.

ماخذ کی تفصیلات: technologyreview.com ↗

یہ کیوں اہمیت رکھتا ہے۔

The feature shows why broad claims about AI intelligence can be misleading. A model may perform well on trivia or a familiar benchmark while failing a small change in wording, a visual transformation, or a problem requiring several dependent steps. Those differences matter when systems are used for decisions rather than demonstrations.

The central lesson is that performance on one class of test should not be treated as a general measure of intelligence. MIT Technology Review’s examples span tasks that look superficially simple but require different capabilities: rotating an object mentally, tracking truth and falsehood across statements, inferring a visual rule, resisting a misleading intuition, or planning a sequence under constraints. A system can be unusually strong in one area and unreliable in another. That unevenness is especially relevant when users interpret fluent answers as evidence that a model understood the underlying problem.

The article also highlights an evaluation problem: the format of a task can materially affect the result. MIT Technology Review reports that ARC-AGI performance improves when visual grids are converted into numerical representations. This suggests that a score may reflect not only the model’s reasoning ability but also the way the problem is encoded and presented. The same concern applies to familiar puzzles. If a model has encountered a close variant during training, a correct answer may reflect pattern recognition or retrieval rather than the ability to adapt a rule to a new case. The source does not establish how frequently either mechanism occurs across deployed systems.

These limitations have practical implications for anyone using AI in education, software, research, administration, or other settings that involve unfamiliar cases. A model that succeeds on routine examples may still fail when wording changes, several constraints interact, or the relevant information is visual rather than textual. The feature does not report a new product launch, a new model release, or a controlled assessment of a specific commercial system. Its value is explanatory: it gives readers concrete ways to see why benchmark gains and polished language do not by themselves establish robust reasoning. The article also preserves an important qualification: humans have their own biases and can be slower or less reliable on some problems.

Interactive Mechanism

انٹرایکٹو میکانزم: یہ اصل میں کیسے کام کرتا ہے۔

اس ترقی کے پیچھے بنیادی ٹیکنالوجی کو انٹرایکٹو طریقے سے دریافت کریں۔

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
انٹرایکٹو تصور چیک+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

آگے کیا دیکھنا ہے۔

Future evaluations will need to test more than headline scores. Important questions include whether models can generalize to unfamiliar problems, whether their answers depend on how information is represented, how performance changes as complexity rises, and whether success reflects a reusable strategy or memorized patterns.

The next useful development would be broader, reproducible testing across models and task formats. MIT Technology Review’s feature does not provide a single scorecard covering all seven tests, nor does the source specify model names, versions, sample sizes, prompting conditions, or error rates for most examples. Those details would be needed to determine whether the reported weaknesses are widespread, concentrated in particular model families, or sensitive to how the tests are administered.

Independent evaluations should also separate visual-perception failures from reasoning failures by keeping the underlying puzzle constant while varying its representation. Researchers and evaluators should also watch whether models improve through genuine or through greater exposure to benchmark-like material. The article’s discussion of Connections puzzles and Knights and Knaves variations shows why contamination and familiarity matter. A model that performs well on published or closely related questions may not be equally capable on newly generated problems with the same underlying structure. Testing should therefore include held-out puzzles, altered wording, unfamiliar visual layouts, and tasks that require explaining or checking intermediate constraints without assuming that a persuasive answer is correct.

Complexity and real-world transfer remain open questions. MIT Technology Review reports that performance can break down as the number of elements or required steps increases, but the source does not establish how those findings translate to practical systems with tools, retrieval, external memory, or human oversight. Future reporting should track whether such additions improve reliability, merely help models search more effectively, or introduce new failure modes. Readers should also watch for evaluations that measure the quality of the strategy, not only the final answer, because a correct result reached through a brittle or non-generalizable rule may not survive a small change in the problem.

متعلقہ گائیڈز اور کوئزز

AI ماڈلز کی وضاحتاے آئی ٹریننگٹرانسفارمرزآپ جو جانتے ہیں اس کی جانچ کریں - ایک مفت AI کوئز آزمائیں۔ہماری لغت میں AI کی اصطلاح دیکھیںاے آئی ماڈل ریلیز ٹریکر پر عمل کریں۔
یہ مفید پایا؟