返回新聞
創新AI Understanding 簡報

麻省理工學院技術評論研究了為什麼兒童學習語言比人工智慧模型更有效

《麻省理工學院技術評論》報導,與大型語言模型相比,兒童獲取語言的數據要少得多,研究人員正在測試兒童的感官、探索和社交學習是否可以指導更有效率的人工智慧。

7 min readRead the original reporting
Source-provided image accompanying MIT Technology Review examines why children learn language far more efficiently than AI models
歸因報告來源記錄
出版商
technologyreview.com
來源連結
technologyreview.comhttps://www.technologyreview.com/2026/08/24/1141740/kids-machines-language-learning/
來源類型
新聞媒體的報道-不是第一方文件。

我們無法獨立確認的內容: 此聲明歸因於指定的商店。我們沒有根據第一方文件對其進行驗證。 (technologyreview.com)

背景60 秒內了解這一點

從這裡開始

關鍵術語

概括
模型在訓練集之外的新的、未見過的資料上的表現如何。
預訓練
在下游適應之前對廣泛資料進行初步大規模模型訓練。
基準測試
用於測量和比較模型性能的標準化測試或資料集。
測試一下自己AI 模型解釋測驗

發生了什麼事

MIT Technology Review reports on the data-efficiency gap between children learning language and large language models trained on enormous text collections. Researchers are using small-data language-model competitions, child headcam footage, and larger recordings of children’s early lives to investigate what machines may be missing.

MIT Technology Review reports that children typically begin producing grammatically correct sentences after hearing roughly 10 million words, with estimates reaching about 30 million at the high end. By contrast, modern large language models may process trillions of tokens during . The article describes this disparity as the “data efficiency gap”: children learn language from a small, embodied stream of experience, while models depend on vastly larger collections of text.

The report points to BabyLM, a competition organized by researchers including Alex Warstadt, Leshem Choshen, and others. MIT Technology Review says the main competition limits models to a developmentally plausible corpus of about 100 million words, with a toddler-scale track using 10 million. The data can include storybooks, dialogue, movie subtitles, Wikipedia, and transcripts of speech directed at children. Models are evaluated with grammar tasks similar to those used in psycholinguistics.

According to MIT Technology Review, the competition has challenged the assumption that curriculum learning—presenting simple material before complex material—would help models learn more like children. The approach was popular in the first round but did not perform as well as expected. The article says the 2024 leading system, GPT-BERT, combined next-token prediction with masked-language modeling and, after training on about 100 million words, exceeded Meta’s Llama 2 70B on one BabyLM . The report also stresses that these models remain much weaker than current commercial systems and do not reproduce a child’s general capabilities.

The article reports that text alone may be an inadequate approximation of childhood experience. Michael Frank’s SAYCam project recorded two hours a week of three babies’ lives between six months and two and a half years. MIT Technology Review says a model trained on 61 hours of this raw footage learned to identify objects and associate them with words, but did not produce the capabilities of a two-year-old. A newer project led by Uri Hasson recorded the first 1,000 days of 17 children’s lives, using cameras and microphones in living areas for about 12 hours a day on most days. The resulting dataset was described in a recent preprint, but the article does not provide enough information here to independently assess that preprint’s methods or findings.

MIT Technology Review reports that researchers are considering active exploration and social learning as additional missing ingredients. Alison Gopnik argues that children choose experiences, experiment, and seek predictable effects on the world, while Elizabeth Bonawitz says children reason about teachers and why information is being provided. The article says a BabyLM track allowing models to learn through interaction with other models did not outperform standard approaches. Meta researchers and academic collaborators have also announced a and challenge involving baby headcam footage, although the report does not establish that this effort has produced a successful child-scale model.

來源詳情: technologyreview.com ↗

為什麼這很重要

The research could affect how AI models are trained, especially for video and minority languages where large datasets are unavailable. It may also provide new experimental tools for studying language development, while the article emphasizes that current models remain far from reproducing a child’s broader abilities.

MIT Technology Review reports that data efficiency has become an engineering constraint as AI developers scale models by increasing training data. The article says frontier models could be on roughly ten times more data than Llama 3.1’s reported 15 trillion tokens, while the supply of easily available internet data may eventually become limited, possibly as early as the 2030s. That timeline is an estimate attributed to the article’s discussion and is not independently confirmed here.

More efficient learning could broaden access to AI development. The report quotes David Samuel, one of GPT-BERT’s architects, saying that Czech and Norwegian have much less training data available than English, while languages such as Sami may have only tens of millions of tokens. If models could learn more effectively from small datasets, universities and communities working in under-resourced languages might be able to build more capable systems without the resources required for hyperscale training. MIT Technology Review presents this as a potential benefit, not a demonstrated outcome.

The research could also matter for multimodal AI. MIT Technology Review reports that adding visual data to BabyLM systems has not yet produced the expected gains, while existing models trained on children’s footage have learned only simple word-object associations. If researchers identify how children combine vision, hearing, movement, attention, and language, those findings could inform systems trained on video and other sensory data. The article does not report a proven method that has already closed the gap.

There is a scientific value beyond model efficiency. The report describes researchers using language models as imperfect experimental stand-ins for language users. Scientists can impose conditions that would be difficult or unethical to impose on children, such as limiting exposure to particular grammatical forms or simulating different degrees of bilingualism. MIT Technology Review says such experiments could test ideas about whether language depends on innate structure, general learning constraints, or environmental experience. These comparisons remain limited because brains are embodied, develop continuously, and contain biological mechanisms that language models do not share.

The central evidence supports a narrower conclusion than claims that AI is approaching human learning. MIT Technology Review reports that models can learn syntax and perform well on selected benchmarks, but also says baby-scale systems remain clunky, many cannot generate text, and current video-trained models are far from childlike. The article does not establish that any existing architecture learns language with human-level efficiency, nor that insights from models will settle longstanding disputes in linguistics or developmental psychology.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

System Requirements:
Best ArchitecturePure RAGRecommended pattern
Hallucination RiskVery LowGrounding efficacy
Update Cost$0 (Vector sync)Ongoing maintenance
Core takeaway: Fine-tuning teaches models how to speak (form, style, syntax); RAG teaches models what to say (verifiable facts). Never use fine-tuning alone for factual memory.
互動式概念檢查+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

接下來看什麼

Watch whether multimodal, interactive, and socially informed training methods improve performance on child-scale data. Also watch whether new datasets such as long-term recordings of children’s lives produce measurable advances, and whether researchers can distinguish useful cognitive insights from comparisons between fundamentally different systems.

The first area to watch is whether models trained on richer sensory data can move beyond object-word associations. MIT Technology Review reports that SAYCam-based systems learned simple words such as “ball” and “cat,” but did not become broadly capable language users. Future evaluations should therefore test grammar, grounding, , interaction, and learning over time rather than relying on isolated recognition results.

Researchers will also need to determine whether active learning produces measurable gains. The article describes children as selecting experiences, experimenting, and seeking information that fills knowledge gaps, but says early social-model experiments did not beat standard models. A meaningful advance would require controlled comparisons showing which forms of exploration or interaction improve learning under the same data and compute limits.

Longitudinal datasets may change the evidence base. MIT Technology Review reports that Hasson’s project recorded 17 children across their first 1,000 days, creating a much larger view of early experience than earlier headcam projects. Important unknowns include how representative the participating families are, how the recordings are annotated, what privacy safeguards govern access, and whether the data can be used to train models without reducing childhood experience to a narrow set of measurable signals. The source does not answer these questions.

BabyLM and related benchmarks should be followed for reproducibility and scope. The article reports a strong result for GPT-BERT on one against Llama 2 70B, but also makes clear that this did not make GPT-BERT generally comparable to a modern commercial model. Readers should watch whether results hold across languages, datasets, model sizes, and evaluation tasks, and whether gains come from genuinely better learning methods or from benchmark-specific design choices.

Finally, the field’s practical priorities may determine whether child-inspired research spreads. MIT Technology Review reports that frontier labs are not broadly racing to copy developmental psychology, while Meta has shown particular interest in video and child headcam research. It remains unknown whether companies will publish enough methods and results to permit independent scrutiny, whether small-language communities will benefit directly, and whether cognitive findings derived from AI models will withstand evidence from real children and developmental studies.

相關指引和測驗

人工智慧模型解釋人工智慧培訓ChatGPT 與大型語言模型AI 的未來測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語關注 AI 模型發布追蹤器
覺得有用嗎?