What happened
MIT Technology Review reports on the data-efficiency gap between children learning language and large language models trained on enormous text collections. Researchers are using small-data language-model competitions, child headcam footage, and larger recordings of children’s early lives to investigate what machines may be missing.
MIT Technology Review reports that children typically begin producing grammatically correct sentences after hearing roughly 10 million words, with estimates reaching about 30 million at the high end. By contrast, modern large language models may process trillions of tokens during pretraining. The article describes this disparity as the “data efficiency gap”: children learn language from a small, embodied stream of experience, while models depend on vastly larger collections of text.
The report points to BabyLM, a competition organized by researchers including Alex Warstadt, Leshem Choshen, and others. MIT Technology Review says the main competition limits models to a developmentally plausible corpus of about 100 million words, with a toddler-scale track using 10 million. The data can include storybooks, dialogue, movie subtitles, Wikipedia, and transcripts of speech directed at children. Models are evaluated with grammar tasks similar to those used in psycholinguistics.
According to MIT Technology Review, the competition has challenged the assumption that curriculum learning—presenting simple material before complex material—would help models learn more like children. The approach was popular in the first round but did not perform as well as expected. The article says the 2024 leading system, GPT-BERT, combined next-token prediction with masked-language modeling and, after training on about 100 million words, exceeded Meta’s Llama 2 70B on one BabyLM benchmark. The report also stresses that these models remain much weaker than current commercial systems and do not reproduce a child’s general capabilities.
The article reports that text alone may be an inadequate approximation of childhood experience. Michael Frank’s SAYCam project recorded two hours a week of three babies’ lives between six months and two and a half years. MIT Technology Review says a model trained on 61 hours of this raw footage learned to identify objects and associate them with words, but did not produce the capabilities of a two-year-old. A newer project led by Uri Hasson recorded the first 1,000 days of 17 children’s lives, using cameras and microphones in living areas for about 12 hours a day on most days. The resulting dataset was described in a recent preprint, but the article does not provide enough information here to independently assess that preprint’s methods or findings.
MIT Technology Review reports that researchers are considering active exploration and social learning as additional missing ingredients. Alison Gopnik argues that children choose experiences, experiment, and seek predictable effects on the world, while Elizabeth Bonawitz says children reason about teachers and why information is being provided. The article says a BabyLM track allowing models to learn through interaction with other models did not outperform standard approaches. Meta researchers and academic collaborators have also announced a benchmark and challenge involving baby headcam footage, although the report does not establish that this effort has produced a successful child-scale model.
Read the primary source: technologyreview.com ↗
Why it matters
The research could affect how AI models are trained, especially for video and minority languages where large datasets are unavailable. It may also provide new experimental tools for studying language development, while the article emphasizes that current models remain far from reproducing a child’s broader abilities.
MIT Technology Review reports that data efficiency has become an engineering constraint as AI developers scale models by increasing training data. The article says frontier models could be pretraining on roughly ten times more data than Llama 3.1’s reported 15 trillion tokens, while the supply of easily available internet data may eventually become limited, possibly as early as the 2030s. That timeline is an estimate attributed to the article’s discussion and is not independently confirmed here.
More efficient learning could broaden access to AI development. The report quotes David Samuel, one of GPT-BERT’s architects, saying that Czech and Norwegian have much less training data available than English, while languages such as Sami may have only tens of millions of tokens. If models could learn more effectively from small datasets, universities and communities working in under-resourced languages might be able to build more capable systems without the resources required for hyperscale training. MIT Technology Review presents this as a potential benefit, not a demonstrated outcome.
The research could also matter for multimodal AI. MIT Technology Review reports that adding visual data to BabyLM systems has not yet produced the expected gains, while existing models trained on children’s footage have learned only simple word-object associations. If researchers identify how children combine vision, hearing, movement, attention, and language, those findings could inform systems trained on video and other sensory data. The article does not report a proven method that has already closed the gap.
There is a scientific value beyond model efficiency. The report describes researchers using language models as imperfect experimental stand-ins for language users. Scientists can impose conditions that would be difficult or unethical to impose on children, such as limiting exposure to particular grammatical forms or simulating different degrees of bilingualism. MIT Technology Review says such experiments could test ideas about whether language depends on innate structure, general learning constraints, or environmental experience. These comparisons remain limited because brains are embodied, develop continuously, and contain biological mechanisms that language models do not share.
The central evidence supports a narrower conclusion than claims that AI is approaching human learning. MIT Technology Review reports that models can learn syntax and perform well on selected benchmarks, but also says baby-scale systems remain clunky, many cannot generate text, and current video-trained models are far from childlike. The article does not establish that any existing architecture learns language with human-level efficiency, nor that insights from models will settle longstanding disputes in linguistics or developmental psychology.
What to watch next
Watch whether multimodal, interactive, and socially informed training methods improve performance on child-scale data. Also watch whether new datasets such as long-term recordings of children’s lives produce measurable advances, and whether researchers can distinguish useful cognitive insights from comparisons between fundamentally different systems.
The first area to watch is whether models trained on richer sensory data can move beyond object-word associations. MIT Technology Review reports that SAYCam-based systems learned simple words such as “ball” and “cat,” but did not become broadly capable language users. Future evaluations should therefore test grammar, grounding, generalization, interaction, and learning over time rather than relying on isolated recognition results.
Researchers will also need to determine whether active learning produces measurable gains. The article describes children as selecting experiences, experimenting, and seeking information that fills knowledge gaps, but says early social-model experiments did not beat standard models. A meaningful advance would require controlled comparisons showing which forms of exploration or interaction improve learning under the same data and compute limits.
Longitudinal datasets may change the evidence base. MIT Technology Review reports that Hasson’s project recorded 17 children across their first 1,000 days, creating a much larger view of early experience than earlier headcam projects. Important unknowns include how representative the participating families are, how the recordings are annotated, what privacy safeguards govern access, and whether the data can be used to train models without reducing childhood experience to a narrow set of measurable signals. The source does not answer these questions.
BabyLM and related benchmarks should be followed for reproducibility and scope. The article reports a strong result for GPT-BERT on one benchmark against Llama 2 70B, but also makes clear that this did not make GPT-BERT generally comparable to a modern commercial model. Readers should watch whether results hold across languages, datasets, model sizes, and evaluation tasks, and whether gains come from genuinely better learning methods or from benchmark-specific design choices.
Finally, the field’s practical priorities may determine whether child-inspired research spreads. MIT Technology Review reports that frontier labs are not broadly racing to copy developmental psychology, while Meta has shown particular interest in video and child headcam research. It remains unknown whether companies will publish enough methods and results to permit independent scrutiny, whether small-language communities will benefit directly, and whether cognitive findings derived from AI models will withstand evidence from real children and developmental studies.


