Apple researchers report scaling law for training models with scarce data
A study of more than 2,000 language-model training runs says scarce target data can be repeated 15–20 times in mixtures, with the best rate varying by scale and compute.
Updated daily396 verified stories
Material AI papers, studies, datasets, evaluations, and scientific advances with direct links to the underlying research.
Every story links to the strongest available evidence: original sources when available, otherwise clearly attributed reporting.
What happened, why it matters, and what to watch — no jargon tax.
When the signal is thin, we publish nothing rather than padding the feed.
A growing stream of verified perspectives for people who need to understand AI without chasing hype.
A study of more than 2,000 language-model training runs says scarce target data can be repeated 15–20 times in mixtures, with the best rate varying by scale and compute.
Apple researchers describe LINK, a pretraining intervention that replaces selected English words with word-level translations from a target language. The paper reports improvements across eight languages and five model sizes, including up to a twofold speedup in reaching equivalent downstream performance.
A position paper reports tacit collusion by DeepSeek-R1 agents in a simulated Bertrand pricing market, even after human prompts against collusion. It argues that observed-behavior certification should precede deployment of reasoning agents in economic markets; the evidence and safeguards remain preliminary.
A systematic review surveys how large language models are being studied for mental-health analysis, risk assessment, therapy support and multimodal monitoring, while stressing unresolved ethical and regulatory challenges.
An ICML 2026 position paper analyzing 500 Hugging Face model cards argues that open-weight foundation models need coordinated model cards, acceptable-use policies, and licenses to address safety and governance gaps.
A position paper proposes the Transition Complexity Profile, a standardized way to describe how unpredictable and long-range the dynamics of game environments are for game-world modeling and reinforcement learning.
A new arXiv benchmark places 15 language-model agents in a 20-year football-management simulation, testing whether they can make consistent decisions when short-term choices affect long-term outcomes.
A new arXiv benchmark reports that changing only the retrieval method raised a fixed model’s exact accuracy on financial reconciliation cases from 2.05% to 72.44%. The study also finds that a correct root-cause label often does not mean the system returned sufficient evidence for an auditable diagnosis.
A new arXiv study presents scaling-law experiments for text-to-image diffusion models across compute budgets from 10^19 to 10^22 FLOPs. Its authors report that image models need substantially more data per parameter than language models to train efficiently.
A new arXiv paper describes a self-supervised EEG method that reports 92.4% accuracy distinguishing Alzheimer’s disease from cognitively normal controls on the ADFTD cohort.
An arXiv paper introduces Agentic ESOpt, a proposed evolution-strategy framework for fine-tuning long-horizon language-model agents with inference-level GPU memory. The authors report gains for Qwen-3.5-27B on WebArena-Lite and improvements in 28 of 36 prompt-optimization settings.
An arXiv paper argues that mean squared error can misjudge irregular time-series forecasts and proposes a continuous-time metric tested across synthetic, semi-synthetic and eight real-world datasets.
One useful briefing each week
Get the week’s verified AI news, original data, useful tools, learning picks, and fresh AI jobs.
Hiring an AI professional or launching a useful AI product? Put it in front of people who came here to learn and act.
Post an AI job Submit an AI tool