已驗證來源
每個故事都連結到最有力的可用證據:可用的原始來源,否則明確歸因的報告。
簡單的英語
發生了什麼、為什麼重要以及值得關注的內容——無需行話。
無填料
當訊號很弱時,我們不會發布任何內容,而是填充提要。
更多故事
9 故事創新
Apple researchers report scaling law for training models with scarce data
A study of more than 2,000 language-model training runs says scarce target data can be repeated 15–20 times in mixtures, with the best rate varying by scale and compute.machinelearning.apple.com創新
Apple Researchers Propose Lexical Substitutions to Improve Multilingual Model Training
Apple researchers describe LINK, a pretraining intervention that replaces selected English words with word-level translations from a target language. The paper reports improvements across eight languages and five model sizes, including up to a twofold speedup in reaching equivalent downstream performance.machinelearning.apple.com政策
立場文件呼籲在人工智慧代理做出市場決策之前進行認證
一份立場文件報告了 DeepSeek-R1 代理人在模擬 Bertrand 定價市場中的默契共謀,即使在人類提示反對共謀之後也是如此。它認為觀察行為認證應該先於經濟市場中推理代理人的部署;證據和保障措施仍處於初步階段。arxiv.org創新
Systematic review maps the growing use of large language models in mental health
A systematic review surveys how large language models are being studied for mental-health analysis, risk assessment, therapy support and multimodal monitoring, while stressing unresolved ethical and regulatory challenges.arxiv.org政策
Model Cards Alone May Not Govern Open-Weight Foundation Models, Position Paper Argues
An ICML 2026 position paper analyzing 500 Hugging Face model cards argues that open-weight foundation models need coordinated model cards, acceptable-use policies, and licenses to address safety and governance gaps.arxiv.org創新
A proposed metric would measure how difficult game worlds are to predict
A position paper proposes the Transition Complexity Profile, a standardized way to describe how unpredictable and long-range the dynamics of game environments are for game-world modeling and reinforcement learning.arxiv.org創新
FM-Bench tests whether AI agents can manage a football club for 20 years
A new arXiv benchmark places 15 language-model agents in a 20-year football-management simulation, testing whether they can make consistent decisions when short-term choices affect long-term outcomes.arxiv.org創新
FinRCA-Bench finds financial AI diagnosis depends heavily on evidence retrieval
A new arXiv benchmark reports that changing only the retrieval method raised a fixed model’s exact accuracy on financial reconciliation cases from 2.05% to 72.44%. The study also finds that a correct root-cause label often does not mean the system returned sufficient evidence for an auditable diagnosis.arxiv.org創新
Abra paper maps compute and data tradeoffs in diffusion image training
A new arXiv study presents scaling-law experiments for text-to-image diffusion models across compute budgets from 10^19 to 10^22 FLOPs. Its authors report that image models need substantially more data per parameter than language models to train efficiently.arxiv.org