已驗證來源
每個故事都連結到最有力的可用證據:可用的原始來源,否則明確歸因的報告。
簡單的英語
發生了什麼、為什麼重要以及值得關注的內容——無需行話。
無填料
當訊號很弱時,我們不會發布任何內容,而是填充提要。
更多故事
9 故事創新
Code agents lose reliability when code is rewritten without changing its meaning, study finds
An arXiv preprint reports that code agents can perform differently on semantically equivalent codebases, with effects varying by model, agent framework and benchmark.arxiv.org創新
Enhanced Fuzzy Logic Model for Power Transformer Fault Diagnosis Using IEEE Key Gas Method Improvements
This study presents an enhanced model combining Fuzzy Logic with the IEEE Key Gas Method (FL-KGM) that introduces refined membership functions, optimized fuzzy rule sets, and a novel separation of CO and CO2 to eliminate diagnostic inconsistencies.arxiv.org創新
New preprint reports gains from looped language models in multi-step tool calling
An arXiv study evaluates looped and conventional language models on three tool-calling benchmarks, reporting stronger results on workflows that require multiple dependent API calls and a potentially more efficient adaptive-computation approach.arxiv.org創新
Adversarial Review tests structured disagreement for agentic code review
A new arXiv paper proposes a three-agent code-review protocol in which a reviewer evaluates an agent’s code and a critic audits that review before edits are made. The authors report improved benchmark results over tested baselines, while also identifying false consensus as a failure mode.arxiv.org創新
ArXiv study reports large Roman Urdu hate-speech gains from LoRA adaptation
An arXiv preprint compares zero-shot and parameter-efficient fine-tuning for hate-speech detection in Roman Urdu. The authors report that LoRA adaptation raised F1 performance from 0.56 to above 0.93 on a corpus containing more than 72,000 annotated comments.arxiv.org安全性
Preliminary FraudBench test finds banking agents vulnerable to adaptive fraud
An arXiv paper introduces FraudBench, a benchmark for testing whether tool-using banking agents can detect fraud that unfolds across conversations. In a preliminary single-trial evaluation, four agents scored 49% to 65% on attack security.arxiv.org創新
Apple researchers report scaling law for training models with scarce data
A study of more than 2,000 language-model training runs says scarce target data can be repeated 15–20 times in mixtures, with the best rate varying by scale and compute.machinelearning.apple.com創新
Apple Researchers Propose Lexical Substitutions to Improve Multilingual Model Training
Apple researchers describe LINK, a pretraining intervention that replaces selected English words with word-level translations from a target language. The paper reports improvements across eight languages and five model sizes, including up to a twofold speedup in reaching equivalent downstream performance.machinelearning.apple.com政策
立場文件呼籲在人工智慧代理做出市場決策之前進行認證
一份立場文件報告了 DeepSeek-R1 代理人在模擬 Bertrand 定價市場中的默契共謀,即使在人類提示反對共謀之後也是如此。它認為觀察行為認證應該先於經濟市場中推理代理人的部署;證據和保障措施仍處於初步階段。arxiv.org