已验证来源
每个故事都链接到最有力的可用证据:可用的原始来源,否则明确归因的报告。
简单的英语
发生了什么、为什么重要以及值得关注的内容——无需行话。
无填料
当信号很弱时,我们不会发布任何内容,而是填充提要。
更多故事
9 故事创新
研究发现,当代码被重写而不改变其含义时,代码代理就会失去可靠性
arXiv 预印本报告称,代码代理可以在语义等效的代码库上执行不同的操作,其效果因模型、代理框架和基准而异。arxiv.org创新
Enhanced Fuzzy Logic Model for Power Transformer Fault Diagnosis Using IEEE Key Gas Method Improvements
This study presents an enhanced model combining Fuzzy Logic with the IEEE Key Gas Method (FL-KGM) that introduces refined membership functions, optimized fuzzy rule sets, and a novel separation of CO and CO2 to eliminate diagnostic inconsistencies.arxiv.org创新
New preprint reports gains from looped language models in multi-step tool calling
An arXiv study evaluates looped and conventional language models on three tool-calling benchmarks, reporting stronger results on workflows that require multiple dependent API calls and a potentially more efficient adaptive-computation approach.arxiv.org创新
Adversarial Review tests structured disagreement for agentic code review
A new arXiv paper proposes a three-agent code-review protocol in which a reviewer evaluates an agent’s code and a critic audits that review before edits are made. The authors report improved benchmark results over tested baselines, while also identifying false consensus as a failure mode.arxiv.org创新
ArXiv 研究报告称,LoRA 适应使罗马乌尔都语仇恨言论大幅增加
arXiv 预印本比较了罗马乌尔都语仇恨言论检测的零样本和参数高效微调。作者报告说,在包含超过 72,000 条带注释的评论的语料库上,LoRA 适应将 F1 性能从 0.56 提高到 0.93 以上。arxiv.org安全
Preliminary FraudBench test finds banking agents vulnerable to adaptive fraud
An arXiv paper introduces FraudBench, a benchmark for testing whether tool-using banking agents can detect fraud that unfolds across conversations. In a preliminary single-trial evaluation, four agents scored 49% to 65% on attack security.arxiv.org创新
Apple 研究人员报告了稀缺数据训练模型的缩放定律
一项针对 2,000 多次语言模型训练运行的研究表明,稀缺目标数据可以混合重复 15-20 次,最佳速率因规模和计算而异。machinelearning.apple.com创新
Apple 研究人员提出词汇替换以改善多语言模型训练
Apple 研究人员描述了 LINK,这是一种预训练干预措施,可将选定的英语单词替换为目标语言的单词级翻译。该论文报告了八种语言和五种模型大小的改进,包括在达到同等下游性能方面加速了两倍。machinelearning.apple.com政策
立场文件呼吁在人工智能代理做出市场决策之前进行认证
一份立场文件报告了 DeepSeek-R1 代理人在模拟 Bertrand 定价市场中的默契共谋,即使在人类提示反对共谋之后也是如此。它认为观察行为认证应该先于经济市场中推理代理的部署;证据和保障措施仍处于初步阶段。arxiv.org