已验证来源
每个故事都链接到最有力的可用证据:可用的原始来源,否则明确归因的报告。
简单的英语
发生了什么、为什么重要以及值得关注的内容——无需行话。
无填料
当信号很弱时,我们不会发布任何内容,而是填充提要。
更多故事
9 故事创新
ArXiv 论文提出对长期 LLM 代理人进行基于里程碑的培训
在训练语言模型代理执行长期、多步骤任务时,MileGPO 使用里程碑发现和本地证据来改进信用分配。作者在 ALFWorld 和 WebShop 上报告了最先进的结果,但声明仍然仅限于论文的实验。arxiv.org创新
Preprint tests whether LLM agents know when to remember, verify or ask
A new benchmark evaluates whether language-model agents correctly decide when interaction-derived information should be saved, checked, used temporarily or clarified with a user.arxiv.org创新
审核语言模型水印中的跨语言公平性
提出了一种新的大语言模型中水印方案的评估框架,重点关注跨语言公平性。arxiv.org政策
Stanford AI Index finds AI policy expanding as sovereignty and investment diverge
Stanford HAI’s 2026 AI Index says national AI strategies are spreading, while data-localization rules, state-backed computing capacity and public investment remain uneven across regions.hai.stanford.edu创新
Together AI 基准测试:GLM-5.3 在第一次尝试时落后于 GPT-5.6 Sol,重试时以一半的价格获胜
对 904 个 DeepSWE 部署的 Together AI 分析报告称,OpenAI 的 GPT-5.6 Sol 在首次尝试编码精度方面领先,而开放权重 GLM-5.3 在允许重试后领先,每次尝试的成本约为每次尝试的一半。together.ai创新
Preprint proposes a locally tokenized AI model for robust time-series watermarking
An arXiv preprint proposes an AI generative model and watermarking method for more reliable multivariate time-series data after editing. Authors report tests across finance, energy and neuroimaging benchmarks, but the abstract gives no numerical results or evidence of deployment.arxiv.org创新
Natural Language Code Retrieval for 1C:Enterprise: An Open Benchmark and Efficient Bi-Encoder
Natural language code retrieval is a rapidly evolving task in computer science. However, the 1C:Enterprise ecosystem combines Russian syntax with highly domain-specific terminology, for which open datasets and specialized models have been virtually non-existent.arxiv.org创新
Time-Series Retrieval for Grounding Multimodal Language Models in Remaining Useful Life
Large language models (LLMs) and agentic AI systems are increasingly being explored for domain-specific maintenance and prognostics tasks, raising the question of whether they can effectively support prognostics and health management (PHM).arxiv.org创新
Nepali-English preprint: text-only AI matched multimodal model on out-of-context misinformation benchmark
A new arXiv preprint introduces NepOOC, a 1,090-pair Nepali-English benchmark for detecting misleading captions attached to authentic images. On this dataset, a text-only mBERT model matched the best tested multimodal system, while image-only models performed near chance.arxiv.org