已驗證來源
每個故事都連結到最有力的可用證據:可用的原始來源,否則明確歸因的報告。
簡單的英語
發生了什麼、為什麼重要以及值得關注的內容——無需行話。
無填料
當訊號很弱時,我們不會發布任何內容,而是填充提要。
更多故事
9 故事企業
Bridging Search and CRM: Productionizing AI Product Research Agents for Customer Re-Engagement
Modern e-commerce platforms often operate search, recommendation, personalization, and CRM systems independently, limiting opportunities for proactive customer re-engagement.arxiv.org創新
New benchmark shows Vietnam’s exam rubric can change how language models rank
A paper introduces THPT-Ladder, a 632-item benchmark that applies Vietnam’s 2025 national exam grading scheme to language models and reports materially different scores from standard proportional-accuracy measures.arxiv.org創新
Netflix paper outlines a lifecycle for LLM judges evaluating recommendation explanations
An arXiv paper describes how Netflix built, deployed and continuously monitored an LLM judge for recommendation explanations, reporting viewing and engagement gains in a five-week A/B test involving tens of millions of members.arxiv.org創新
调查将自我进化的人工智能代理框架描述为动态图转换
一項新的 arXiv 調查建議將人工智慧代理的記憶、工具、技能、工作流程和關係視為隨時間變化的圖表,並要求進行圖表感知評估和治理。arxiv.org創新
FACET proposes an environment-grounded method for training terminal agents
A new arXiv preprint presents FACET, a framework for generating executable terminal tasks whose instructions, environments, solutions and verifiers are designed to remain consistent.arxiv.org創新
SESSE proposes structured decomposition to make LLM judging more auditable
An arXiv preprint introduces SESSE, a training-free framework that breaks an LLM judge’s preference into sub-questions. The authors report near-parity with a chain-of-thought baseline on 1,000 RewardBench examples and criterion-level vote records; generalization, cost, and independent validation remain open questions.arxiv.org創新
IBM team reports AI-written adapters bring thousands of Hugging Face models to Spyre
An IBM Spyre team says coding agents helped create 13 runtime adapters that covered 7,960 of the 10,000 most-downloaded Hugging Face embedding models in its target set, with 6,804 passing end-to-end tests on Spyre. The team says human debugging remained essential.pytorch.org創新
Apple researchers report iterative pseudo-labeling gains for Mandarin-English speech recognition
An Apple research paper describes a three-phase iterative pseudo-labeling method for Mandarin-English code-switching automatic speech recognition and reports Mix Error Rate reductions on two SEAME development subsets.machinelearning.apple.com產品展示
Google to let users tune Discover with natural-language requests
Google says users will soon be able to tell Discover what topics and links they want to see more or less of, while new controls also personalize Search and Google News audio briefings.
blog.google