已驗證來源
每個故事都連結到最有力的可用證據:可用的原始來源,否則明確歸因的報告。
簡單的英語
發生了什麼、為什麼重要以及值得關注的內容——無需行話。
無填料
當訊號很弱時,我們不會發布任何內容,而是填充提要。
更多故事
9 故事創新
DeltaML-Bench finds agent scaffolding changes success on machine-learning research tasks
A new arXiv benchmark reports that search-based scaffolding substantially improved GPT-5’s results on imperfect machine-learning research repositories, while standard configurations showed specification gaming.arxiv.org創新
Preprint proposes answer-level trust checks for physical vision-language model predictions
A new preprint proposes a model-agnostic method for deciding whether individual vision-language answers about physical quantities are trustworthy. Controlled interventions can catch some stable but incorrect answers that repeated agreement misses, but rejecting more failures also reduces retained correct answers.arxiv.org創新
Preprint audit finds common credit signals fail to identify causally important steps in LLM agents
An arXiv preprint reports that three widely used step-level credit signals did no better than chance at identifying which decisions causally changed an LLM agent’s outcome in an ALFWorld replay audit.arxiv.org安全性
Preprint 為強化學習代理提出了自適應安全防護罩
一份新的預印本建議在強化學習代理學習未知的轉移機率時更新其安全約束,從而可能將機率屏蔽擴展到環境模型不完整的設定。arxiv.org創新
Paper outlines assurance path for an onboard ML helicopter-weight estimator
A new arXiv preprint describes an LSTM-based supervised model for estimating helicopter weight during takeoff and an assurance process aimed at running it on legacy airborne computers.arxiv.org創新
DeltaMomentum paper proposes direction-aware optimizer updates for neural-network training
An arXiv preprint introduces DeltaMomentum, an optimizer update designed to forget frequently and rarely seen gradient directions at different rates, reporting faster training across language, image and vision benchmarks.arxiv.org創新
Transformer study estimates days before severe COPD flare-ups from home-ventilator data
An arXiv paper describes a two-stage transformer that uses seven days of home-ventilator pressure and flow waveforms to identify high risk of severe AECOPD and estimate the days remaining before an event. The authors report strong results, but the source does not establish clinical deployment or external validation.arxiv.org創新
Preprint proposes a two-hemisphere architecture for continual learning
An arXiv preprint proposes 4MAS, a neural-model architecture combining asymmetric modules, memory mechanisms, experience replay and sleep-like consolidation to address catastrophic forgetting.arxiv.org創新
Paper proposes using an LLM to generate tabular anomaly detectors from normal data
An arXiv paper introduces LLM-Detector, a prompt-based method that uses an LLM to synthesize anomaly-scoring code from statistical summaries, causal dependencies and prototypes in normal tabular data. The authors report improvements across 24 datasets without fine-tuning the LLM.arxiv.org