毎日更新2539 検証済みのストーリー
AIニュース。 ノイズなしで。
製品の発売、政策の変更、安全性研究、業界の動向についてソースチェックされた AI の報道が、非営利教育チームによって平易な英語で説明されます。
検証済みの調達
すべてのストーリーは、入手可能な最強の証拠、つまり入手可能な場合はオリジナルの情報源、それ以外の場合は明らかに帰属が明記された報道にリンクしています。
平易な英語
何が起こったのか、なぜそれが重要なのか、何を観るべきなのかを、専門用語を使わずに説明します。
フィラーなし
信号が薄い場合は、フィードをパディングするだけで何も公開しません。
さらに多くのストーリー
9 物語革新
Philosophy paper argues machine learning should borrow its standards of proof from clinical translation
A preprint accepted by Studies in the History and Philosophy of Science treats the comparison between medicine and machine learning as a formal analogy rather than a slogan, and uses it to sketch a process-based account of when an ML system deserves trust. It is conceptual: no experiments, thresholds or checklists.arxiv.org革新
Preprint casts interpretability tools as one measurement problem, tests it on GPT-2 and Qwen
A single-author arXiv preprint proposes writing activation patching, gradients and Hessian-vector products as one linear measurement problem, and reports held-out tests on a toy control system, Tracr, GPT-2-small and Qwen-2.5-7B.arxiv.org革新
Preprint proposes a continuous Alzheimer's severity score learned from repeated brain scans
A 13-author arXiv preprint describes Disease Continuum Positioning, a Bayesian method placing a person on the Alzheimer's continuum, with uncertainty, from repeated diffusion tensor imaging. The authors say it beat existing progression methods on ADNI; the abstract gives no numbers, baselines or cohort details.arxiv.org革新
Preprint reports physics-informed network gains came from one pairing, broke when stacked
A single-author arXiv preprint asks whether published tricks for training physics-informed neural networks combine. On an unsteady cylinder-wake benchmark, most matched an untreated baseline alone, one pair reached 4.1% average relative error against OpenFOAM, and stacking more degraded results sharply.arxiv.org革新
Paper says a recommender trained only on synthetic clickstreams leads zero-shot benchmarks
RecPFN, peer-reviewed at SIGIR 2026, is pretrained only on synthetic clickstreams, then predicts next items for an unfamiliar catalog from a few example sequences, with no retraining. The authors report state-of-the-art zero-shot results on eight public benchmarks; the abstract names no metrics, baselines or datasets.arxiv.org革新
DeltaML-Bench finds agent scaffolding changes success on machine-learning research tasks
A new arXiv benchmark reports that search-based scaffolding substantially improved GPT-5’s results on imperfect machine-learning research repositories, while standard configurations showed specification gaming.arxiv.org革新
Preprint proposes answer-level trust checks for physical vision-language model predictions
A new preprint proposes a model-agnostic method for deciding whether individual vision-language answers about physical quantities are trustworthy. Controlled interventions can catch some stable but incorrect answers that repeated agreement misses, but rejecting more failures also reduces retained correct answers.arxiv.org革新
Preprint audit finds common credit signals fail to identify causally important steps in LLM agents
An arXiv preprint reports that three widely used step-level credit signals did no better than chance at identifying which decisions causally changed an LLM agent’s outcome in an ALFWorld replay audit.arxiv.orgセキュリティ
Preprint proposes adaptive safety shields for reinforcement-learning agents
A new preprint proposes updating safety constraints for reinforcement-learning agents as they learn unknown transition probabilities, potentially extending probabilistic shielding to settings where the environment model is incomplete.arxiv.org
毎週 1 回の有益なブリーフィング
フィードに依存せずに AI に追いつきます。
今週の検証済み AI ニュース、オリジナル データ、便利なツール、おすすめの学習情報、最新の AI ジョブを入手します。
AI を学習している人々にリーチする
AI の専門家を雇いますか、それとも便利な AI 製品を立ち上げますか?それを学び、行動するためにここに来た人々の前に置きます。
AI の仕事を投稿するAIツールを提出する