毎日更新2396 検証済みのストーリー
AIニュース。 ノイズなしで。
製品の発売、政策の変更、安全性研究、業界の動向についてソースチェックされた AI の報道が、非営利教育チームによって平易な英語で説明されます。
検証済みの調達
すべてのストーリーは、入手可能な最強の証拠、つまり入手可能な場合はオリジナルの情報源、それ以外の場合は明らかに帰属が明記された報道にリンクしています。
平易な英語
何が起こったのか、なぜそれが重要なのか、何を観るべきなのかを、専門用語を使わずに説明します。
フィラーなし
信号が薄い場合は、フィードをパディングするだけで何も公開しません。
さらに多くのストーリー
9 物語革新
Evaluating AI Agent Skill Performance with NVIDIA SkillEvaluator
NVIDIA SkillEvaluator measures the impact of verified skills on AI agent performance through a three-tier evaluation process.
developer.nvidia.com革新
Paper Introduces "Shadow Evaluations": AI Agents Did the Engineering but Failed Two Research Questions
A 24-author preprint had frontier AI agents attempt the central research questions of two unpublished NeurIPS 2026 submissions, then had the papers' own authors grade the results. The agents handled the engineering unaided over six days but were unambiguously rejected on the research.arxiv.org製品
Warp Opens Early Access to "Factories," a Config-as-Code System for Running Fleets of Coding Agents
Warp is taking early-access requests for Warp Factories, which defines fleets of coding agents as code — repos, models, permissions and human checkpoints in one YAML file, driven by CLI, API, SDK and MCP. Its automation and cost figures are vendor claims: no pricing, general availability date or independent testing.warp.dev革新
ベンチマークによると、上位のマルチモーダル モデルは、ペンの音と手の動きから単語を読み取る際のスコアが 10% 未満である
新しい arXiv 論文では、インクが見えない状態で、モデルがペンで擦る音声と手の動きのビデオから書かれた単語を推測するテストが紹介されています。著者らは、人間の順序文字精度は 80% 以上で、主要なモデルは 10% 以下であり、モデルに両方のモダリティを与えると結果が悪化することが多いと報告しています。arxiv.org革新
Paper Says Agent-Aware Cache Management Cuts First-Token Delay Up to 45% in Multi-Agent Serving
A new arXiv preprint describes CacheScout, a layer built on the open-source vLLM server that decides what to keep in a model's key-value cache based on which agent is likely to run next. The authors report double-digit latency and throughput gains; the workloads, models, and hardware are not stated in the abstract.arxiv.org革新
Audit of an OpenAI AI-Generated Proof Finds a Reversed Condition, and Publishes a Repair
Two researchers say a lemma proof in Chapter 6 of OpenAI's mathematics document has a polarity error: a test in terms of average success where the next step needs a large conditional failure. They give a counterexample and a corrected proof, and caution that this is not verification of the chapter's main theorem.arxiv.org革新
Replication Study Says FLOPs Still Mispredict AI Runtime, and the Proposed Fix Fails on Newer Hardware
A preprint by two researchers reproduces an earlier study on why equal FLOP counts do not mean equal execution time. It confirms the underlying claim but reports that the α-FLOPs correction formula generally underestimates runtime on newer hardware, which shows jumps and oscillations the formula does not capture.arxiv.orgエンタープライズ
Benchmark Paper Finds Four Ways to Query Enterprise Data With LLMs All Score Under 26%
A new arXiv preprint pits four architectures for natural-language querying of enterprise databases against each other on a synthetic bilingual benchmark. None answered more than about a quarter of cases correctly, and the design that scored highest was not the safest or the cheapest.arxiv.org革新
Paper は、より強力なモデルへの収束を測定することにより、ラベルを使用せずに AI セキュリティ エージェントを評価することを提案しています
新しい arXiv プレプリントでは、セキュリティ チームは、不足または古いことが多いラベル付きのベンチマークではなく、より強力な「教師」モデルとのギャップをどの程度縮めるかを測定することで、メモリまたは検索機能を備えた AI エージェントが学習しているかどうかを判断できると主張しています。同様のパワーを備えたモデルから判断すると、使用可能な信号は得られませんでした。arxiv.org
毎週 1 回の有益なブリーフィング
フィードに依存せずに AI に追いつきます。
今週の検証済み AI ニュース、オリジナル データ、便利なツール、おすすめの学習情報、最新の AI ジョブを入手します。
AI を学習している人々にリーチする
AI の専門家を雇いますか、それとも便利な AI 製品を立ち上げますか?それを学び、行動するためにここに来た人々の前に置きます。
AI の仕事を投稿するAIツールを提出する