毎日更新1982 検証済みのストーリー
AIニュース。 ノイズなしで。
製品の発売、政策の変更、安全性研究、業界の動向についてソースチェックされた AI の報道が、非営利教育チームによって平易な英語で説明されます。
検証済みの調達
すべてのストーリーは、入手可能な最強の証拠、つまり入手可能な場合はオリジナルの情報源、それ以外の場合は明らかに帰属が明記された報道にリンクしています。
平易な英語
何が起こったのか、なぜそれが重要なのか、何を観るべきなのかを、専門用語を使わずに説明します。
フィラーなし
信号が薄い場合は、フィードをパディングするだけで何も公開しません。
さらに多くのストーリー
9 物語革新
Replication Study Says FLOPs Still Mispredict AI Runtime, and the Proposed Fix Fails on Newer Hardware
A preprint by two researchers reproduces an earlier study on why equal FLOP counts do not mean equal execution time. It confirms the underlying claim but reports that the α-FLOPs correction formula generally underestimates runtime on newer hardware, which shows jumps and oscillations the formula does not capture.arxiv.orgエンタープライズ
Benchmark Paper Finds Four Ways to Query Enterprise Data With LLMs All Score Under 26%
A new arXiv preprint pits four architectures for natural-language querying of enterprise databases against each other on a synthetic bilingual benchmark. None answered more than about a quarter of cases correctly, and the design that scored highest was not the safest or the cheapest.arxiv.org革新
Paper は、より強力なモデルへの収束を測定することにより、ラベルを使用せずに AI セキュリティ エージェントを評価することを提案しています
新しい arXiv プレプリントでは、セキュリティ チームは、不足または古いことが多いラベル付きのベンチマークではなく、より強力な「教師」モデルとのギャップをどの程度縮めるかを測定することで、メモリまたは検索機能を備えた AI エージェントが学習しているかどうかを判断できると主張しています。同様のパワーを備えたモデルから判断すると、使用可能な信号は得られませんでした。arxiv.org革新
New Benchmark Tests Whether AI Assistants Can Remember a Year of Phone Use
A 17-author technical report posted to arXiv introduces MobileMem, a benchmark and framework for on-device long-term memory built from a year-scale collection of mobile experiences. The abstract describes the design but reports no scores, and key details about the underlying data remain undisclosed.arxiv.org革新
Paper Reports Brain-Like Modular Organization Emerging Inside Large Language Models
A new arXiv preprint says large language models develop functionally specialized internal structure that lines up with distinct human brain networks, based on circuit analyses across 46 tasks in four cognitive domains. The abstract page leaves key methodological details unstated.arxiv.org革新
Paper Finds Late Layers of a Mixture-of-Experts Model Tolerate Heavy Expert Masking
A preprint reports that disabling low-magnitude experts in the last five layers of a 35-billion-parameter Mixture-of-Experts model preserved far more usable code-translation outputs than spreading the same cuts across all layers. It covers one model and one benchmark, and the abstract reports no unmasked baseline.arxiv.orgセキュリティ
CoreBreak の欠陥により、モデルが呼び出されなくてもエージェント ツールが実行される
Cloud Security Alliance の調査ノートには、Amazon Bedrock AgentCore、Google のエージェント開発キット、および Vercel の AI SDK ハーネス パッケージに存在する欠陥パターンである CoreBreak について説明されており、モデルをターンせずにツールを実行でき、モデル レベルのガードレールに検査対象が何も残らない状態になっています。labs.cloudsecurityalliance.orgポリシー
Anthropic Claude のテキスト透かしの仕組みの詳細
Anthropic によると、将来の Claude モデルには、EU AI 法に準拠するために、Google DeepMind の SynthID テキストに基づいた統計透かしが埋め込まれる予定です。同社は、文字、トークン、ユーザー ID は一切追加しておらず、完全な書き換えによって無効になると述べています。
anthropic.com産業
カーソルは、SpaceXによる買収が正式に終了したと発表
Cursor は、SpaceX が AI コーディング ツールの買収を完了し、SpaceXAI とのモデル トレーニング パートナーシップで 4 月に開始したとされるプロセスを完了したとの短い投稿を公開しました。この投稿では、いわゆる世界最大の GPU フリートへのアクセスを約束していますが、条件、スケジュール、製品の変更については明らかにしていません。
cursor.com
毎週 1 回の有益なブリーフィング
フィードに依存せずに AI に追いつきます。
今週の検証済み AI ニュース、オリジナル データ、便利なツール、おすすめの学習情報、最新の AI ジョブを入手します。
AI を学習している人々にリーチする
AI の専門家を雇いますか、それとも便利な AI 製品を立ち上げますか?それを学び、行動するためにここに来た人々の前に置きます。
AI の仕事を投稿するAIツールを提出する