Оновлюється щодня2273 перевірені історії
Новини AI. Без шуму.
Перевірений джерелом штучний інтелект висвітлює випуск продуктів, зміни в політиці, дослідження безпеки та зміни в галузі, пояснені простою англійською мовою некомерційною освітньою командою.
Перевірене джерело
Кожна історія пов’язана з найвагомішими наявними доказами: оригінальними джерелами, якщо вони доступні, в іншому випадку чітко зазначеними повідомленнями.
Звичайна англійська
Що сталося, чому це важливо та що дивитися — без жаргону.
Без наповнювача
Коли сигнал слабкий, ми нічого не публікуємо, а лише доповнюємо канал.
Більше історій
9 історіїІнновація
Benchmark каже, що найкращі мультимодальні моделі набирають менше 10% результатів при читанні слів із звуків ручки та рухів руки
Нова стаття arXiv представляє тест, у якому моделі повинні вивести написане слово на основі аудіо, написаного ручкою, і відео рухів руки, без видимих чорнила. Автори повідомляють, що точність упорядкованих літер у людей перевищує 80%, а у провідних моделей – нижче 10%, і що використання моделей обох способів часто погіршує результати.arxiv.orgІнновація
Paper Says Agent-Aware Cache Management Cuts First-Token Delay Up to 45% in Multi-Agent Serving
A new arXiv preprint describes CacheScout, a layer built on the open-source vLLM server that decides what to keep in a model's key-value cache based on which agent is likely to run next. The authors report double-digit latency and throughput gains; the workloads, models, and hardware are not stated in the abstract.arxiv.orgІнновація
Аудит доказу, створеного штучним інтелектом OpenAI, знаходить зворотний стан і публікує ремонт
Два дослідники кажуть, що доказ леми в розділі 6 математичного документа OpenAI містить помилку полярності: тест з точки зору середнього успіху, де наступний крок потребує великої умовної невдачі. Вони наводять контрприклад і виправлений доказ і застерігають, що це не перевірка основної теореми розділу.arxiv.orgІнновація
Replication Study Says FLOPs Still Mispredict AI Runtime, and the Proposed Fix Fails on Newer Hardware
A preprint by two researchers reproduces an earlier study on why equal FLOP counts do not mean equal execution time. It confirms the underlying claim but reports that the α-FLOPs correction formula generally underestimates runtime on newer hardware, which shows jumps and oscillations the formula does not capture.arxiv.orgпідприємство
Benchmark Paper Finds Four Ways to Query Enterprise Data With LLMs All Score Under 26%
A new arXiv preprint pits four architectures for natural-language querying of enterprise databases against each other on a synthetic bilingual benchmark. None answered more than about a quarter of cases correctly, and the design that scored highest was not the safest or the cheapest.arxiv.orgІнновація
Paper Proposes Grading AI Security Agents Without Labels by Measuring Convergence to a Stronger Model
A new arXiv preprint argues security teams can judge whether a memory- or retrieval-equipped AI agent is learning by measuring how far it closes the gap to a stronger "teacher" model, rather than on labeled benchmarks that are often scarce or stale. Judging by a similarly powered model gave no usable signal.arxiv.orgІнновація
New Benchmark Tests Whether AI Assistants Can Remember a Year of Phone Use
A 17-author technical report posted to arXiv introduces MobileMem, a benchmark and framework for on-device long-term memory built from a year-scale collection of mobile experiences. The abstract describes the design but reports no scores, and key details about the underlying data remain undisclosed.arxiv.orgІнновація
Paper Reports Brain-Like Modular Organization Emerging Inside Large Language Models
A new arXiv preprint says large language models develop functionally specialized internal structure that lines up with distinct human brain networks, based on circuit analyses across 46 tasks in four cognitive domains. The abstract page leaves key methodological details unstated.arxiv.orgІнновація
Paper Finds Late Layers of a Mixture-of-Experts Model Tolerate Heavy Expert Masking
A preprint reports that disabling low-magnitude experts in the last five layers of a 35-billion-parameter Mixture-of-Experts model preserved far more usable code-translation outputs than spreading the same cuts across all layers. It covers one model and one benchmark, and the abstract reports no unmasked baseline.arxiv.org
Один корисний брифінг щотижня
Будьте в курсі ШІ, не живучи в стрічці.
Отримуйте перевірені новини про штучний інтелект, оригінальні дані, корисні інструменти, підбірки для навчання та нові вакансії зі штучного інтелекту.
Охопіть людей, які вивчають ШІ
Найняти професіонала ШІ чи запустити корисний продукт ШІ? Покажіть це людям, які прийшли сюди вчитися та діяти.
Опублікувати роботу ШІНадішліть інструмент ШІ