Обновляется ежедневно2273 проверенные истории
Новости ИИ. Без шума.
Проверенные источники искусственного интеллекта освещают запуски продуктов, изменения в политике, исследования в области безопасности и отраслевые шаги, объясненные простым языком командой некоммерческого образования.
Проверенный источник
Каждая история связана с самыми убедительными имеющимися доказательствами: первоначальными источниками, если таковые имеются, или явно приписываемыми сообщениями.
Простой английский
Что произошло, почему это важно и что смотреть — без жаргона.
Без наполнителя
Когда сигнал слабый, мы ничего не публикуем, а лишь дополняем канал.
Больше историй
9 историиИнновации
Benchmark Says Top Multimodal Models Score Under 10% at Reading Words From Pen Sounds and Hand Motion
A new arXiv paper introduces a test in which models must infer a written word from pen-scratch audio and hand-movement video, with no ink visible. The authors report humans above 80% ordered letter accuracy and leading models below 10% — and that giving models both modalities often made results worse.arxiv.orgИнновации
Paper Says Agent-Aware Cache Management Cuts First-Token Delay Up to 45% in Multi-Agent Serving
A new arXiv preprint describes CacheScout, a layer built on the open-source vLLM server that decides what to keep in a model's key-value cache based on which agent is likely to run next. The authors report double-digit latency and throughput gains; the workloads, models, and hardware are not stated in the abstract.arxiv.orgИнновации
Аудит OpenAI доказательства, сгенерированного искусственным интеллектом, обнаруживает обратное состояние и публикует исправление
Два исследователя говорят, что доказательство леммы в главе 6 математического документа OpenAI содержит ошибку полярности: тест с точки зрения среднего успеха, где следующий шаг требует большого условного провала. Они приводят контрпример и исправленное доказательство и предупреждают, что это не проверка основной теоремы главы.arxiv.orgИнновации
Replication Study Says FLOPs Still Mispredict AI Runtime, and the Proposed Fix Fails on Newer Hardware
A preprint by two researchers reproduces an earlier study on why equal FLOP counts do not mean equal execution time. It confirms the underlying claim but reports that the α-FLOPs correction formula generally underestimates runtime on newer hardware, which shows jumps and oscillations the formula does not capture.arxiv.orgПредприятие
Benchmark Paper Finds Four Ways to Query Enterprise Data With LLMs All Score Under 26%
A new arXiv preprint pits four architectures for natural-language querying of enterprise databases against each other on a synthetic bilingual benchmark. None answered more than about a quarter of cases correctly, and the design that scored highest was not the safest or the cheapest.arxiv.orgИнновации
Paper Proposes Grading AI Security Agents Without Labels by Measuring Convergence to a Stronger Model
A new arXiv preprint argues security teams can judge whether a memory- or retrieval-equipped AI agent is learning by measuring how far it closes the gap to a stronger "teacher" model, rather than on labeled benchmarks that are often scarce or stale. Judging by a similarly powered model gave no usable signal.arxiv.orgИнновации
New Benchmark Tests Whether AI Assistants Can Remember a Year of Phone Use
A 17-author technical report posted to arXiv introduces MobileMem, a benchmark and framework for on-device long-term memory built from a year-scale collection of mobile experiences. The abstract describes the design but reports no scores, and key details about the underlying data remain undisclosed.arxiv.orgИнновации
Paper Reports Brain-Like Modular Organization Emerging Inside Large Language Models
A new arXiv preprint says large language models develop functionally specialized internal structure that lines up with distinct human brain networks, based on circuit analyses across 46 tasks in four cognitive domains. The abstract page leaves key methodological details unstated.arxiv.orgИнновации
Paper Finds Late Layers of a Mixture-of-Experts Model Tolerate Heavy Expert Masking
A preprint reports that disabling low-magnitude experts in the last five layers of a 35-billion-parameter Mixture-of-Experts model preserved far more usable code-translation outputs than spreading the same cuts across all layers. It covers one model and one benchmark, and the abstract reports no unmasked baseline.arxiv.org
Один полезный брифинг каждую неделю
Идите в ногу с ИИ, не живя в ленте.
Получайте проверенные новости об искусственном интеллекте за неделю, оригинальные данные, полезные инструменты, обучающие материалы и свежие вакансии в области искусственного интеллекта.
Охватите людей, которые изучают ИИ
Нанимаете специалиста по искусственному интеллекту или запускаете полезный продукт в области искусственного интеллекта? Покажите это людям, которые пришли сюда учиться и действовать.
Опубликовать вакансию ИИОтправить инструмент искусственного интеллекта