Cập nhật hàng ngày2273 câu chuyện đã được xác minh
Tin tức AI. Không có tiếng ồn.
Dựa trên AI được kiểm tra nguồn về việc ra mắt sản phẩm, thay đổi chính sách, nghiên cứu an toàn và các động thái trong ngành, được giải thích bằng tiếng Anh đơn giản bởi một nhóm giáo dục phi lợi nhuận.
Nguồn cung cấp đã được xác minh
Mỗi câu chuyện đều liên kết với bằng chứng mạnh mẽ nhất hiện có: nguồn gốc khi có, các báo cáo được ghi nhận rõ ràng.
Tiếng Anh đơn giản
Điều gì đã xảy ra, tại sao nó lại quan trọng và những gì cần xem — không có biệt ngữ.
Không có phụ
Khi tín hiệu yếu, chúng tôi không xuất bản gì ngoài việc đệm nguồn cấp dữ liệu.
Thêm câu chuyện
9 câu chuyệnĐổi mới
Benchmark Says Top Multimodal Models Score Under 10% at Reading Words From Pen Sounds and Hand Motion
A new arXiv paper introduces a test in which models must infer a written word from pen-scratch audio and hand-movement video, with no ink visible. The authors report humans above 80% ordered letter accuracy and leading models below 10% — and that giving models both modalities often made results worse.arxiv.orgĐổi mới
Paper Says Agent-Aware Cache Management Cuts First-Token Delay Up to 45% in Multi-Agent Serving
A new arXiv preprint describes CacheScout, a layer built on the open-source vLLM server that decides what to keep in a model's key-value cache based on which agent is likely to run next. The authors report double-digit latency and throughput gains; the workloads, models, and hardware are not stated in the abstract.arxiv.orgĐổi mới
Audit of an OpenAI AI-Generated Proof Finds a Reversed Condition, and Publishes a Repair
Two researchers say a lemma proof in Chapter 6 of OpenAI's mathematics document has a polarity error: a test in terms of average success where the next step needs a large conditional failure. They give a counterexample and a corrected proof, and caution that this is not verification of the chapter's main theorem.arxiv.orgĐổi mới
Replication Study Says FLOPs Still Mispredict AI Runtime, and the Proposed Fix Fails on Newer Hardware
A preprint by two researchers reproduces an earlier study on why equal FLOP counts do not mean equal execution time. It confirms the underlying claim but reports that the α-FLOPs correction formula generally underestimates runtime on newer hardware, which shows jumps and oscillations the formula does not capture.arxiv.orgDoanh nghiệp
Benchmark Paper Finds Four Ways to Query Enterprise Data With LLMs All Score Under 26%
A new arXiv preprint pits four architectures for natural-language querying of enterprise databases against each other on a synthetic bilingual benchmark. None answered more than about a quarter of cases correctly, and the design that scored highest was not the safest or the cheapest.arxiv.orgĐổi mới
Paper Proposes Grading AI Security Agents Without Labels by Measuring Convergence to a Stronger Model
A new arXiv preprint argues security teams can judge whether a memory- or retrieval-equipped AI agent is learning by measuring how far it closes the gap to a stronger "teacher" model, rather than on labeled benchmarks that are often scarce or stale. Judging by a similarly powered model gave no usable signal.arxiv.orgĐổi mới
New Benchmark Tests Whether AI Assistants Can Remember a Year of Phone Use
A 17-author technical report posted to arXiv introduces MobileMem, a benchmark and framework for on-device long-term memory built from a year-scale collection of mobile experiences. The abstract describes the design but reports no scores, and key details about the underlying data remain undisclosed.arxiv.orgĐổi mới
Bài báo Báo cáo Tổ chức mô-đun giống như bộ não đang nổi lên bên trong các mô hình ngôn ngữ lớn
Bản in trước arXiv mới cho biết các mô hình ngôn ngữ lớn phát triển cấu trúc bên trong chuyên biệt về mặt chức năng phù hợp với các mạng não riêng biệt của con người, dựa trên các phân tích mạch qua 46 nhiệm vụ trong bốn lĩnh vực nhận thức. Trang tóm tắt không nêu rõ các chi tiết chính về phương pháp luận.arxiv.orgĐổi mới
Bài viết tìm thấy các lớp muộn của mô hình hỗn hợp các chuyên gia chấp nhận việc che giấu chuyên gia nặng
Một bản in trước báo cáo rằng việc vô hiệu hóa các chuyên gia có cường độ thấp trong năm lớp cuối cùng của mô hình Hỗn hợp các chuyên gia gồm 35 tỷ tham số sẽ bảo toàn được nhiều đầu ra dịch mã có thể sử dụng được hơn so với việc trải rộng các phần cắt giống nhau trên tất cả các lớp. Nó bao gồm một mô hình và một điểm chuẩn, và các báo cáo trừu tượng không có đường cơ sở rõ ràng.arxiv.org
Mỗi tuần một buổi họp hữu ích
Theo dõi AI mà không sống trong nguồn tin.
Nhận tin tức AI đã được xác minh trong tuần, dữ liệu gốc, công cụ hữu ích, lựa chọn học tập và các công việc AI mới.
Tiếp cận những người đang học AI
Thuê một chuyên gia AI hay ra mắt một sản phẩm AI hữu ích? Hãy đặt nó trước những người đến đây để học hỏi và hành động.
Đăng tuyển dụng AIGửi một công cụ AI