Paper proposes adaptive graph-or-language handoffs for multi-agent LLMs
Routed Graph Handoff uses a lightweight language-model router to choose between structured dependency graphs and natural-language messages when AI agents delegate tasks.
Cập nhật hàng ngày1796 câu chuyện đã được xác minh
Dựa trên AI được kiểm tra nguồn về việc ra mắt sản phẩm, thay đổi chính sách, nghiên cứu an toàn và các động thái trong ngành, được giải thích bằng tiếng Anh đơn giản bởi một nhóm giáo dục phi lợi nhuận.
Mỗi câu chuyện đều liên kết với bằng chứng mạnh mẽ nhất hiện có: nguồn gốc khi có, các báo cáo được ghi nhận rõ ràng.
What happened, why it matters, and what to watch — without the jargon.
Khi tín hiệu yếu, chúng tôi không xuất bản gì ngoài việc đệm nguồn cấp dữ liệu.
Source-checked AI stories, newest first, for people who need to understand AI without chasing hype.
Routed Graph Handoff uses a lightweight language-model router to choose between structured dependency graphs and natural-language messages when AI agents delegate tasks.
A new benchmark evaluates whether language models can generate and improve GPU code for database-style queries, with the paper reporting speedups of up to 2.11× on one H100 GPU and 2.54× across four H100 GPUs.
A paper accepted at EMNLP 2026 reports that flipping routing-layer bits in mixture-of-experts language models can sharply inflate output length, potentially creating an availability and inference-cost risk.
An arXiv preprint proposes fixing a report’s evidence, numbers, directions and permitted wording before an LLM generates connective prose. The authors report higher cross-run reproducibility than a hybrid template in fMRI and randomized-trial reporting, alongside lower observed token use and median latency in one…
A newly submitted paper presents ClueWeaver, a reward-trained framework that divides long-narrative question answering between evidence retrieval and interpretation in compact, locally deployable language models.
A new arXiv paper proposes Golden-GRPO Injection, a three-stage self-learning framework intended to help large language models absorb newly injected knowledge and apply it across paraphrases, combined documents and reasoning tasks.
A new preprint reports that BanglaMamba, a state space model for Bangla fake-news detection, uses less GPU memory and processes data faster than tested BERT-based systems, while trailing BanglaBERT on accuracy and external-dataset performance.
A new preprint reports that direct messages and peer relays can shift the stances of language-model agents, and argues that evaluations should track exposure histories and actions, not text alone.
A new preprint reports that several normalized metrics used to compare language models across languages can introduce biases linked to tokenization, encoding and orthographic differences. It argues that sentence-level negative log-likelihood over semantically equivalent sequences produces more consistent comparisons.
A new dataset evaluates large language models as diagnostic agents in interactive clinical encounters rather than through isolated questions. MTDiag combines cases from emergency medicine, clinical records and published case reports, and proposes knowledge-grounded metrics beyond diagnostic accuracy.
Researchers created HealthBench-Psych, a 610-conversation mental-health subset of OpenAI’s HealthBench, and released its evaluation pipeline, model responses, grades and analysis code.
A survey accepted to Findings of EMNLP 2026 catalogs 80 methods for adapting foundation models without human labels, preference data, stronger teachers or executable verifiers. It warns that internal learning signals can either improve a model or recursively amplify its errors.
Mỗi tuần một buổi họp hữu ích
Nhận tin tức AI đã được xác minh trong tuần, dữ liệu gốc, công cụ hữu ích, lựa chọn học tập và các công việc AI mới.
Thuê một chuyên gia AI hay ra mắt một sản phẩm AI hữu ích? Hãy đặt nó trước những người đến đây để học hỏi và hành động.
Đăng tuyển dụng AI Gửi một công cụ AI