Preprint proposes structured memory for medical AI agents
MSM-Mem gives medical AI agents semantic, episodic and visual memory so they can retrieve prior clinical experiences and refine future reasoning, according to a new arXiv preprint.
Updated daily402 verified stories
Material AI papers, studies, datasets, evaluations, and scientific advances with direct links to the underlying research.
Every story links to the strongest available evidence: original sources when available, otherwise clearly attributed reporting.
What happened, why it matters, and what to watch — no jargon tax.
When the signal is thin, we publish nothing rather than padding the feed.
A growing stream of verified perspectives for people who need to understand AI without chasing hype.
MSM-Mem gives medical AI agents semantic, episodic and visual memory so they can retrieve prior clinical experiences and refine future reasoning, according to a new arXiv preprint.
A new arXiv preprint reports that reinforcement learning on factual data containing no personal information made memorized email addresses easier to extract from tested language models.
A new paper introduces BanglaVeilGuard, a benchmark and lightweight prompt guard designed to test Bangla-language models across standard, Romanized, code-mixed, noisy and dialectal forms. The authors report sharply lower attack success in their evaluation, but also find that the guard over-refuses some benign…
Technology.org reports that Britain and Ukraine signed a defense-AI partnership giving British researchers access to Ukraine’s Avengers AI Labs and its battlefield data platform. The report says the first projects will focus on military-facility sensing and low-power chips for drones, though the claims are not…
A new arXiv benchmark evaluates chemistry-focused large language models across varied instructions, molecular representations and eight task categories, reporting substantial prompt sensitivity, representation dependence and uneven performance.
A preprint introduces Evidence-State Reliability, a measure for testing whether intermediate evidence remains usable as language-model pipelines process degraded inputs. In one controlled evaluation using GLM-5.2, parser-valid outputs persisted while stage success declined.
A new arXiv study evaluates 53 language models across 11 safety datasets and reports that model strengths vary sharply by harm category, while conversational safety remains unresolved.
ToSCA proposes a hierarchical reinforcement-learning framework that separates strategic choices from token-level response generation. The authors report improved strategy selection and response quality against several baselines in daily and emotional-support conversations.
An arXiv study reports that language-model systems can enter a failure pattern when search retrieves documents they previously generated. In the authors’ simulations, 79.6% of 1,528 runs ended in what the paper calls “RAG collapse.”
A new arXiv paper identifies a gap in LLM unlearning: an agent may stop recalling information from its weights but recover it through web search, retrieval or database tools. The authors propose a two-stage method to reduce both forms of recovery while preserving legitimate tool use.
A new arXiv paper proposes replacing raw retrieved passages with compact, keyword-grounded facts to reduce the risk that prompt injections cause retrieval-augmented generation systems to expose sensitive database contents.
An arXiv preprint introduces MCite-RL, a framework that uses iterative retrieval, reasoning, recursive image cropping and citation-focused reinforcement learning to improve both multimodal answers and the accuracy of their visual evidence links.
One useful briefing each week
Get the week’s verified AI news, original data, useful tools, learning picks, and fresh AI jobs.
Hiring an AI professional or launching a useful AI product? Put it in front of people who came here to learn and act.
Post an AI job Submit an AI tool