Preprint tests whether LLM agents know when to remember, verify or ask
A new benchmark evaluates whether language-model agents correctly decide when interaction-derived information should be saved, checked, used temporarily or clarified with a user.
Aggiornato quotidianamente1795 storie verificate
Copertura controllata dall'IA su lanci di prodotti, cambiamenti di politiche, ricerche sulla sicurezza e mosse del settore, spiegata in un linguaggio semplice da un team educativo non profit.
Ogni articolo è collegato alle prove più forti disponibili: fonti originali quando disponibili, altrimenti segnalazioni chiaramente attribuite.
What happened, why it matters, and what to watch — without the jargon.
Quando il segnale è debole, non pubblichiamo nulla invece di riempire il feed.
Source-checked AI stories, newest first, for people who need to understand AI without chasing hype.
A new benchmark evaluates whether language-model agents correctly decide when interaction-derived information should be saved, checked, used temporarily or clarified with a user.
A new evaluation framework for watermarking schemes in large language models is proposed, focusing on cross-lingual fairness.
Stanford HAI’s 2026 AI Index says national AI strategies are spreading, while data-localization rules, state-backed computing capacity and public investment remain uneven across regions.
A Together AI analysis of 904 DeepSWE rollouts reports that OpenAI's GPT-5.6 Sol leads on first-attempt coding accuracy while the open-weight GLM-5.3 leads once retries are allowed, at roughly half the cost per attempt.
An arXiv preprint proposes an AI generative model and watermarking method for more reliable multivariate time-series data after editing. Authors report tests across finance, energy and neuroimaging benchmarks, but the abstract gives no numerical results or evidence of deployment.
Natural language code retrieval is a rapidly evolving task in computer science. However, the 1C:Enterprise ecosystem combines Russian syntax with highly domain-specific terminology, for which open datasets and specialized models have been virtually non-existent.
Large language models (LLMs) and agentic AI systems are increasingly being explored for domain-specific maintenance and prognostics tasks, raising the question of whether they can effectively support prognostics and health management (PHM).
A new arXiv preprint introduces NepOOC, a 1,090-pair Nepali-English benchmark for detecting misleading captions attached to authentic images. On this dataset, a text-only mBERT model matched the best tested multimodal system, while image-only models performed near chance.
A new arXiv preprint tests financial named-entity recognition across filings, news and social media, finding that some confidence measures remain useful under domain shift while others deteriorate sharply.
A study explores whether conversational AI can shift how majority-group members categorize and relate to Latine immigrants.
A new arXiv preprint reports that injecting a new subject every few hundred tokens can increase language-model outputs’ judged surprise and connection. The effect did not produce integrated documents or improve the best result in an online bin-packing test, exposing limits in how long generations are evaluated.
A preliminary arXiv study finds that responses from language models have become less diverse across open-ended creativity tasks, raising questions about homogenization in human-AI creative work.
Un briefing utile ogni settimana
Ricevi le notizie verificate sull'IA della settimana, dati originali, strumenti utili, scelte per l'apprendimento e nuovi lavori in IA.
Assumere un professionista dell'IA o lanciare un prodotto AI utile? Mettilo davanti a chi è venuto qui per imparare e agire.
Pubblica un lavoro AI Invia uno strumento AI