SKILL.state proposes explicit execution state for longer-running AI agents
A paper accepted at EMNLP presents a runtime architecture that replaces growing agent conversation histories with mutable structured execution state.
Updated daily1517 verified stories
Source-checked AI coverage of product launches, policy shifts, safety research, and industry moves, explained in plain English by a nonprofit education team.
Every story links to the strongest available evidence: original sources when available, otherwise clearly attributed reporting.
What happened, why it matters, and what to watch — no jargon tax.
When the signal is thin, we publish nothing rather than padding the feed.
A growing stream of verified perspectives for people who need to understand AI without chasing hype.
A paper accepted at EMNLP presents a runtime architecture that replaces growing agent conversation histories with mutable structured execution state.
A new Chinese-language benchmark evaluates large language models by the decisions and evidence states they face in financial risk workflows, rather than by general capability scores alone.
A new benchmark evaluates whether AI agents can turn raw biological data into research outputs across complete computational biology workflows. In 20 tasks, 13 frontier models scored between 0.00 and 0.48, with performance declining as datasets and sequences of analysis steps grew.
Researchers introduce FLARE, a system that pairs an LLM-based agent with the Lean proof assistant to check whether proposed mixed-integer linear programming reformulations preserve the original problem. On the paper’s 20-problem, 109-formulation benchmark, the authors report 100% accuracy on the NP-hard subset and…
Researchers propose a framework that compares AI systems with workplace tasks using shared cognitive-capability profiles, based on evaluations of six AI systems and task requirements gathered from 410 employees.
A new arXiv benchmark evaluates whether large language models can discover distinct crashes in open-source software without being given a predefined vulnerability target. In the authors’ tests, Claude Opus 4.8 triggered crashes in 60 of 77 challenges, but achieved only 196 of a possible 579 points.
OpenAI says it has launched commercial operations in Brazil, with a São Paulo team supporting businesses, developers, researchers and public institutions as the company expands ChatGPT, Codex and enterprise adoption.
A new benchmark of 11,586 bilingual, diagram-based physics problems reports that the strongest of 18 tested multimodal language models achieved only 33.7% answer accuracy.
A new arXiv benchmark evaluates LLM agents on geospatial planning using maps, tools and local social-media posts. Its authors report a sharp decline on complex tasks, with a 40.2% pass rate.
Socure raised $156 million at a reported $5.2 billion valuation and agreed to acquire AI fraud-investigation startup Fravity, according to Crunchbase News.
MiniMax says revenue from its Open Platform and other AI-based enterprise services rose 703.1% year over year to $73.9 million in the first half of 2026. The unaudited results also show higher gross profit but a wider adjusted net loss.
Ai2 says Providence Swedish will run its AutoDiscovery platform on protected cancer-research data after a breast-cancer signal was validated in a separate dataset and lab work.
One useful briefing each week
Get the week’s verified AI news, original data, useful tools, learning picks, and fresh AI jobs.
Hiring an AI professional or launching a useful AI product? Put it in front of people who came here to learn and act.
Post an AI job Submit an AI tool