SKILL.state proposes explicit execution state for longer-running AI agents
A paper accepted at EMNLP presents a runtime architecture that replaces growing agent conversation histories with mutable structured execution state.
Inasasishwa kila siku1529 Hadithi zilizothibitishwa
Chanjo ya AI iliyokaguliwa na chanzo ya uzinduzi wa bidhaa, mabadiliko ya sera, utafiti wa usalama, na hatua za tasnia, zilizofafanuliwa kwa Kiingereza wazi na timu ya elimu isiyo ya faida.
Kila hadithi inaunganisha na ushahidi wenye nguvu zaidi unaopatikana: vyanzo vya asili vinapopatikana, vinginevyo vinahusishwa wazi.
Nini kilitokea, kwa nini ni muhimu, na nini cha kutazama - hakuna ushuru wa jargon.
Wakati ishara ni nyembamba, hatuchapishi chochote badala ya kuweka malisho.
Mtiririko unaokua wa mitazamo iliyothibitishwa kwa watu wanaohitaji kuelewa AI bila kufuata hype.
A paper accepted at EMNLP presents a runtime architecture that replaces growing agent conversation histories with mutable structured execution state.
A new Chinese-language benchmark evaluates large language models by the decisions and evidence states they face in financial risk workflows, rather than by general capability scores alone.
A new benchmark evaluates whether AI agents can turn raw biological data into research outputs across complete computational biology workflows. In 20 tasks, 13 frontier models scored between 0.00 and 0.48, with performance declining as datasets and sequences of analysis steps grew.
Researchers introduce FLARE, a system that pairs an LLM-based agent with the Lean proof assistant to check whether proposed mixed-integer linear programming reformulations preserve the original problem. On the paper’s 20-problem, 109-formulation benchmark, the authors report 100% accuracy on the NP-hard subset and…
Researchers propose a framework that compares AI systems with workplace tasks using shared cognitive-capability profiles, based on evaluations of six AI systems and task requirements gathered from 410 employees.
A new arXiv benchmark evaluates whether large language models can discover distinct crashes in open-source software without being given a predefined vulnerability target. In the authors’ tests, Claude Opus 4.8 triggered crashes in 60 of 77 challenges, but achieved only 196 of a possible 579 points.
OpenAI says it has launched commercial operations in Brazil, with a São Paulo team supporting businesses, developers, researchers and public institutions as the company expands ChatGPT, Codex and enterprise adoption.
A new benchmark of 11,586 bilingual, diagram-based physics problems reports that the strongest of 18 tested multimodal language models achieved only 33.7% answer accuracy.
A new arXiv benchmark evaluates LLM agents on geospatial planning using maps, tools and local social-media posts. Its authors report a sharp decline on complex tasks, with a 40.2% pass rate.
Socure raised $156 million at a reported $5.2 billion valuation and agreed to acquire AI fraud-investigation startup Fravity, according to Crunchbase News.
MiniMax says revenue from its Open Platform and other AI-based enterprise services rose 703.1% year over year to $73.9 million in the first half of 2026. The unaudited results also show higher gross profit but a wider adjusted net loss.
Ai2 says Providence Swedish will run its AutoDiscovery platform on protected cancer-research data after a breast-cancer signal was validated in a separate dataset and lab work.
Muhtasari mmoja muhimu kila wiki
Pata habari za AI zilizothibitishwa za wiki, data asili, zana muhimu, chaguo za kujifunza na kazi mpya za AI.
Kuajiri mtaalamu wa AI au kuzindua bidhaa muhimu ya AI? Iweke mbele ya watu waliokuja hapa kujifunza na kuchukua hatua.
Chapisha kazi ya AI Peana zana ya AI