Emelitere kwa ụbọchị2037 ezi akụkọ
Akụkọ AI. Enweghị mkpọtụ.
Mkpuchi AI nyochara isi mmalite nke mmalite ngwaahịa, mgbanwe amụma, nyocha nchekwa, na mmegharị ụlọ ọrụ, nke otu agụmakwụkwọ na-anaghị akwụ ụgwọ kọwara n'asụsụ bekee dị larịị.
Isi mmalite enwetara
Akụkọ ọ bụla na-ejikọta na ihe akaebe siri ike dị: isi mmalite mgbe ọ dị, ma ọ bụghị nke akọwapụtara nke ọma.
Bekee dị larịị
Gịnị mere, ihe mere o ji dị mkpa, na ihe na-ekiri - na-enweghị jargon.
Enweghị ndochi
Mgbe mgbaàmà ahụ dị gịrịgịrị, anyị na-ebipụta ihe ọ bụla kama ịkwanye ndepụta.
Akụkọ ndị ọzọ
9 akụkọIhe ohuru ohuru
Benchmark finds local AI agents can handle many hardware-design tool calls, but reliability varies
A new arXiv benchmark finds that open-source AI agents can complete many dependency-ordered hardware-design operations through MCP tools, while tool descriptions, context length and agent configuration strongly affect reliability.arxiv.orgIhe ohuru ohuru
SKILL.state proposes explicit execution state for longer-running AI agents
A paper accepted at EMNLP presents a runtime architecture that replaces growing agent conversation histories with mutable structured execution state.arxiv.orgIhe ohuru ohuru
FinRiskAtlas finds broad AI scores can miss weaknesses in financial risk review
A new Chinese-language benchmark evaluates large language models by the decisions and evidence states they face in financial risk workflows, rather than by general capability scores alone.arxiv.orgIhe ohuru ohuru
BixBench3 finds AI agents struggle with full-scale computational biology studies
A new benchmark evaluates whether AI agents can turn raw biological data into research outputs across complete computational biology workflows. In 20 tasks, 13 frontier models scored between 0.00 and 0.48, with performance declining as datasets and sequences of analysis steps grew.arxiv.orgIhe ohuru ohuru
FLARE uses an LLM and Lean to verify optimization reformulations
Researchers introduce FLARE, a system that pairs an LLM-based agent with the Lean proof assistant to check whether proposed mixed-integer linear programming reformulations preserve the original problem. On the paper’s 20-problem, 109-formulation benchmark, the authors report 100% accuracy on the NP-hard subset and…arxiv.orgIhe ohuru ohuru
A preprint proposes profiling AI and workplace tasks by cognitive capabilities
Researchers propose a framework that compares AI systems with workplace tasks using shared cognitive-capability profiles, based on evaluations of six AI systems and task requirements gathered from 410 employees.arxiv.orgNchekwa
FuzzingBrain-Bench tests whether LLMs can find unexpected software crashes
A new arXiv benchmark evaluates whether large language models can discover distinct crashes in open-source software without being given a predefined vulnerability target. In the authors’ tests, Claude Opus 4.8 triggered crashes in 60 of 77 challenges, but achieved only 196 of a possible 579 points.arxiv.orgỤlọ ọrụ mmepụta ihe
OpenAI opens commercial operations in Brazil as ChatGPT use expands
OpenAI says it has launched commercial operations in Brazil, with a São Paulo team supporting businesses, developers, researchers and public institutions as the company expands ChatGPT, Codex and enterprise adoption.
openai.comIhe ohuru ohuru
PhysElite paper reports that leading multimodal LLM solved 33.7% of Olympiad-level physics problems
A new benchmark of 11,586 bilingual, diagram-based physics problems reports that the strongest of 18 tested multimodal language models achieved only 33.7% answer accuracy.arxiv.org
Otu nkowa okwu bara uru kwa izu
Jigide AI na-ebighị na nri.
Nweta ozi AI enwetara nke izu, data izizi, ngwa bara uru, nhọrọ mmụta, yana ọrụ AI ọhụrụ.
Gakwuru ndị na-amụ AI
Ịnweta onye ọkachamara AI ma ọ bụ ịmalite ngwaahịa AI bara uru? Tinye ya n'ihu ndị bịara ebe a ịmụta na ime ihe.
Biputere ọrụ AINyefee ngwa AI