በየቀኑ ይዘምናል።2273 የተረጋገጡ ታሪኮች
AI ዜና. ያለ ጫጫታ።
በምንጭ የተረጋገጠ የኤአይአይ ሽፋን የምርት ጅምር፣ የፖሊሲ ሽግግሮች፣ የደህንነት ጥናት እና የኢንዱስትሪ እንቅስቃሴዎች፣ ለትርፍ ያልተቋቋመ የትምህርት ቡድን በግልፅ እንግሊዝኛ ተብራርቷል።
የተረጋገጠ ምንጭ
እያንዳንዱ ታሪክ በጣም ጠንካራ ከሆኑ ማስረጃዎች ጋር ይገናኛል፡ ኦሪጅናል ምንጮች ሲገኝ፣ ያለበለዚያ በግልጽ ከተመዘገበ ሪፖርት።
ግልጽ እንግሊዝኛ
ምን እንደተከሰተ ፣ ለምን አስፈላጊ ነው ፣ እና ምን መታየት እንዳለበት - ያለ ጃርጎን።
ምንም መሙያ የለም።
ምልክቱ ቀጭን ሲሆን ምግቡን ከማሸግ ይልቅ ምንም ነገር አናተምም።
ተጨማሪ ታሪኮች
9 ታሪኮችፈጠራ
ቤንችማርክ አለ ከፍተኛ መልቲሞዳል ሞዴሎች ከብዕር ድምፅ እና የእጅ እንቅስቃሴ ቃላትን በማንበብ ከ10% በታች ውጤት ያስመዘገቡ
A new arXiv paper introduces a test in which models must infer a written word from pen-scratch audio and hand-movement video, with no ink visible. The authors report humans above 80% ordered letter accuracy and leading models below 10% — and that giving models both modalities often made results worse.arxiv.orgፈጠራ
ወኪል-አዋቂ መሸጎጫ አስተዳደር በብዙ ወኪል ማገልገል ላይ እስከ 45% የሚደርስ የመጀመሪያ-ቶከን መዘግየትን ይቀንሳል ይላል ወረቀት
አዲስ የarXiv ቅድመ-ህትመት CacheScoutን ይገልፃል ፣በክፍት ምንጭ vLLM አገልጋይ ላይ የተገነባው ንብርብር በአምሳያው ቁልፍ-እሴት መሸጎጫ ውስጥ ምን ማቆየት እንዳለበት የሚወስነው በሚቀጥለው የትኛው ወኪል እንደሚሰራ ነው። ደራሲዎቹ ባለ ሁለት አሃዝ መዘግየት እና የውጤት ግኝቶችን ሪፖርት አድርገዋል። የሥራ ጫናዎች፣ ሞዴሎች እና ሃርድዌር በአብስትራክት ውስጥ አልተገለጹም።arxiv.orgፈጠራ
የOpenAI በ AI የመነጨ ማረጋገጫ የተገላቢጦሽ ሁኔታ ሲያገኝ እና ጥገናን አትሟል።
Two researchers say a lemma proof in Chapter 6 of OpenAI's mathematics document has a polarity error: a test in terms of average success where the next step needs a large conditional failure. They give a counterexample and a corrected proof, and caution that this is not verification of the chapter's main theorem.arxiv.orgፈጠራ
የማባዛት ጥናት FLOPs አሁንም AI Runtimeን በተሳሳተ መንገድ ይተነብያሉ፣ እና የታቀደው ማስተካከያ በአዲሱ ሃርድዌር ላይ አልተሳካም ይላል።
የሁለት ተመራማሪዎች ቅድመ ህትመት ለምን እኩል የ FLOP ቆጠራዎች እኩል የአፈፃፀም ጊዜ ማለት አይደለም የሚለውን ቀደም ሲል የተደረገ ጥናትን ይደግማል። ዋናውን የይገባኛል ጥያቄ ያረጋግጣል ነገር ግን የ α-FLOPs እርማት ፎርሙላ በአጠቃላይ በአዲሱ ሃርድዌር ላይ ያለውን የሩጫ ጊዜ እንደሚገምተው ዘግቧል፣ ይህም ቀመሩ የማይይዘው ዝላይ እና ማወዛወዝን ያሳያል።arxiv.orgድርጅት
Benchmark Paper Finds Four Ways to Query Enterprise Data With LLMs All Score Under 26%
A new arXiv preprint pits four architectures for natural-language querying of enterprise databases against each other on a synthetic bilingual benchmark. None answered more than about a quarter of cases correctly, and the design that scored highest was not the safest or the cheapest.arxiv.orgፈጠራ
Paper Proposes Grading AI Security Agents Without Labels by Measuring Convergence to a Stronger Model
A new arXiv preprint argues security teams can judge whether a memory- or retrieval-equipped AI agent is learning by measuring how far it closes the gap to a stronger "teacher" model, rather than on labeled benchmarks that are often scarce or stale. Judging by a similarly powered model gave no usable signal.arxiv.orgፈጠራ
New Benchmark Tests Whether AI Assistants Can Remember a Year of Phone Use
A 17-author technical report posted to arXiv introduces MobileMem, a benchmark and framework for on-device long-term memory built from a year-scale collection of mobile experiences. The abstract describes the design but reports no scores, and key details about the underlying data remain undisclosed.arxiv.orgፈጠራ
Paper Reports Brain-Like Modular Organization Emerging Inside Large Language Models
A new arXiv preprint says large language models develop functionally specialized internal structure that lines up with distinct human brain networks, based on circuit analyses across 46 tasks in four cognitive domains. The abstract page leaves key methodological details unstated.arxiv.orgፈጠራ
Paper Finds Late Layers of a Mixture-of-Experts Model Tolerate Heavy Expert Masking
A preprint reports that disabling low-magnitude experts in the last five layers of a 35-billion-parameter Mixture-of-Experts model preserved far more usable code-translation outputs than spreading the same cuts across all layers. It covers one model and one benchmark, and the abstract reports no unmasked baseline.arxiv.org
በየሳምንቱ አንድ ጠቃሚ አጭር መግለጫ
በምግብ ውስጥ ሳይኖሩ ከ AI ጋር ይቀጥሉ.
የሳምንቱን የተረጋገጠ የኤአይ ዜና፣ ኦሪጅናል ውሂብ፣ ጠቃሚ መሳሪያዎችን፣ የመማሪያ ምርጫዎችን እና ትኩስ AI ስራዎችን ያግኙ።
AI የሚማሩ ሰዎችን ይድረሱ
የ AI ባለሙያ መቅጠር ወይም ጠቃሚ የ AI ምርት ማስጀመር? ለመማር እና ለመስራት ወደዚህ በመጡ ሰዎች ፊት አስቀምጠው።
AI ሥራ ይለጥፉየ AI መሳሪያ አስገባ