검증된 출처
모든 기사는 가장 강력한 증거와 연결됩니다: 원가 자료가 있을 경우 그렇고, 그 외에는 명확히 출처가 명확히 기재된 보도입니다.
평범한 영어
무슨 일이 일어났는지, 왜 중요한지, 무엇을 보아야 하는지 등을 전문 용어 없이 설명합니다.
채우기 없는 이야기
신호가 얇을 때는 피드를 채우기보다는 아무것도 발행하지 않습니다.
더 많은 이야기
9 이야기들혁신
코드 에이전트는 의미를 변경하지 않고 코드를 재작성할 때 신뢰성을 잃는다고 연구가 밝혔습니다
arXiv 프리프린트는 코드 에이전트가 의미적으로 동등한 코드베이스에서 다르게 동작할 수 있으며, 그 효과는 모델, 에이전트 프레임워크, 벤치마크에 따라 다르다고 보고합니다.arxiv.org혁신
IEEE 핵심 가스 방법 개선을 이용한 전력 변압기 고장 진단을 위한 향상된 퍼지 논리 모델
본 연구는 퍼지 논리와 IEEE Key Gas 방법(FL-KGM)을 결합한 향상된 모델을 제시하며, 정제된 소속 함수, 최적화된 퍼지 규칙 집합, 그리고 진단 불일치를 제거하기 위한 CO와 CO2의 새로운 분리를 도입합니다.arxiv.org혁신
새로운 프리프린트 보고서는 다중 단계 도구 호출에서 반복 언어 모델의 이점을 얻었습니다
arXiv 연구는 세 가지 도구 호출 벤치마크에서 반복 및 기존 언어 모델을 평가하며, 여러 의존 API 호출과 잠재적으로 더 효율적인 적응형 계산 접근법이 필요한 워크플로우에서 더 강력한 결과를 보고했습니다.arxiv.org혁신
대립적 검토 테스트는 에이전트 코드 검토를 위한 구조화된 불일치
새로운 arXiv 논문은 리뷰어가 에이전트의 코드를 평가하고 비평가가 편집 전에 검토하는 감사를 하는 3인 에이전트 코드 검토 프로토콜을 제안합니다. 저자들은 테스트된 베이스라인 대비 벤치마크 결과가 개선되었다고 보고하며, 허위 합의를 실패 모드로 식별했습니다.arxiv.org혁신
ArXiv study reports large Roman Urdu hate-speech gains from LoRA adaptation
An arXiv preprint compares zero-shot and parameter-efficient fine-tuning for hate-speech detection in Roman Urdu. The authors report that LoRA adaptation raised F1 performance from 0.56 to above 0.93 on a corpus containing more than 72,000 annotated comments.arxiv.org보안
Preliminary FraudBench test finds banking agents vulnerable to adaptive fraud
An arXiv paper introduces FraudBench, a benchmark for testing whether tool-using banking agents can detect fraud that unfolds across conversations. In a preliminary single-trial evaluation, four agents scored 49% to 65% on attack security.arxiv.org혁신
Apple researchers report scaling law for training models with scarce data
A study of more than 2,000 language-model training runs says scarce target data can be repeated 15–20 times in mixtures, with the best rate varying by scale and compute.machinelearning.apple.com혁신
Apple Researchers Propose Lexical Substitutions to Improve Multilingual Model Training
Apple researchers describe LINK, a pretraining intervention that replaces selected English words with word-level translations from a target language. The paper reports improvements across eight languages and five model sizes, including up to a twofold speedup in reaching equivalent downstream performance.machinelearning.apple.com정책
Position paper calls for certification before AI agents make market decisions
A position paper reports tacit collusion by DeepSeek-R1 agents in a simulated Bertrand pricing market, even after human prompts against collusion. It argues that observed-behavior certification should precede deployment of reasoning agents in economic markets; the evidence and safeguards remain preliminary.arxiv.org
매주 한 번씩 유용한 브리핑을 듣습니다
피드에 살지 않고도 AI를 따라가세요.
이번 주의 검증된 AI 뉴스, 원본 데이터, 유용한 도구, 학습 추천, 그리고 새로운 AI 일자리를 받아보세요.
AI를 배우는 사람들에게 다가가세요
AI 전문가를 고용하거나 유용한 AI 제품을 출시하는 것? 배우고 행동하기 위해 이곳에 온 사람들 앞에 그것을 보여주세요.
AI 공고를 올리세요AI 도구를 제출하세요