검증된 출처
모든 기사는 가장 강력한 증거와 연결됩니다: 원가 자료가 있을 경우 그렇고, 그 외에는 명확히 출처가 명확히 기재된 보도입니다.
평범한 영어
무슨 일이 일어났는지, 왜 중요한지, 무엇을 보아야 하는지 등을 전문 용어 없이 설명합니다.
채우기 없는 이야기
신호가 얇을 때는 피드를 채우기보다는 아무것도 발행하지 않습니다.
더 많은 이야기
9 이야기들혁신
새로운 프리프린트 보고서는 다중 단계 도구 호출에서 반복 언어 모델의 이점을 얻었습니다
arXiv 연구는 세 가지 도구 호출 벤치마크에서 반복 및 기존 언어 모델을 평가하며, 여러 의존 API 호출과 잠재적으로 더 효율적인 적응형 계산 접근법이 필요한 워크플로우에서 더 강력한 결과를 보고했습니다.arxiv.org혁신
대립적 검토 테스트는 에이전트 코드 검토를 위한 구조화된 불일치
새로운 arXiv 논문은 리뷰어가 에이전트의 코드를 평가하고 비평가가 편집 전에 검토하는 감사를 하는 3인 에이전트 코드 검토 프로토콜을 제안합니다. 저자들은 테스트된 베이스라인 대비 벤치마크 결과가 개선되었다고 보고하며, 허위 합의를 실패 모드로 식별했습니다.arxiv.org혁신
ArXiv study reports large Roman Urdu hate-speech gains from LoRA adaptation
An arXiv preprint compares zero-shot and parameter-efficient fine-tuning for hate-speech detection in Roman Urdu. The authors report that LoRA adaptation raised F1 performance from 0.56 to above 0.93 on a corpus containing more than 72,000 annotated comments.arxiv.org보안
Preliminary FraudBench test finds banking agents vulnerable to adaptive fraud
An arXiv paper introduces FraudBench, a benchmark for testing whether tool-using banking agents can detect fraud that unfolds across conversations. In a preliminary single-trial evaluation, four agents scored 49% to 65% on attack security.arxiv.org혁신
Apple researchers report scaling law for training models with scarce data
A study of more than 2,000 language-model training runs says scarce target data can be repeated 15–20 times in mixtures, with the best rate varying by scale and compute.machinelearning.apple.com혁신
Apple Researchers Propose Lexical Substitutions to Improve Multilingual Model Training
Apple researchers describe LINK, a pretraining intervention that replaces selected English words with word-level translations from a target language. The paper reports improvements across eight languages and five model sizes, including up to a twofold speedup in reaching equivalent downstream performance.machinelearning.apple.com정책
Position paper calls for certification before AI agents make market decisions
A position paper reports tacit collusion by DeepSeek-R1 agents in a simulated Bertrand pricing market, even after human prompts against collusion. It argues that observed-behavior certification should precede deployment of reasoning agents in economic markets; the evidence and safeguards remain preliminary.arxiv.org혁신
Systematic review maps the growing use of large language models in mental health
A systematic review surveys how large language models are being studied for mental-health analysis, risk assessment, therapy support and multimodal monitoring, while stressing unresolved ethical and regulatory challenges.arxiv.org정책
Model Cards Alone May Not Govern Open-Weight Foundation Models, Position Paper Argues
An ICML 2026 position paper analyzing 500 Hugging Face model cards argues that open-weight foundation models need coordinated model cards, acceptable-use policies, and licenses to address safety and governance gaps.arxiv.org
매주 한 번씩 유용한 브리핑을 듣습니다
피드에 살지 않고도 AI를 따라가세요.
이번 주의 검증된 AI 뉴스, 원본 데이터, 유용한 도구, 학습 추천, 그리고 새로운 AI 일자리를 받아보세요.
AI를 배우는 사람들에게 다가가세요
AI 전문가를 고용하거나 유용한 AI 제품을 출시하는 것? 배우고 행동하기 위해 이곳에 온 사람들 앞에 그것을 보여주세요.
AI 공고를 올리세요AI 도구를 제출하세요