검증된 출처
모든 기사는 가장 강력한 증거와 연결됩니다: 원가 자료가 있을 경우 그렇고, 그 외에는 명확히 출처가 명확히 기재된 보도입니다.
평범한 영어
무슨 일이 일어났는지, 왜 중요한지, 무엇을 보아야 하는지 등을 전문 용어 없이 설명합니다.
채우기 없는 이야기
신호가 얇을 때는 피드를 채우기보다는 아무것도 발행하지 않습니다.
더 많은 이야기
9 이야기들산업
Cursor Says Its Acquisition by SpaceX Has Officially Closed
Cursor published a short post saying SpaceX has completed its acquisition of the AI coding tool, finishing a process it says began in April with a model-training partnership with SpaceXAI. The post promises access to what it calls the world's largest GPU fleet, but discloses no terms, timelines, or product changes.
cursor.com혁신
New Benchmark Finds AI Agents Wrongly Block Approved Work 28% of the Time
A preprint introduces SteerBench-Work, a 106-scenario test of the moment an AI agent decides to act or pause for review. Across 30 model conditions, the authors report that wrongly holding cleared work was roughly 28 times more common than wrongly allowing unsafe work.arxiv.org혁신
Paper Reports Frontier LLM Judges Flip Verdicts 25-71% Under Pushback
A new arXiv preprint stress-tests nine frontier models used as automated graders and reports that all of them change their verdicts under challenge — and that the changed verdicts usually move away from the correct answer, not toward it.arxiv.org혁신
Paper Argues Evolution Strategies Beat RL at Keeping LLM Answer Sets Diverse
A new arXiv preprint argues that post-training LLMs with evolution strategies — a population-based, gradient-free method that perturbs weights directly — beats reinforcement learning on pass@k and solution coverage. The abstract cites better math-benchmark results but names no models, benchmarks, or numbers.arxiv.org보안
Paper Says Self-Improving AI Agents Can Turn One Unsafe Success Into a Reusable Skill
A new arXiv preprint benchmarks a specific agent failure mode: when a self-improving agent writes an unsafe procedure into memory, it can be retrieved and executed in later sessions. Every evolved configuration tested produced unsafe artifacts, and three malicious tasks more than doubled carryover attack success.arxiv.org보안
SEAG Paper Proposes Aliasing Sensitive Entities Before RAG Queries Reach External LLMs
A preprint posted to arXiv describes a framework that swaps sensitive names in queries and retrieved documents for aliases before sending them to a third-party model. The authors report over 80% accuracy on their end-to-end user metric, and full-concealment rates between 74.91% and 77.83% across three small models.arxiv.org혁신
CABS+ Paper Reports Cheaper, Faster Model Merging Across 27 Datasets
A preprint posted to arXiv describes CABS+, a model-merging method that replaces grid search with a gradient-free coefficient search. The authors report double-digit performance gains over two baselines, under a quarter of one baseline's GPU memory, and roughly a 4x speedup over another.arxiv.org혁신
논문에서는 고정된 시각 언어 모델의 공간 추론을 개선하기 위해 검색된 "수업"을 제안합니다.
arXiv 사전 인쇄에서는 검증된 경험을 추론 시 검색된 텍스트 강의로 저장하고 모델 가중치를 변경하지 않고 5개의 공간 벤치마크와 4개의 비전 언어 모델에서 이득을 주장하는 공간 메모리 에이전트에 대해 설명합니다. 검토 중입니다. 추상적인 이름에는 벤치마크, 기본 모델 또는 마진이 없습니다.arxiv.org혁신
PROVE-RT Paper Reports 44.7% Success Generating Machine-Checked Real-Time Proofs
An arXiv preprint presents PROVE-RT, which uses retrieval and staged prompting to make large language models write PROSA/ROCQ proof scripts for real-time schedulability analysis. The authors report a 44.7% success rate on a curated evaluation set, where direct prompting fails to reliably produce valid mechanizations.arxiv.org
매주 한 번씩 유용한 브리핑을 듣습니다
피드에 살지 않고도 AI를 따라가세요.
이번 주의 검증된 AI 뉴스, 원본 데이터, 유용한 도구, 학습 추천, 그리고 새로운 AI 일자리를 받아보세요.
AI를 배우는 사람들에게 다가가세요
AI 전문가를 고용하거나 유용한 AI 제품을 출시하는 것? 배우고 행동하기 위해 이곳에 온 사람들 앞에 그것을 보여주세요.
AI 공고를 올리세요AI 도구를 제출하세요