Updated daily1866 verified stories
AI News. Without the noise.
Source-checked AI coverage of product launches, policy shifts, safety research, and industry moves, explained in plain English by a nonprofit education team.
Verified sourcing
Every story links to the strongest available evidence: original sources when available, otherwise clearly attributed reporting.
Plain English
What happened, why it matters, and what to watch — without the jargon.
No filler
When the signal is thin, we publish nothing rather than padding the feed.
More stories
9 storiesIndustry
Cursor Says Its Acquisition by SpaceX Has Officially Closed
Cursor published a short post saying SpaceX has completed its acquisition of the AI coding tool, finishing a process it says began in April with a model-training partnership with SpaceXAI. The post promises access to what it calls the world's largest GPU fleet, but discloses no terms, timelines, or product changes.
cursor.comInnovation
New Benchmark Finds AI Agents Wrongly Block Approved Work 28% of the Time
A preprint introduces SteerBench-Work, a 106-scenario test of the moment an AI agent decides to act or pause for review. Across 30 model conditions, the authors report that wrongly holding cleared work was roughly 28 times more common than wrongly allowing unsafe work.arxiv.orgInnovation
Paper Reports Frontier LLM Judges Flip Verdicts 25-71% Under Pushback
A new arXiv preprint stress-tests nine frontier models used as automated graders and reports that all of them change their verdicts under challenge — and that the changed verdicts usually move away from the correct answer, not toward it.arxiv.orgInnovation
Paper Argues Evolution Strategies Beat RL at Keeping LLM Answer Sets Diverse
A new arXiv preprint argues that post-training LLMs with evolution strategies — a population-based, gradient-free method that perturbs weights directly — beats reinforcement learning on pass@k and solution coverage. The abstract cites better math-benchmark results but names no models, benchmarks, or numbers.arxiv.orgSecurity
Paper Says Self-Improving AI Agents Can Turn One Unsafe Success Into a Reusable Skill
A new arXiv preprint benchmarks a specific agent failure mode: when a self-improving agent writes an unsafe procedure into memory, it can be retrieved and executed in later sessions. Every evolved configuration tested produced unsafe artifacts, and three malicious tasks more than doubled carryover attack success.arxiv.orgSecurity
SEAG Paper Proposes Aliasing Sensitive Entities Before RAG Queries Reach External LLMs
A preprint posted to arXiv describes a framework that swaps sensitive names in queries and retrieved documents for aliases before sending them to a third-party model. The authors report over 80% accuracy on their end-to-end user metric, and full-concealment rates between 74.91% and 77.83% across three small models.arxiv.orgInnovation
CABS+ Paper Reports Cheaper, Faster Model Merging Across 27 Datasets
A preprint posted to arXiv describes CABS+, a model-merging method that replaces grid search with a gradient-free coefficient search. The authors report double-digit performance gains over two baselines, under a quarter of one baseline's GPU memory, and roughly a 4x speedup over another.arxiv.orgInnovation
Paper Proposes Retrieved "Lessons" to Improve Spatial Reasoning in Frozen Vision-Language Models
An arXiv preprint describes Spatial Memory Agent, which stores verified experience as text lessons retrieved at inference time, claiming gains across five spatial benchmarks and four vision-language models without changing model weights. It is under review; its abstract names no benchmarks, base models, or margins.arxiv.orgInnovation
PROVE-RT Paper Reports 44.7% Success Generating Machine-Checked Real-Time Proofs
An arXiv preprint presents PROVE-RT, which uses retrieval and staged prompting to make large language models write PROSA/ROCQ proof scripts for real-time schedulability analysis. The authors report a 44.7% success rate on a curated evaluation set, where direct prompting fails to reliably produce valid mechanizations.arxiv.org
One useful briefing each week
Keep up with AI without living in the feed.
Get the week’s verified AI news, original data, useful tools, learning picks, and fresh AI jobs.
Reach people who are learning AI
Hiring an AI professional or launching a useful AI product? Put it in front of people who came here to learn and act.
Post an AI jobSubmit an AI tool