Updated daily2000 verified stories
AI News. Without the noise.
Source-checked AI coverage of product launches, policy shifts, safety research, and industry moves, explained in plain English by a nonprofit education team.
Verified sourcing
Every story links to the strongest available evidence: original sources when available, otherwise clearly attributed reporting.
Plain English
What happened, why it matters, and what to watch — without the jargon.
No filler
When the signal is thin, we publish nothing rather than padding the feed.
More stories
9 storiesInnovation
Study maps 46 language models built for Portuguese
A systematic mapping study catalogs 46 Portuguese language models and compares architectures, training resources, licensing, code, data, and weights. The authors say the field is growing but difficult to assess because information is spread across papers, technical reports, repositories, and project documentation.arxiv.orgInnovation
Interpretability study links Qwen3-4B’s overconfident answers to a certainty-biased mechanism
An arXiv study reports that Qwen3-4B favors certainty over uncertainty in controlled reasoning tasks and identifies model features that may drive the imbalance. The authors say targeted interventions reduced overconfident errors, but the abstract does not disclose effect sizes or testing details.arxiv.orgInnovation
The Deontic Gap: Large Language Models and the Modal Language of Obligation
Modal auxiliaries such as must, should, and have to mark necessity and obligation within the contexts of speaker authority and interpersonal stance.arxiv.orgSecurity
Paper proposes obscuring refusal signals to resist abliteration attacks
An arXiv preprint introduces a weight-editing method intended to make safety refusals harder to extract and remove. The paper reports stronger post-abliteration refusal scores on two open models, with different tradeoffs in general-purpose performance.arxiv.orgInnovation
Paper reports a temporal method for detecting hallucinations at the token level
An arXiv preprint describes a hallucination detector that combines text statistics, entailment signals and language-model surprisal across sequences instead of judging tokens independently. Its BiGRU model reached an AUC of 0.840 on RAGTruth, according to the paper.arxiv.orgInnovation
Entity tracking emerges at 410 million parameters, exceeds humans across naturalistic narratives
Understanding language requires tracking entities across discourse - i.e., knowing where things are and how they change, even when not explicitly stated.arxiv.orgInnovation
Paper reports compiler-guided search improves Lean theorem-proving efficiency
An arXiv preprint proposes an adaptive proof-search method for context-dependent Lean 4 projects. Its authors report a 12.8-percentage-point average pass-rate improvement within a pass@32 budget while using 21.9% fewer LLM calls than pass@k baselines.arxiv.orgInnovation
Benchmark Separates Geometry Solving From Diagram Construction
A new open-source benchmark tests whether foundation models can construct faithful geometry diagrams, not merely solve the underlying problems. Its authors report that evaluated models achieved an average compile success rate of 36.14%, exposing a gap between mathematical answers and usable visual constructions.arxiv.orgInnovation
Alignment Is All You Need: Instruction-Free Training for General Audio-Language Models
A new approach to training multimodal large language models (MLLMs) eliminates the need for extensive task-specific supervision.arxiv.org
One useful briefing each week
Keep up with AI without living in the feed.
Get the week’s verified AI news, original data, useful tools, learning picks, and fresh AI jobs.
Reach people who are learning AI
Hiring an AI professional or launching a useful AI product? Put it in front of people who came here to learn and act.
Post an AI jobSubmit an AI tool