SSKG method makes LLM student simulations track mastery more faithfully, paper reports
An arXiv preprint reports that a stochastic knowledge-graph method made three LLMs produce more distinguishable simulated students across 379 SAT Algebra items.
Updated daily401 verified stories
Material AI papers, studies, datasets, evaluations, and scientific advances with direct links to the underlying research.
Every story links to the strongest available evidence: original sources when available, otherwise clearly attributed reporting.
What happened, why it matters, and what to watch — no jargon tax.
When the signal is thin, we publish nothing rather than padding the feed.
A growing stream of verified perspectives for people who need to understand AI without chasing hype.
An arXiv preprint reports that a stochastic knowledge-graph method made three LLMs produce more distinguishable simulated students across 379 SAT Algebra items.
A new preprint reports that optimized lightweight hybrid models classified oral-cancer images with 83.2% average sensitivity and 86.0% average specificity on a held-out test set, suggesting a possible edge-AI approach for primary-care triage.
A new arXiv preprint describes HIRA, a retrieval-based AI classification system that routes uncertain documents to a local language model or human reviewer. The authors report higher Macro-F1 scores while reducing the number of documents requiring human correction and the number of language-model calls.
A new arXiv preprint proposes training GUI agents with trajectory-level feedback, rewarding concise successful executions and distinguishing failures by how they diverge from successful behavior.
The University of Chicago will make its social science Core classes technology-free beginning this fall, according to The College Fix, citing a divisional memo obtained by The Chicago Maroon.
A new preprint introduces a benchmark for measuring whether large language models can deliberately control their internal activation patterns. The authors report that most tested models could alter the direction and magnitude of residual-stream activations through natural-language instructions, sometimes evading…
An arXiv preprint introduces CompPO, a reinforcement-learning method that uses a language model’s own attention patterns to assign training credit across tokens. The authors report higher held-out accuracy and greater stability than tuned GRPO in experiments on Qwen3-4B and Llama-3.1-8B-Instruct.
A new benchmark built from live scientific-agent requests finds that no tested model consistently met the study’s acceptance threshold. The paper reports stronger communication than scientific accuracy, with overclaiming the most common failure tag.
A new arXiv preprint presents ATHENA, a multi-agent system that reuses architecture knowledge across hospitals when searching for Transformer designs for electronic-health-record prediction. The authors report that it matched or outperformed four neural-architecture-search baselines in 9 of 12 hospital-task…
The University of Chicago Law School will bar laptops, tablets and phones from required first-year classes and add live oral discussions to second-year research papers as part of an AI-focused education strategy, the Hyde Park Herald reports.
Tech Xplore reports that a benchmark of 22 language models found trade-offs between factual accuracy and social representation when models recommend experts. Retrieval improved factuality, while prompting improved representation, but no tested intervention improved both.
AFP reports that Ukraine will give Britain access to the Avengers dataset, containing millions of battlefield images used to train autonomous Ukrainian drones. The data is expected to be tested at a UK defense site to help protect bases and other sensitive infrastructure.
One useful briefing each week
Get the week’s verified AI news, original data, useful tools, learning picks, and fresh AI jobs.
Hiring an AI professional or launching a useful AI product? Put it in front of people who came here to learn and act.
Post an AI job Submit an AI tool