Paper reports edge LLM agents can cut thinking compute by 43%-65% with calibrated deferral
A revised arXiv paper proposes TSDS, a framework that stops local reasoning when an action stabilizes and sends uncertain actions to a cloud model. The authors report 43%-65% lower per-episode thinking compute on three of four tested tasks while maintaining stated guarantees on expected reward and cloud-call rate.