Gated-memory routing aims to reduce the cost of multi-agent LLM systems
A newly submitted paper proposes retaining only useful reasoning steps when multiple language-model agents collaborate. Its abstract reports higher average benchmark accuracy and a 31.9% reduction in HumanEval inference cost compared with the strongest baseline.