GraphEcho: การประเมินตัวแทนกราฟ LLM
GraphEcho tests whether LLM agents mistake repeated encounters for additional corroboration. GraphEcho is a benchmark designed to evaluate large language model (LLM) graph agents. It tests whether these agents can distinguish between repeated encounters and additional corroboration. เกณฑ์มาตรฐานแตกต่างกันไปตามจำนวนเส้นทาง...