뉴스로 돌아가기
혁신AI Understanding 브리핑

TOPAS는 다중 에이전트 LLM 작업을 단축하기 위해 워크플로우 인식 스케줄링을 제안합니다.

새로운 arXiv 문서에서는 캐시된 LLM 접두어를 공동으로 관리하고 다단계 에이전트 워크플로에서 실행을 요청하는 스케줄러인 TOPAS를 소개합니다. 저자는 합성 워크플로와 두 가지 MetaGPT 소프트웨어 개발 워크로드에 대해 테스트한 기준보다 작업 완료 시간이 더 낮다고 보고합니다.

6 min readRead the primary source
Primary-source image accompanying TOPAS proposes workflow-aware scheduling to shorten multi-agent LLM jobs
기본 소스 문서녹음된 소스
출판사
arxiv.org
소스 링크
arxiv.orghttps://arxiv.org/abs/2608.25523
소스 유형
기본 문서 — 우리가 직접 읽는 공식 발표, 논문, 서류 또는 자사 페이지입니다.
맥락60초 안에 이해하세요

여기서 시작하세요

주요 용어

대형 언어 모델(LLM)
텍스트를 생성하고 분석하기 위해 대규모 텍스트 말뭉치를 학습한 언어 모델입니다.
메모리(에이전트 메모리)
AI 에이전트는 연속성을 향상하기 위해 여러 단계 또는 세션에서 사용하는 저장된 컨텍스트입니다.
시스템 프롬프트
모델의 동작, 정책 및 응답 스타일을 설정하는 우선순위가 높은 명령입니다.
자신을 테스트해 보세요AI 에이전트 퀴즈

무슨 일이 일어났나요?

Researchers propose TOPAS, a Task-Oriented Prefix-Aware Scheduler for serving multi-agent large language model workflows. It decides both which agent prefixes remain in a shared key-value cache and which requests should run, balancing cache reuse against progress through the overall workflow.

The paper focuses on a specific systems problem in multi-agent LLM serving: prefix caching. When an agent repeatedly uses a long or other shared prefix, retaining its key-value cache can accelerate a later model call. But the same cache occupies GPU memory that could otherwise be used to batch concurrent requests. The authors describe this as a tradeoff between immediate prefix locality and progress through a multi-stage workflow. In that framing, cache contents are part of the scheduling state rather than a separate implementation detail.

TOPAS addresses that tradeoff by making two decisions together. It chooses which agent prefixes to retain under a shared key-value-cache budget and which pending requests to schedule for execution. Its scoring method evaluates possible post-decision states by balancing the expected reduction in each task’s longest remaining service path against the near-term benefit of reusing downstream prefixes. The calculation also accounts for prefix movement and preemption costs. The scheduler includes a task-level aging mechanism intended to prevent some tasks from waiting indefinitely. These components connect memory residency, request order, and workflow progress in one policy.

The authors implemented TOPAS within the SGLang framework and evaluated it on three synthetic directed acyclic graph workloads and two MetaGPT software-development workflows. The paper reports that, compared with the best-performing baseline for each workload and metric, TOPAS reduced mean and 99th-percentile job-completion time by as much as 39.8% and 49.4% on the synthetic workloads. On the MetaGPT-SOP workload, it reports a 9.8% reduction in mean job-completion time. On MetaGPT-TL, it reports reductions of 22.0% in mean completion time and 26.6% at the 99th percentile. These figures are claims made by the paper, not independently verified results in the supplied source. The evaluation therefore describes the reported behavior of the implementation across the listed test workloads.

The proposal is notable for combining cache selection with execution selection at the workflow level. That combination is the central systems development described in the paper, while the reported percentages provide the evidence offered for its effect on completion time.

소스 세부정보: arxiv.org ↗

왜 중요한가요?

Multi-agent systems often issue dependent model calls, so a scheduling decision that helps one request can delay later stages. The paper reports substantial reductions in mean and tail job-completion time on its evaluated workloads, suggesting that serving infrastructure—not only model quality—can materially affect the performance and cost of agent systems.

The practical importance lies in the structure of agent workflows. A multi-agent task is often not one model request but a chain or graph of calls, with later calls waiting for earlier outputs. A scheduler that optimizes only the next request or only cache reuse can improve a local metric while delaying a critical downstream path. TOPAS instead treats the workflow’s remaining path as part of the scheduling decision, which is a direct attempt to align serving behavior with the completion of the user’s overall task. This makes the unit of optimization the workflow job rather than an isolated model call.

The reported reductions are potentially meaningful because job-completion time affects how quickly an agent can respond and how many concurrent workflows a serving system can support. Improvements in cache use and scheduling could also reduce pressure to add hardware for a given workload, although the source does not measure infrastructure savings, energy use, operating cost, response quality, or user satisfaction. The paper therefore provides evidence about scheduling latency in its test setup, not a complete case for deployment economics. Those unmeasured dimensions remain separate from the latency results.

The work also highlights a limitation of evaluating agent performance solely at the model-call level. For multi-stage systems, a model may produce the same outputs while users experience different delays because requests compete for memory and execution slots. By making prefix state and workflow dependencies explicit, the paper offers a systems perspective that may be useful to engineers designing agent-serving infrastructure. Still, the source does not establish that TOPAS improves reliability, accuracy, fairness among users, or safety; its reported target is job-completion time. The distinction matters because faster completion does not by itself demonstrate broader system quality.

More generally, the paper connects infrastructure behavior to the user-visible timing of compound AI tasks. That connection helps explain why workflow-aware scheduling can matter even when the underlying model and generated outputs remain unchanged, while keeping the source’s evidence limited to its reported evaluation.

Interactive Mechanism

대화형 메커니즘: 실제로 작동하는 방식

이 개발의 이면에 있는 기본 기술을 대화식으로 살펴보세요.

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
대화형 개념 확인+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

다음에 무엇을 볼 것인가

The results are from a single arXiv preprint and a limited evaluation: three synthetic directed acyclic graphs and two MetaGPT workflows. Important open questions include performance on production traffic, different model-serving stacks, larger or more diverse workflows, and the effect of implementation overhead and cache policies outside the tested settings.

The first question is whether the reported gains hold beyond the paper’s five workload families. Three workloads are synthetic, while the remaining two are identified as MetaGPT software-development workflows. The source does not provide the workload sizes, traffic distributions, model configurations, hardware details, baseline names, or variance across repeated runs in the supplied text. Those details are important for judging how broadly the percentages apply. Without them, the numerical comparisons are difficult to transfer directly to another serving environment.

Evaluation on real production traces would be especially informative. Production systems may have irregular arrival patterns, cancellations, heterogeneous prompt lengths, multiple models, changing cache residency, and service-level objectives that are not represented by a fixed workflow graph. It would also be useful to know how TOPAS behaves when prefix movement and preemption are expensive, when the cache budget is very small, or when workflows contain branches that do not execute predictably. Such tests would probe the same tradeoffs under conditions more varied than those described in the supplied evaluation.

The paper is an arXiv submission dated Aug. 26, 2026, and the source identifies it as an eight-page paper. The supplied record does not establish peer review, independent replication, public code availability, or deployment by a serving provider. Follow-up work should test those issues, compare TOPAS with additional schedulers, report resource and quality tradeoffs, and examine whether the aging mechanism introduces different delays across tasks. Until then, the strongest supported conclusion is that the authors report a promising scheduling method in a bounded experimental evaluation. The scope of that conclusion should remain tied to the evidence and workload descriptions available here.

The most useful next evidence would therefore combine broader workloads with transparent experimental details and measurements beyond completion time. Those additions would clarify both reproducibility and the practical limits of the reported scheduling approach without changing what the current paper claims.

관련 가이드 및 퀴즈

AI 에이전트AI 모델 설명트랜스포머알고 있는 내용을 테스트해 보세요. 무료 AI 퀴즈를 시도해 보세요.용어집에서 AI 용어를 찾아보세요.AI 모델 출시 추적기를 따르세요.
이것이 유용하다고 생각하시나요?