뉴스로 돌아가기
혁신AI Understanding 브리핑

GLANCE는 한 번의 패스로 토큰 블록 초안을 작성하여 비전 언어 모델 디코딩 속도를 높입니다.

새로운 arXiv 논문에서는 비전 언어 모델의 융합 상태를 사용하여 전체 토큰 블록 초안을 한 번에 작성하는 추측 디코딩 방법인 GLANCE를 소개합니다. 저자는 테스트된 작업 부하에서 무손실 그리디 출력과 최대 2.93배 더 빠른 디코딩을 보고했습니다.

5 min readRead the primary source
Source-provided image accompanying GLANCE speeds up vision-language model decoding by drafting token blocks in one pass
기본 소스 문서녹음된 소스
출판사
arxiv.org
소스 링크
arxiv.orghttps://arxiv.org/abs/2609.00355
소스 유형
기본 문서 — 우리가 직접 읽는 공식 발표, 논문, 서류 또는 자사 페이지입니다.
맥락60초 안에 이해하세요

여기서 시작하세요

주요 용어

비전-언어 모델(VLM)
시각적 정보와 텍스트 정보를 공동으로 처리하는 다중 모드 모델입니다.
토큰
단어 조각이나 기호와 같은 언어 모델에 의해 처리되는 텍스트 덩어리입니다.
메모리(에이전트 메모리)
AI 에이전트는 연속성을 향상하기 위해 여러 단계 또는 세션에서 사용하는 저장된 컨텍스트입니다.
자신을 테스트해 보세요AI 모델 설명 퀴즈

무슨 일이 일어났나요?

Researchers introduced GLANCE, a one-pass block drafter for speculative decoding in vision-language models. The method reads the target model’s already-fused visual and language representation, generates a block of candidate tokens in one forward pass, and has the target model verify those candidates in a single pass. The authors report that audited prompts reproduced greedy decoding exactly and that GLANCE reached up to 2.93× the speed of autoregressive decoding under their stated test conditions.

A paper submitted to arXiv on Aug. 31 introduces GLANCE, which it describes as a one-pass block drafter for speculative decoding in vision-language models. Speculative decoding normally uses a smaller drafter to propose several tokens and then asks a larger target model to verify them. The paper argues that this design creates a problem for vision-language systems: a small autoregressive drafter may not be able to process the image at every step, even though visual information can make the next text tokens more predictable.

GLANCE changes the drafter’s role. According to the source, a block-diffusion head reads the target model’s already-fused vision-language state and fills an entire candidate block in one forward pass. A wide candidate tree is then checked in one target-model pass. This is the central technical claim: visual information is already present in the state supplied to the drafter, while the proposed tokens are generated as a block rather than sequentially. The paper calls the approach lossless because it targets an unmodified vision-language model and reports that every audited prompt reproduced greedy decoding exactly.

The authors report several comparisons under what they call one engine and one round budget. GLANCE decoded up to 2.93 times faster than autoregressive decoding, used one draft pass per round where a production EAGLE3-VL head used eight, and accepted blocks 2.7 times longer than an EAGLE-3 head trained on the same corpus. The paper also reports a relationship between accepted block length and the target model’s next- entropy. That relationship became steeper as tasks relied more heavily on grounding, across five tasks, and the authors say it transferred across targets and modalities. These are claims from the preprint, not independently established results.

소스 세부정보: arxiv.org ↗

왜 중요한가요?

The work addresses a practical bottleneck in multimodal AI: generating text one at a time can make image-grounded responses expensive and slow. If the reported results generalize, GLANCE could reduce inference latency and compute for tasks that copy or closely follow visual content, while preserving the target model’s output rather than trading accuracy for speed.

The immediate significance is computational. Vision-language models combine image processing with language generation, but the generation stage can still proceed by token. For responses that repeat text visible in an image or follow a strongly constrained visual prompt, that sequential process may perform many nearly predictable steps. A method that drafts a longer run in one pass could reduce latency and the amount of target-model computation required for each response.

The source’s distinction between grounded and free-running text is important. GLANCE is not presented as a universal replacement for autoregressive decoding. The authors say its advantage is strongest in a “verbatim-copy regime,” where visual grounding makes long sequences predictable, while free-running text still favors a chain. That boundary gives the result practical meaning: performance may depend less on whether a model is multimodal in general than on how much the particular task constrains the answer through visual evidence.

The claimed losslessness could also matter for deployment decisions. Speedups often require accepting a possible change in outputs, but this paper says the target model’s greedy result was reproduced on its audited prompts. If confirmed across broader tests, that would make the technique relevant to applications that need the original model’s decoding behavior while seeking lower response time or inference cost. The source does not establish those downstream savings, however, and it does not report production deployment, user-facing availability, or effects on quality beyond exact reproduction in the audited evaluation.

Interactive Mechanism

대화형 메커니즘: 실제로 작동하는 방식

이 개발의 이면에 있는 기본 기술을 대화식으로 살펴보세요.

Document Size:128K tokens
Needle Placement Depth (Location in document):50% into text
Attention Context Buffer Map:
Target Fact (50%)
Equivalent Pages~320Standard book pages
Retrieval Accuracy99.9%Needle recall score
RAM / KV Cache5.1 GBMemory overhead
Prompt CachingActive~80% discount on reuse
Core takeaway: Million-token context windows allow querying whole codebases or legal archives in one prompt. However, KV cache memory scales with context length, making prompt caching crucial for real-time production.
대화형 개념 확인+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

다음에 무엇을 볼 것인가

The paper is a newly submitted preprint, so its claims still need independent replication. The abstract does not specify the hardware, target models, datasets, prompt counts, absolute latency, or the five evaluated tasks. The reported maximum speedup may therefore depend heavily on workload and implementation choices; the authors themselves say free-running text still favors sequential drafting.

The main question is reproducibility. The source identifies the paper, authors, submission date, and code availability, but its abstract does not provide the implementation URL, hardware configuration, model names, dataset sizes, prompt counts, or detailed task definitions. Those details are needed to determine whether the 2.93× figure reflects a broad improvement or a best-case result on a particular workload. The wording “up to” signals that the maximum is not necessarily typical.

Comparisons also require careful interpretation. GLANCE is compared with autoregressive decoding and with EAGLE-based heads under a stated engine and round budget, but the abstract does not explain whether all systems used identical kernels, batching, memory settings, or verification policies. The claim that GLANCE accepted 2.7 times longer blocks than an EAGLE-3 head trained on the same corpus is informative, yet accepted length alone does not determine end-to-end latency. Verification cost, memory use, hardware utilization, and failure or rejection rates will all matter.

Further evaluation should test the method on more target models, modalities, languages, image types, and response styles, including cases where the answer is open-ended rather than copied from the visual input. The entropy relationship is potentially useful because it could help predict when block drafting will pay off, but the source gives no uncertainty interval or independent validation for that proposed law. Until those measurements are available, GLANCE is best understood as a promising research result from a preprint rather than a proven general-purpose acceleration standard.

관련 가이드 및 퀴즈

AI 모델 설명트랜스포머AI 트레이닝AI의 미래알고 있는 내용을 테스트해 보세요. 무료 AI 퀴즈를 시도해 보세요.용어집에서 AI 용어를 찾아보세요.AI 모델 출시 추적기를 따르세요.
이것이 유용하다고 생각하시나요?