返回新聞
創新AI Understanding 簡報

SHIFT-LLM 報告了一種在修剪 LLM 層後恢復準確性的免訓練方法

新的預印本描述了 SHIFT-LLM,這是一種後剪枝校正方法,它使用輕量級線性適配器來近似從大型語言模型中刪除的計算。作者報告稱,在七個零樣本基準測試中,Llama-3.1-8B-Instruct 的準確率提高了高達 15.7 點。

5 min readRead the primary source
Primary-source image accompanying SHIFT-LLM reports a training-free way to recover accuracy after pruning LLM layers
主要來源文件來源記錄
出版商
arxiv.org
來源連結
arxiv.orghttps://arxiv.org/abs/2608.25068
來源類型
主要文件-我們直接閱讀的官方公告、文件、文件或第一方頁面。
背景60 秒內了解這一點

從這裡開始

關鍵術語

大語言模型(LLM)
在海量文本語料庫上訓練來產生和分析文本的語言模型。
修剪
刪除不太重要的模型權重或神經元以減少尺寸和計算量。
記憶體(代理記憶體)
AI 代理程式跨步驟或會話使用儲存的上下文來提高連續性。
測試一下自己AI 模型解釋測驗

發生了什麼事

A research team introduced SHIFT-LLM, a training-free correction framework for large language models whose Transformer blocks have been removed to reduce inference costs. The method inserts a Linear Residual Adapter at each site and calibrates it with closed-form least-squares regression using a small held-out dataset. The paper reports evaluations across five model families, six layer-selection criteria and seven zero-shot benchmarks, with accuracy recovery in most tested configurations.

The paper describes depth as a way to reduce the inference cost of large language models by removing entire Transformer blocks. According to the authors, this creates a distribution shift: the hidden states produced after a removed block no longer match the distributions expected by the downstream layers. The resulting mismatch can cause substantial accuracy loss even though the remaining model still contains most of its original structure. SHIFT-LLM is presented as a correction framework for this specific problem, making the interaction between layer removal and downstream model behavior the central subject of the work.

The proposed component is a Linear Residual Adapter, or LRA, placed at each location where a block has been pruned. The adapter preserves the identity path of the original residual block and adds a lightweight affine residual correction. The authors say this correction is fitted through closed-form least-squares regression on a small held-out calibration set. It does not require gradient computation. In the paper’s description, the LRA is intended to approximate the residual update that would have been produced by the removed block, while avoiding the attention and feed-forward computations associated with that block.

The source reports that the method was tested on five model families, six criteria for selecting layers to remove and seven zero-shot benchmarks. It says SHIFT-LLM consistently recovered accuracy lost through depth across most configurations, with a maximum improvement of 15.7 percentage points on Llama-3.1-8B-Instruct. The source does not provide the benchmark names, baseline scores, pruning ratios, hardware measurements or per-model results in the supplied text. It also does not say that the method has been independently reproduced or integrated into a commercial product.

來源詳情: arxiv.org ↗

為什麼這很重要

Removing entire layers can lower the cost of running an LLM but can also disrupt the internal representations expected by later layers. SHIFT-LLM is designed to address that tradeoff with lightweight corrections rather than restoring the removed attention and feed-forward computations. If the reported results generalize, the approach could make compressed language models more practical where inference cost, memory or latency are constrained.

The practical issue addressed by the paper is a familiar compression tradeoff: a smaller or shallower model may be cheaper to run, but its accuracy can fall because removing layers changes the internal signals passed through the network. The paper’s contribution is to target that failure mode directly rather than treating the pruned model as an ordinary smaller model. This matters because inference cost is influenced not only by parameter count but also by the amount of attention and feed-forward computation performed for each input.

SHIFT-LLM’s reported calibration procedure is potentially useful because it does not depend on gradient-based retraining. The authors say only a few hundred calibration samples are needed and that the adapters can be fitted with closed-form regression. That could reduce the data, memory and engineering burden associated with recovering quality after . The proposed low-rank factorization and exact merging across consecutive pruned layers are also presented as ways to add further compression. These are claims about the method’s design and reported experiments, not evidence that every deployment will achieve the same savings.

The work also highlights why model compression cannot be evaluated solely by counting removed parameters or layers. A pruned model may be structurally smaller while still suffering from internal distribution mismatches that affect output quality. If the reported recovery is robust, lightweight correction layers could give developers another option between running a full model and accepting the accuracy loss of aggressive . The public significance remains bounded, however: the supplied source does not quantify energy use, serving cost, end-to-end latency or performance in high-stakes applications.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Document Size:128K tokens
Needle Placement Depth (Location in document):50% into text
Attention Context Buffer Map:
Target Fact (50%)
Equivalent Pages~320Standard book pages
Retrieval Accuracy99.9%Needle recall score
RAM / KV Cache5.1 GBMemory overhead
Prompt CachingActive~80% discount on reuse
Core takeaway: Million-token context windows allow querying whole codebases or legal archives in one prompt. However, KV cache memory scales with context length, making prompt caching crucial for real-time production.
互動式概念檢查+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

接下來看什麼

The main open questions are how the method performs outside the reported benchmarks, how much calibration data and added computation are required in practice, and whether its gains persist across tasks, languages and deployment environments. The source is an arXiv preprint and does not establish independent replication, production availability, detailed latency results or the full experimental conditions behind the reported maximum gain.

The first issue to watch is reproducibility. The source identifies the paper as version one of an arXiv submission and summarizes its results, but the supplied text does not include the underlying tables, ablation studies or implementation details. Independent tests would need to examine whether the reported gains hold under different percentages, different calibration-set sizes and different choices of which layers to remove. They would also need to compare SHIFT-LLM against retraining, distillation, ordinary fine-tuning and other compression methods under consistent compute budgets.

Deployment measurements will be important. The paper says LRAs avoid the expensive attention and feed-forward computations of removed blocks and support low-rank factorization and merging, but the source does not report actual latency, memory use, throughput or hardware results. Added adapters may have different effects depending on the inference framework and whether consecutive operations can be fused. A practical evaluation should therefore measure total serving cost, not just accuracy recovery or the number of parameters removed.

The scope of generalization is another unknown. The reported evaluation covers five model families and seven zero-shot benchmarks, but the supplied abstract does not identify the tasks, languages, model sizes beyond the named Llama-3.1-8B-Instruct result, or the behavior of the method on instruction following, long-context use, coding or safety evaluations. The source also does not establish whether calibration data must resemble deployment inputs. Future papers or released code could clarify these limitations and show whether the approach is broadly useful or mainly effective for particular architectures and patterns.

相關指引和測驗

人工智慧模型解釋變形金剛人工智慧培訓AI 的未來測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語關注 AI 模型發布追蹤器
覺得有用嗎?