返回新聞
創新AI Understanding 簡報

研究人員提出 Transformer 前饋層的令牌自適應活化混合

修訂後的 arXiv 預印本提出了混合激活,這是一種前饋層設計,允許語言模型在每個標記的激活函數中進行選擇。作者報告了從 0.12B 到 2B 參數的模型中終端訓練損失較低,但在來源記錄中沒有提供確切的結果。

5 min readRead the primary source
Source-page capture accompanying Researchers propose token-adaptive activation mixing for Transformer feedforward layers
主要來源文件來源記錄
出版商
arxiv.org
來源連結
arxiv.orghttps://arxiv.org/abs/2605.26647
來源類型
主要文件-我們直接閱讀的官方公告、文件、文件或第一方頁面。
背景60 秒內了解這一點

從這裡開始

關鍵術語

變壓器
一種神經架構,利用注意力並行地對序列之間的關係進行建模。
代幣
由語言模型處理的文字區塊,例如單字或符號。
記憶體(代理記憶體)
AI 代理程式跨步驟或會話使用儲存的上下文來提高連續性。
測試一下自己AI 模型解釋測驗

發生了什麼事

A paper revised on arXiv on August 28 proposes Mixture of Activations, or MoA, for -based language models. The design uses lightweight, input-dependent gates to mix several activation functions for individual tokens while sharing the same linear projections. The authors also introduce learnable activations that combine functions without -dependent gates.

The authoritative source is an arXiv record for “More Expressive Feedforward Layers: Part I. -Adaptive Mixing of Activations.” It says the paper was first submitted on May 26, 2026, and revised on August 28, placing the substantive update inside the current news window. The authors are Mingze Wang, Jinbo Wang, Yikuan Xia, Kai Shen, and Shu Zhong. The source identifies the work as a machine-learning and artificial-intelligence paper, with language models as its direct subject.

The paper starts from a limitation the authors identify in common feedforward network, or FFN, layers. Earlier designs used functions such as ReLU and GELU, while later gated designs include SwiGLU, but the source says most FFNs still apply one fixed nonlinear transformation to every . The proposed Mixture of Activations instead uses a dictionary of activation functions and lightweight gates whose values depend on the input. The same linear projections are shared, while the mixture of nonlinear functions can vary from token to token.

The source also describes learnable activations, abbreviated LA, as an input-independent counterpart. LA forms linear combinations of activation functions in both ReLU-type and SwiGLU-type FFNs. The authors state that their finite-width theory establishes strict expressive separations: LA strictly contains fixed-activation FFNs, while MoA strictly contains LA because it adds input-dependent nonlinear hybridization. These are theoretical expressivity claims about the functions the layers can represent, not evidence that a deployed language model is generally more capable or reliable.

For empirical testing, the authors say they conducted extensive pretraining experiments on dense and mixture-of-experts language models ranging from 0.12 billion to 2 billion parameters. They report testing different budgets, optimizers, and learning-rate schedules. According to the abstract, MoA consistently achieved lower terminal loss and showed more favorable scaling behavior than well-tuned baselines, with minimal parameter and computational overhead. The source record does not provide the numerical loss differences, the names of the datasets or baselines, hardware details, or the precise meaning of “minimal” overhead.

來源詳情: arxiv.org ↗

為什麼這很重要

Feedforward layers contain a large share of the parameters and nonlinear behavior in many language models. If -adaptive activation mixing improves training without materially increasing compute or parameters, it could offer a relatively small architectural change with implications for model scaling. The evidence remains limited to the authors’ reported pretraining experiments and does not establish better downstream performance or production efficiency.

The proposal targets a consequential part of current language-model architecture. The source says FFN layers account for a large fraction of model parameters and nonlinear expressivity. That makes them an important place to seek improvements: a change to the nonlinear transformation could affect how efficiently a model uses its existing width and projections, rather than requiring an entirely different model family.

MoA’s design is potentially practical because it retains shared linear projections and adds lightweight input-dependent gating. In principle, this could allow different tokens to receive different nonlinear treatment while preserving much of the surrounding FFN structure. However, the source does not establish that existing models can be upgraded without retraining, nor does it report implementation details sufficient to determine whether the added gates improve real-world throughput, memory use, or energy consumption.

The reported lower terminal loss is relevant because training loss is a central measure of how well a model fits its pretraining data. The claimed scaling behavior could also matter if the method continues to improve as model size or training compute increases. But terminal loss is not the same as usefulness to people. The source gives no results for factuality, reasoning, coding, multilingual performance, safety, calibration, robustness, or downstream task accuracy, so the practical significance of the reported training gains remains uncertain.

The evidence is also preliminary. It comes from a revised preprint, and the source record presents the authors’ theoretical and empirical conclusions rather than independent validation. The tested range ends at 2 billion parameters, which is materially smaller than many widely deployed language models. The record does not state whether the experiments used multiple random seeds, how the baselines were tuned, whether the method changes inference latency, or whether its gains persist outside the reported pretraining settings.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
互動式概念檢查+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

接下來看什麼

The important next evidence is the size and consistency of the reported gains, especially beyond the tested 0.12B-to-2B range. Readers should look for exact compute, memory, latency, and parameter overheads; results on downstream tasks; independent replication; and evidence that the method remains useful across architectures, datasets, optimizers, and longer training runs.

The next useful check is quantitative detail. The full paper or later work should show exact loss curves, confidence or run-to-run variation, parameter counts, operation counts, memory requirements, and wall-clock training costs for MoA, LA, and fixed-activation baselines. Because the abstract says the comparisons were against well-tuned baselines, the tuning procedure and compute budget will be important for judging whether the advantage comes from the architecture rather than unequal optimization.

Scale is another unresolved issue. The source reports models from 0.12B to 2B parameters, but it does not say whether the same pattern holds in substantially larger dense or mixture-of-experts systems. Future experiments should test larger models, longer training runs, additional budgets, and different data mixtures. They should also clarify whether the claimed favorable scaling behavior is measured by loss at a fixed compute budget, by loss at a fixed parameter count, or by another comparison.

Operational costs deserve close attention. MoA introduces a -dependent gate and a dictionary of activation functions, even if the added parameter and compute burden is described as minimal. Independent measurements should determine whether this produces meaningful changes in accelerator utilization, memory traffic, batching, inference latency, training stability, or energy use. A small theoretical overhead can have a larger systems cost if it disrupts efficient hardware execution.

Finally, readers should look for independent replication and broader evaluation. Results on downstream language, coding, reasoning, and multilingual tasks would show whether lower pretraining loss transfers to capabilities that users notice. Evaluations of reliability and safety would test whether -adaptive nonlinear behavior introduces new failure patterns. Until those results are available, the paper is best understood as a promising architectural research claim, not a demonstrated production improvement or an immediately available product.

相關指引和測驗

人工智慧模型解釋變形金剛人工智慧培訓測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語關注 AI 模型發布追蹤器
覺得有用嗎?