返回新闻
创新AI Understanding 简报

研究人员提出 Transformer 前馈层的令牌自适应激活混合

修订后的 arXiv 预印本提出了混合激活,这是一种前馈层设计,允许语言模型在每个标记的激活函数中进行选择。作者报告了从 0.12B 到 2B 参数的模型中终端训练损失较低,但在源记录中没有提供确切的结果。

5 min readRead the primary source
Source-page capture accompanying Researchers propose token-adaptive activation mixing for Transformer feedforward layers
主要来源文件来源记录
出版商
arxiv.org
来源链接
arxiv.orghttps://arxiv.org/abs/2605.26647
来源类型
主要文件——我们直接阅读的官方公告、文件、文件或第一方页面。
背景60 秒内了解这一点

从这里开始

关键术语

变压器
一种神经架构,利用注意力并行地对序列之间的关系进行建模。
代币
由语言模型处理的文本块,例如单词或符号。
内存(代理内存)
AI 代理跨步骤或会话使用存储的上下文来提高连续性。
测试一下自己AI 模型解释测验

发生了什么

A paper revised on arXiv on August 28 proposes Mixture of Activations, or MoA, for -based language models. The design uses lightweight, input-dependent gates to mix several activation functions for individual tokens while sharing the same linear projections. The authors also introduce learnable activations that combine functions without -dependent gates.

The authoritative source is an arXiv record for “More Expressive Feedforward Layers: Part I. -Adaptive Mixing of Activations.” It says the paper was first submitted on May 26, 2026, and revised on August 28, placing the substantive update inside the current news window. The authors are Mingze Wang, Jinbo Wang, Yikuan Xia, Kai Shen, and Shu Zhong. The source identifies the work as a machine-learning and artificial-intelligence paper, with language models as its direct subject.

The paper starts from a limitation the authors identify in common feedforward network, or FFN, layers. Earlier designs used functions such as ReLU and GELU, while later gated designs include SwiGLU, but the source says most FFNs still apply one fixed nonlinear transformation to every . The proposed Mixture of Activations instead uses a dictionary of activation functions and lightweight gates whose values depend on the input. The same linear projections are shared, while the mixture of nonlinear functions can vary from token to token.

The source also describes learnable activations, abbreviated LA, as an input-independent counterpart. LA forms linear combinations of activation functions in both ReLU-type and SwiGLU-type FFNs. The authors state that their finite-width theory establishes strict expressive separations: LA strictly contains fixed-activation FFNs, while MoA strictly contains LA because it adds input-dependent nonlinear hybridization. These are theoretical expressivity claims about the functions the layers can represent, not evidence that a deployed language model is generally more capable or reliable.

For empirical testing, the authors say they conducted extensive pretraining experiments on dense and mixture-of-experts language models ranging from 0.12 billion to 2 billion parameters. They report testing different budgets, optimizers, and learning-rate schedules. According to the abstract, MoA consistently achieved lower terminal loss and showed more favorable scaling behavior than well-tuned baselines, with minimal parameter and computational overhead. The source record does not provide the numerical loss differences, the names of the datasets or baselines, hardware details, or the precise meaning of “minimal” overhead.

来源详情: arxiv.org ↗

为什么这很重要

Feedforward layers contain a large share of the parameters and nonlinear behavior in many language models. If -adaptive activation mixing improves training without materially increasing compute or parameters, it could offer a relatively small architectural change with implications for model scaling. The evidence remains limited to the authors’ reported pretraining experiments and does not establish better downstream performance or production efficiency.

The proposal targets a consequential part of current language-model architecture. The source says FFN layers account for a large fraction of model parameters and nonlinear expressivity. That makes them an important place to seek improvements: a change to the nonlinear transformation could affect how efficiently a model uses its existing width and projections, rather than requiring an entirely different model family.

MoA’s design is potentially practical because it retains shared linear projections and adds lightweight input-dependent gating. In principle, this could allow different tokens to receive different nonlinear treatment while preserving much of the surrounding FFN structure. However, the source does not establish that existing models can be upgraded without retraining, nor does it report implementation details sufficient to determine whether the added gates improve real-world throughput, memory use, or energy consumption.

The reported lower terminal loss is relevant because training loss is a central measure of how well a model fits its pretraining data. The claimed scaling behavior could also matter if the method continues to improve as model size or training compute increases. But terminal loss is not the same as usefulness to people. The source gives no results for factuality, reasoning, coding, multilingual performance, safety, calibration, robustness, or downstream task accuracy, so the practical significance of the reported training gains remains uncertain.

The evidence is also preliminary. It comes from a revised preprint, and the source record presents the authors’ theoretical and empirical conclusions rather than independent validation. The tested range ends at 2 billion parameters, which is materially smaller than many widely deployed language models. The record does not state whether the experiments used multiple random seeds, how the baselines were tuned, whether the method changes inference latency, or whether its gains persist outside the reported pretraining settings.

Interactive Mechanism

互动机制:它实际上是如何运作的

以交互方式探索这一发展背后的基础技术。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
交互式概念检查+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

接下来看什么

The important next evidence is the size and consistency of the reported gains, especially beyond the tested 0.12B-to-2B range. Readers should look for exact compute, memory, latency, and parameter overheads; results on downstream tasks; independent replication; and evidence that the method remains useful across architectures, datasets, optimizers, and longer training runs.

The next useful check is quantitative detail. The full paper or later work should show exact loss curves, confidence or run-to-run variation, parameter counts, operation counts, memory requirements, and wall-clock training costs for MoA, LA, and fixed-activation baselines. Because the abstract says the comparisons were against well-tuned baselines, the tuning procedure and compute budget will be important for judging whether the advantage comes from the architecture rather than unequal optimization.

Scale is another unresolved issue. The source reports models from 0.12B to 2B parameters, but it does not say whether the same pattern holds in substantially larger dense or mixture-of-experts systems. Future experiments should test larger models, longer training runs, additional budgets, and different data mixtures. They should also clarify whether the claimed favorable scaling behavior is measured by loss at a fixed compute budget, by loss at a fixed parameter count, or by another comparison.

Operational costs deserve close attention. MoA introduces a -dependent gate and a dictionary of activation functions, even if the added parameter and compute burden is described as minimal. Independent measurements should determine whether this produces meaningful changes in accelerator utilization, memory traffic, batching, inference latency, training stability, or energy use. A small theoretical overhead can have a larger systems cost if it disrupts efficient hardware execution.

Finally, readers should look for independent replication and broader evaluation. Results on downstream language, coding, reasoning, and multilingual tasks would show whether lower pretraining loss transfers to capabilities that users notice. Evaluations of reliability and safety would test whether -adaptive nonlinear behavior introduces new failure patterns. Until those results are available, the paper is best understood as a promising architectural research claim, not a demonstrated production improvement or an immediately available product.

相关指南和测验

人工智能模型解释变形金刚人工智能培训测试你所知道的——尝试免费的人工智能测验在我们的词汇表中查找人工智能术语关注 AI 模型发布跟踪器
觉得这有用吗?