Retour aux Actualités
InnovationBriefing AI Understanding

Les chercheurs proposent un mélange d'activation adaptatif aux jetons pour les couches de rétroaction de Transformer

Une préimpression révisée de arXiv propose un mélange d'activations, une conception de couche de rétroaction qui permet aux modèles de langage de choisir parmi les fonctions d'activation pour chaque jeton. Les auteurs signalent une perte de formation terminale plus faible dans les modèles allant de 0,12B à 2B, mais ne fournissent aucun résultat exact dans l'enregistrement source.

5 min readRead the primary source
Source-page capture accompanying Researchers propose token-adaptive activation mixing for Transformer feedforward layers
Document de source principaleSource enregistrée
Éditeur
arxiv.org
Lien source
arxiv.orghttps://arxiv.org/abs/2605.26647
Type de source
Document principal : une annonce officielle, un document, un dépôt ou une page de première partie que nous lisons directement.
ContexteComprenez cela en 60 secondes

Commencez ici

Termes clés

Transformateur
Une architecture neuronale qui utilise l'attention pour modéliser les relations entre les séquences en parallèle.
Jeton
Morceau de texte traité par des modèles de langage, tel qu'un mot ou un symbole.
Mémoire (mémoire de l'agent)
Contexte stocké qu'un agent IA utilise au fil des étapes ou des sessions pour améliorer la continuité.
Testez-vousQuiz sur les modèles d'IA expliqués

Que s'est-il passé

A paper revised on arXiv on August 28 proposes Mixture of Activations, or MoA, for -based language models. The design uses lightweight, input-dependent gates to mix several activation functions for individual tokens while sharing the same linear projections. The authors also introduce learnable activations that combine functions without -dependent gates.

The authoritative source is an arXiv record for “More Expressive Feedforward Layers: Part I. -Adaptive Mixing of Activations.” It says the paper was first submitted on May 26, 2026, and revised on August 28, placing the substantive update inside the current news window. The authors are Mingze Wang, Jinbo Wang, Yikuan Xia, Kai Shen, and Shu Zhong. The source identifies the work as a machine-learning and artificial-intelligence paper, with language models as its direct subject.

The paper starts from a limitation the authors identify in common feedforward network, or FFN, layers. Earlier designs used functions such as ReLU and GELU, while later gated designs include SwiGLU, but the source says most FFNs still apply one fixed nonlinear transformation to every . The proposed Mixture of Activations instead uses a dictionary of activation functions and lightweight gates whose values depend on the input. The same linear projections are shared, while the mixture of nonlinear functions can vary from token to token.

The source also describes learnable activations, abbreviated LA, as an input-independent counterpart. LA forms linear combinations of activation functions in both ReLU-type and SwiGLU-type FFNs. The authors state that their finite-width theory establishes strict expressive separations: LA strictly contains fixed-activation FFNs, while MoA strictly contains LA because it adds input-dependent nonlinear hybridization. These are theoretical expressivity claims about the functions the layers can represent, not evidence that a deployed language model is generally more capable or reliable.

For empirical testing, the authors say they conducted extensive pretraining experiments on dense and mixture-of-experts language models ranging from 0.12 billion to 2 billion parameters. They report testing different budgets, optimizers, and learning-rate schedules. According to the abstract, MoA consistently achieved lower terminal loss and showed more favorable scaling behavior than well-tuned baselines, with minimal parameter and computational overhead. The source record does not provide the numerical loss differences, the names of the datasets or baselines, hardware details, or the precise meaning of “minimal” overhead.

Détails de la source: arxiv.org ↗

Pourquoi c'est important

Feedforward layers contain a large share of the parameters and nonlinear behavior in many language models. If -adaptive activation mixing improves training without materially increasing compute or parameters, it could offer a relatively small architectural change with implications for model scaling. The evidence remains limited to the authors’ reported pretraining experiments and does not establish better downstream performance or production efficiency.

The proposal targets a consequential part of current language-model architecture. The source says FFN layers account for a large fraction of model parameters and nonlinear expressivity. That makes them an important place to seek improvements: a change to the nonlinear transformation could affect how efficiently a model uses its existing width and projections, rather than requiring an entirely different model family.

MoA’s design is potentially practical because it retains shared linear projections and adds lightweight input-dependent gating. In principle, this could allow different tokens to receive different nonlinear treatment while preserving much of the surrounding FFN structure. However, the source does not establish that existing models can be upgraded without retraining, nor does it report implementation details sufficient to determine whether the added gates improve real-world throughput, memory use, or energy consumption.

The reported lower terminal loss is relevant because training loss is a central measure of how well a model fits its pretraining data. The claimed scaling behavior could also matter if the method continues to improve as model size or training compute increases. But terminal loss is not the same as usefulness to people. The source gives no results for factuality, reasoning, coding, multilingual performance, safety, calibration, robustness, or downstream task accuracy, so the practical significance of the reported training gains remains uncertain.

The evidence is also preliminary. It comes from a revised preprint, and the source record presents the authors’ theoretical and empirical conclusions rather than independent validation. The tested range ends at 2 billion parameters, which is materially smaller than many widely deployed language models. The record does not state whether the experiments used multiple random seeds, how the baselines were tuned, whether the method changes inference latency, or whether its gains persist outside the reported pretraining settings.

Interactive Mechanism

Mécanisme interactif : comment cela fonctionne réellement

Explorez de manière interactive la technologie sous-jacente à ce développement.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Vérification de concept interactive+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Que regarder ensuite

The important next evidence is the size and consistency of the reported gains, especially beyond the tested 0.12B-to-2B range. Readers should look for exact compute, memory, latency, and parameter overheads; results on downstream tasks; independent replication; and evidence that the method remains useful across architectures, datasets, optimizers, and longer training runs.

The next useful check is quantitative detail. The full paper or later work should show exact loss curves, confidence or run-to-run variation, parameter counts, operation counts, memory requirements, and wall-clock training costs for MoA, LA, and fixed-activation baselines. Because the abstract says the comparisons were against well-tuned baselines, the tuning procedure and compute budget will be important for judging whether the advantage comes from the architecture rather than unequal optimization.

Scale is another unresolved issue. The source reports models from 0.12B to 2B parameters, but it does not say whether the same pattern holds in substantially larger dense or mixture-of-experts systems. Future experiments should test larger models, longer training runs, additional budgets, and different data mixtures. They should also clarify whether the claimed favorable scaling behavior is measured by loss at a fixed compute budget, by loss at a fixed parameter count, or by another comparison.

Operational costs deserve close attention. MoA introduces a -dependent gate and a dictionary of activation functions, even if the added parameter and compute burden is described as minimal. Independent measurements should determine whether this produces meaningful changes in accelerator utilization, memory traffic, batching, inference latency, training stability, or energy use. A small theoretical overhead can have a larger systems cost if it disrupts efficient hardware execution.

Finally, readers should look for independent replication and broader evaluation. Results on downstream language, coding, reasoning, and multilingual tasks would show whether lower pretraining loss transfers to capabilities that users notice. Evaluations of reliability and safety would test whether -adaptive nonlinear behavior introduces new failure patterns. Until those results are available, the paper is best understood as a promising architectural research claim, not a demonstrated production improvement or an immediately available product.

Guides et quiz associés

Modèles d'IA expliquésTransformateursFormation IATestez ce que vous savez : essayez un quiz gratuit sur l'IARecherchez un terme d'IA dans notre glossaireSuivez le suivi des versions du modèle AI
Vous avez trouvé cela utile ?