Researchers propose token-adaptive activation mixing for Transformer feedforward layers
A revised arXiv preprint proposes Mixture of Activations, a feedforward-layer design that lets language models choose among activation functions for each token. The authors report lower terminal training loss in models from 0.12B to 2B parameters, but provide no exact results in the source record.