Thanks for the great survey! Could you please include a discussion of this work from Microsoft and UIUC? It proposes a general modular activation mechanism, SMA, that unifies previous works on MoE, adaptive computation, dynamic routing and sparse attention, and further applies SMA to develop a novel architecture, SeqBoat, to achieve SoTA quality-efficiency trade-off on Long Range Arena.
Thanks for the great survey! Could you please include a discussion of this work from Microsoft and UIUC? It proposes a general modular activation mechanism, SMA, that unifies previous works on MoE, adaptive computation, dynamic routing and sparse attention, and further applies SMA to develop a novel architecture, SeqBoat, to achieve SoTA quality-efficiency trade-off on Long Range Arena.