🤖 AI Summary
This work proposes a novel mixture-of-experts (MoE) adapter that overcomes the limitations of existing parameter-efficient fine-tuning methods, which are confined to fixed spatial or frequency domains and struggle to adapt to task-, layer-, or token-specific optimal representations. By introducing the fractional Fourier transform (FrFT) into the MoE architecture for the first time, each expert is equipped with a learnable FrFT order, enabling continuous interpolation between spatial and frequency domains and dynamic selection of the most compact low-rank update space. This design naturally induces expert decorrelation through learnable domain selection, substantially enhancing multi-task compositionality with minimal computational overhead. Experiments on LLaMA-3.1-8B and Qwen2.5-7B demonstrate consistent superiority over strong baselines such as FlyLoRA and FourierMoE across commonsense, mathematical, coding, and knowledge-intensive tasks, while maintaining low active parameter counts and revealing interpretable patterns of order specialization at both task and layer levels.
📝 Abstract
Parameter-efficient fine-tuning (PEFT) reparameterizes weight updates in a fixed basis: low-rank adapters operate in the spatial domain, while a recent line of spectral methods operates in a fixed Fourier domain. We argue that the choice of domain is itself a design degree of freedom that should be learned, and that no single basis is optimal across tasks, layers, or tokens. We introduce Fractional-Fourier Mixture of Experts, a mixture-of-experts adapter in which every expert carries a learnable fractional-Fourier order that continuously interpolates between the spatial domain (recovering vanilla LoRA) and the Fourier domain (recovering a spectral adapter). Routing tokens through experts that occupy different points on this spatial-spectral continuum lets the model place each low-rank update in the domain where it is most compact, and -- because fractional-Fourier operators of different orders are mutually incoherent -- makes the experts naturally decorrelated, which reduces interference and improves multi-task composition. The order is a single scalar per expert, trained with a separate optimizer, and the transform is computed with an $\mathcal{O}(d\log d)$ chirp--FFT surrogate, so Fractional-Fourier Mixture of Experts adds negligible cost over standard MoE-LoRA. Across commonsense, mathematical, code, and knowledge benchmarks on LLaMA-3.1-8B and Qwen2.5-7B, Fractional-Fourier Mixture of Experts improves over strong MoE-LoRA and spectral baselines -- including FlyLoRA, FourierMoE, and HMoRA -- while keeping the active-parameter budget small, and analysis shows that the learned orders specialize by task and layer in interpretable ways.