Muon Can Outperform Dedicated Continual Learning Methods
研究使用Muon优化器正交化更新以解决持续学习中的遗忘问题,对比了IncLoRA、O-LoRA和ELLA方法,发现Muon在标准CL基准上表现良好。
研究使用Muon优化器正交化更新以解决持续学习中的遗忘问题,对比了IncLoRA、O-LoRA和ELLA方法,发现Muon在标准CL基准上表现良好。
研究解决了等变网络中Adam优化器性能不佳的问题,通过单独归一化每个块更新的方法,提高了其在等变线性层中的表现。
本文探讨了使用Hyper-V套接字作为恶意软件分析沙箱的实时数据提取通道,对比了其与WinSock TCP套接字在阻塞和枚举方面的优势,并比较了两者的数据吞吐量。
While large foundation models demonstrate strong performance in solving time-dependent partial differential equations, their high computational cost limits their practicality as replacements for efficient numerical solvers. This work proposes the Teacher Rollout Extension (TREX) framework, which leverages knowledge distillation to transfer capabilities from a pretrained teacher model to a lightweight student model. By using long-horizon synthetic trajectories generated by the teacher to augment limited downstream data, TREX enables sampling of rollout trajectories without requiring prior knowledge of the initial condition distribution. This exposes the student model to both long-term dynamics and local recovery behaviors, while allowing integration of task-specific inductive biases—such as equivariance. Combined with noise injection and an equivariant network architecture, the resulting student model achieves several orders of magnitude fewer parameters, over tenfold faster inference, and accuracy comparable to or exceeding that of the teacher.
This work investigates the intrinsic mechanisms by which foundation models detect deepfakes, addressing why pretrained representations effectively distinguish authentic from synthetic media. The study reveals that forged samples consistently elicit lower-magnitude feature responses across diverse foundation models and systematically demonstrates— for the first time—that this amplitude discrepancy serves as a key signal for authenticity verification, rooted in semantic shift. Building on this insight, the authors reformulate deepfake detection as an anomaly detection task, showing that simple statistics of feature magnitudes alone enable efficient zero-shot detection. The proposed approach achieves performance on par with complex specialized models across both image and video modalities, with detection capability scaling favorably with model size, thereby confirming that large-scale foundation models inherently possess strong zero-shot potential for forgery identification.
研究使用Muon优化器正交化更新以解决持续学习中的遗忘问题,对比了IncLoRA、O-LoRA和ELLA方法,发现Muon在标准CL基准上表现良好。
研究解决了等变网络中Adam优化器性能不佳的问题,通过单独归一化每个块更新的方法,提高了其在等变线性层中的表现。
本文探讨了使用Hyper-V套接字作为恶意软件分析沙箱的实时数据提取通道,对比了其与WinSock TCP套接字在阻塞和枚举方面的优势,并比较了两者的数据吞吐量。
While large foundation models demonstrate strong performance in solving time-dependent partial differential equations, their high computational cost limits their practicality as replacements for efficient numerical solvers. This work proposes the Teacher Rollout Extension (TREX) framework, which leverages knowledge distillation to transfer capabilities from a pretrained teacher model to a lightweight student model. By using long-horizon synthetic trajectories generated by the teacher to augment limited downstream data, TREX enables sampling of rollout trajectories without requiring prior knowledge of the initial condition distribution. This exposes the student model to both long-term dynamics and local recovery behaviors, while allowing integration of task-specific inductive biases—such as equivariance. Combined with noise injection and an equivariant network architecture, the resulting student model achieves several orders of magnitude fewer parameters, over tenfold faster inference, and accuracy comparable to or exceeding that of the teacher.
This work investigates the intrinsic mechanisms by which foundation models detect deepfakes, addressing why pretrained representations effectively distinguish authentic from synthetic media. The study reveals that forged samples consistently elicit lower-magnitude feature responses across diverse foundation models and systematically demonstrates— for the first time—that this amplitude discrepancy serves as a key signal for authenticity verification, rooted in semantic shift. Building on this insight, the authors reformulate deepfake detection as an anomaly detection task, showing that simple statistics of feature magnitudes alone enable efficient zero-shot detection. The proposed approach achieves performance on par with complex specialized models across both image and video modalities, with detection capability scaling favorably with model size, thereby confirming that large-scale foundation models inherently possess strong zero-shot potential for forgery identification.