Score
Designs and implements a dual-branch evaluation framework that trains two parallel model branches under controlled, identical settings (shared decoder architecture, training objective, optimizer, and curriculum) so differences can be attributed solely to differences in input modality or branch input. Builds experiments and analyses to isolate and quantify performance and robustness differences (for example under input corruption) between the branches.
This work addresses the challenge in large language model fine-tuning where parameter and data selection, often guided by independent scoring mechanisms, fail to coordinate effectively, leading to computational redundancy. The authors unify these two processes into a bilevel optimization framework sharing a common validation objective and propose DualSFT, a novel method that constructs a gradient interaction matrix to establish a row-column correspondence between parameter importance and data utility. This enables, for the first time, joint closed-form scoring and co-extraction of parameter masks and data subsets. By integrating first- and second-order validation approximations with a single-pass dual-scoring strategy, DualSFT significantly outperforms sequential baselines across 3B–9B scale models, simultaneously enhancing task performance and the trade-off between stability and plasticity under fixed computational budgets.
To address the vulnerability of deep neural networks to adversarial attacks, this paper proposes BORT, a multi-branch robust training framework that enhances adversarial robustness from the perspective of deep feature distribution. Its core innovation is a branch orthogonality loss that enforces orthogonality among feature subspaces of parallel branches, thereby improving invariance to input perturbations. Crucially, BORT requires only the original training data—introducing no auxiliary data, inference overhead, or pre-trained modules. Integrated with ℓ∞-constrained adversarial training, BORT is comprehensively evaluated on CIFAR-10, CIFAR-100, and SVHN. On CIFAR-10 and CIFAR-100, it achieves robust accuracies of 67.3% and 41.5%, respectively—surpassing state-of-the-art methods by 7.23% and 9.07%. These results significantly outperform all existing data-free approaches and even surpass several methods relying on large-scale auxiliary datasets.
This work addresses the challenge in large language models of simultaneously unlearning harmful information while preserving useful knowledge. To this end, the authors propose DualOptim, a framework that models a shared base state for general representations and employs an incremental state to capture task-specific residual updates for forgetting and retention. A gradient conflict-aware dual-state optimization mechanism dynamically balances these two objectives by aligning their update directions. Furthermore, the method is enhanced with low-bit quantization to yield DualOptim+ 8bit, a memory-efficient variant operating at 8-bit precision. Experimental results demonstrate that DualOptim consistently outperforms existing approaches across diverse scenarios—including synthetic and real-world unlearning, safety alignment, and multi-task learning—achieving a superior trade-off between effective forgetting and knowledge retention.
To address the high computational complexity and poor interpretability of cross-modal interactions in audio-visual speech recognition (AVSR) under noisy conditions, this paper introduces Branchformer—the first application of this architecture to AVSR—proposing a novel two-stage, customized unified encoder-decoder framework: “unimodal-first, then fusion.” By incorporating modality-specific branch scoring and layer-level structural pruning, the method achieves parameter-efficient and interpretable audio-visual joint modeling. Evaluated on multi-scenario English and Spanish benchmarks, it attains word error rates (WER) of 2.5% and 9.1%, respectively—significantly outperforming comparable large models while reducing parameter count substantially and achieving state-of-the-art performance. Key contributions include: (1) the pioneering adaptation of Branchformer to AVSR; (2) a new architectural paradigm that jointly optimizes model lightweighting and cross-modal interpretability; and (3) enhanced end-to-end cross-modal collaborative modeling capability.
Existing methods struggle to effectively compare internal representations across large language models with different architectures, hindering the discovery of potentially safety-critical behaviors in newly developed models. This work addresses this challenge by extending the Crosscoder framework to enable unsupervised cross-architectural contrastive analysis and introduces a Dedicated Feature Crosscoder (DFC) architecture designed to more precisely disentangle and identify behaviorally distinctive features unique to each model. Experiments conducted on models such as Qwen3, Llama3.1, and GPT-OSS successfully uncover meaningful behavioral differences—including political bias and copyright refusal mechanisms—in a fully unsupervised manner, thereby demonstrating the effectiveness and practical utility of the proposed approach.
This work addresses the suppression of minority-class learning in deep neural networks under severe class imbalance, which arises from shared feature representations. To mitigate this issue, the authors propose a Class-Specific Branch Attention (CSBA) mechanism. By analyzing inter-layer gradient flows and constructing a gradient conflict matrix based on class-specific gradient cosine similarity, they reveal—through the lens of optimization dynamics—for the first time how majority classes dominate and suppress gradients of minority classes. A lightweight channel reweighting module is integrated into a multi-branch convolutional architecture to implicitly decouple features and gradients in a class-aware manner. Experiments demonstrate that the proposed method significantly improves minority-class performance without compromising overall accuracy: the F1 score for the Physical-Damage class increases from 0.261 to 0.522, and Macro-F1 on CIFAR-10-LT rises from 0.595 to 0.655.
This work addresses the reliability challenges in large language models (LLMs) for test-time output prediction, where errors often stem from code execution failures or hallucinated pseudocode. To mitigate these issues, the authors propose DuET, a dual-execution framework that uniquely integrates direct execution of generated code with LLM-based simulated execution of pseudocode. By reconciling the outputs of both execution paths through a functional majority voting mechanism, DuET effectively compensates for the limitations inherent in either approach alone. Evaluated on the LiveCodeBench benchmark, the method demonstrates substantial performance gains, achieving a 13.6 percentage point improvement in Pass@1 over current state-of-the-art techniques, thereby exhibiting enhanced robustness and prediction accuracy.
Existing methods for generating unlearnable examples exhibit insufficient robustness against advanced defenses. This work proposes DUNE, the first approach to extend perturbation optimization into both spatial and color domains, jointly generating cross-domain perturbations through a dual-branch architecture. It further introduces an enhanced ensemble strategy based on pretrained models to significantly boost perturbation strength and generalization of interference. By effectively inducing models to learn spurious features, DUNE severely degrades their generalization performance. Evaluated on CIFAR-10 and ImageNet against seven state-of-the-art defense mechanisms, DUNE reduces the average test accuracy of victim models to 14.95%–50.82%, substantially outperforming twelve existing state-of-the-art methods.
This work addresses the representational imbalance between shared and private modalities in multimodal sentiment analysis, which often undermines modality-specific discriminability and cross-modal complementarity. To mitigate this issue, the authors propose a dual-branch rebalancing framework that integrates three key components: Temporal-Structural Factorization (TSF) to suppress redundant shared representations, Anchor-Guided Private Routing (AGPR) to enhance the discriminative capacity of private features, and Bidirectional Rebalancing Fusion (BRF) for context-aware representation integration. This approach is the first to systematically alleviate the imbalance between shared and private branches. Extensive experiments on CMU-MOSI, CMU-MOSEI, and MIntRec demonstrate substantial performance gains over current state-of-the-art baselines, confirming the efficacy of the proposed rebalancing mechanism in advancing multimodal sentiment analysis.