controlled dual-branch evaluation

Designs and implements a dual-branch evaluation framework that trains two parallel model branches under controlled, identical settings (shared decoder architecture, training objective, optimizer, and curriculum) so differences can be attributed solely to differences in input modality or branch input. Builds experiments and analyses to isolate and quantify performance and robustness differences (for example under input corruption) between the branches.

controlleddual-branchevaluation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.02
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge in large language model fine-tuning where parameter and data selection, often guided by independent scoring mechanisms, fail to coordinate effectively, leading to computational redundancy. The authors unify these two processes into a bilevel optimization framework sharing a common validation objective and propose DualSFT, a novel method that constructs a gradient interaction matrix to establish a row-column correspondence between parameter importance and data utility. This enables, for the first time, joint closed-form scoring and co-extraction of parameter masks and data subsets. By integrating first- and second-order validation approximations with a single-pass dual-scoring strategy, DualSFT significantly outperforms sequential baselines across 3B–9B scale models, simultaneously enhancing task performance and the trade-off between stability and plasticity under fixed computational budgets.

bilevel optimizationcoordinated selectiondata selection

Two Heads are Better than One: Robust Learning Meets Multi-branch Models

Aug 17, 2022
DH
Dong Huang
🏛️ National University of Singapore | The University of Hong Kong

To address the vulnerability of deep neural networks to adversarial attacks, this paper proposes BORT, a multi-branch robust training framework that enhances adversarial robustness from the perspective of deep feature distribution. Its core innovation is a branch orthogonality loss that enforces orthogonality among feature subspaces of parallel branches, thereby improving invariance to input perturbations. Crucially, BORT requires only the original training data—introducing no auxiliary data, inference overhead, or pre-trained modules. Integrated with ℓ∞-constrained adversarial training, BORT is comprehensively evaluated on CIFAR-10, CIFAR-100, and SVHN. On CIFAR-10 and CIFAR-100, it achieves robust accuracies of 67.3% and 41.5%, respectively—surpassing state-of-the-art methods by 7.23% and 9.07%. These results significantly outperform all existing data-free approaches and even surpass several methods relying on large-scale auxiliary datasets.

Enhancing adversarial training without additional dataImproving DNN robustness against adversarial examplesOptimizing multi-branch models for orthogonal solution spaces

This work addresses the challenge in large language models of simultaneously unlearning harmful information while preserving useful knowledge. To this end, the authors propose DualOptim, a framework that models a shared base state for general representations and employs an incremental state to capture task-specific residual updates for forgetting and retention. A gradient conflict-aware dual-state optimization mechanism dynamically balances these two objectives by aligning their update directions. Furthermore, the method is enhanced with low-bit quantization to yield DualOptim+ 8bit, a memory-efficient variant operating at 8-bit precision. Experimental results demonstrate that DualOptim consistently outperforms existing approaches across diverse scenarios—including synthetic and real-world unlearning, safety alignment, and multi-task learning—achieving a superior trade-off between effective forgetting and knowledge retention.

forgettinglarge language modelsmachine unlearning

Tailored Design of Audio-Visual Speech Recognition Models using Branchformers

Jul 09, 2024
DG
David Gimeno-G'omez
🏛️ Universitat Polit`ecnica de Val`encia

To address the high computational complexity and poor interpretability of cross-modal interactions in audio-visual speech recognition (AVSR) under noisy conditions, this paper introduces Branchformer—the first application of this architecture to AVSR—proposing a novel two-stage, customized unified encoder-decoder framework: “unimodal-first, then fusion.” By incorporating modality-specific branch scoring and layer-level structural pruning, the method achieves parameter-efficient and interpretable audio-visual joint modeling. Evaluated on multi-scenario English and Spanish benchmarks, it attains word error rates (WER) of 2.5% and 9.1%, respectively—significantly outperforming comparable large models while reducing parameter count substantially and achieving state-of-the-art performance. Key contributions include: (1) the pioneering adaptation of Branchformer to AVSR; (2) a new architectural paradigm that jointly optimizes model lightweighting and cross-modal interpretability; and (3) enhanced end-to-end cross-modal collaborative modeling capability.

Achieve state-of-the-art recognition rates efficiently.Optimize cross-modal architecture for AVSR.Reduce model complexity and computational cost.

Existing methods struggle to effectively compare internal representations across large language models with different architectures, hindering the discovery of potentially safety-critical behaviors in newly developed models. This work addresses this challenge by extending the Crosscoder framework to enable unsupervised cross-architectural contrastive analysis and introduces a Dedicated Feature Crosscoder (DFC) architecture designed to more precisely disentangle and identify behaviorally distinctive features unique to each model. Experiments conducted on models such as Qwen3, Llama3.1, and GPT-OSS successfully uncover meaningful behavioral differences—including political bias and copyright refusal mechanisms—in a fully unsupervised manner, thereby demonstrating the effectiveness and practical utility of the proposed approach.

behavioral differencescross-architecturelarge language models

Latest Papers

What's happening recently
View more

This work addresses the suppression of minority-class learning in deep neural networks under severe class imbalance, which arises from shared feature representations. To mitigate this issue, the authors propose a Class-Specific Branch Attention (CSBA) mechanism. By analyzing inter-layer gradient flows and constructing a gradient conflict matrix based on class-specific gradient cosine similarity, they reveal—through the lens of optimization dynamics—for the first time how majority classes dominate and suppress gradients of minority classes. A lightweight channel reweighting module is integrated into a multi-branch convolutional architecture to implicitly decouple features and gradients in a class-aware manner. Experiments demonstrate that the proposed method significantly improves minority-class performance without compromising overall accuracy: the F1 score for the Physical-Damage class increases from 0.261 to 0.522, and Macro-F1 on CIFAR-10-LT rises from 0.595 to 0.655.

class imbalancegradient interferenceminority-class learning

This work addresses the reliability challenges in large language models (LLMs) for test-time output prediction, where errors often stem from code execution failures or hallucinated pseudocode. To mitigate these issues, the authors propose DuET, a dual-execution framework that uniquely integrates direct execution of generated code with LLM-based simulated execution of pseudocode. By reconciling the outputs of both execution paths through a functional majority voting mechanism, DuET effectively compensates for the limitations inherent in either approach alone. Evaluated on the LiveCodeBench benchmark, the method demonstrates substantial performance gains, achieving a 13.6 percentage point improvement in Pass@1 over current state-of-the-art techniques, thereby exhibiting enhanced robustness and prediction accuracy.

code generationLLM reliabilitypseudocode execution

Existing methods for generating unlearnable examples exhibit insufficient robustness against advanced defenses. This work proposes DUNE, the first approach to extend perturbation optimization into both spatial and color domains, jointly generating cross-domain perturbations through a dual-branch architecture. It further introduces an enhanced ensemble strategy based on pretrained models to significantly boost perturbation strength and generalization of interference. By effectively inducing models to learn spurious features, DUNE severely degrades their generalization performance. Evaluated on CIFAR-10 and ImageNet against seven state-of-the-art defense mechanisms, DUNE reduces the average test accuracy of victim models to 14.95%–50.82%, substantially outperforming twelve existing state-of-the-art methods.

DefensesModel TrainingPerturbation

This work addresses the representational imbalance between shared and private modalities in multimodal sentiment analysis, which often undermines modality-specific discriminability and cross-modal complementarity. To mitigate this issue, the authors propose a dual-branch rebalancing framework that integrates three key components: Temporal-Structural Factorization (TSF) to suppress redundant shared representations, Anchor-Guided Private Routing (AGPR) to enhance the discriminative capacity of private features, and Bidirectional Rebalancing Fusion (BRF) for context-aware representation integration. This approach is the first to systematically alleviate the imbalance between shared and private branches. Extensive experiments on CMU-MOSI, CMU-MOSEI, and MIntRec demonstrate substantial performance gains over current state-of-the-art baselines, confirming the efficacy of the proposed rebalancing mechanism in advancing multimodal sentiment analysis.

Branch RedundancyModality HeterogeneityMultimodal Sentiment Analysis

Hot Scholars

LW

Lihong Wang

Associate Professor of Psychiatry,University of Connecticut Health Center
depressionstressfMRIcognitive impairment
WW

Wenshuo Wang

Professor, Beijing Institute of Technology (BIT) | Research Fellow, UC Berkeley, CMU, McGill
Human-Robot InteractionAutonomous DrivingBayesian LearningHuman Factors
SD

Shuangli Du

xi'an university of technology
deep learning