Institution profile

Inventec Corporation

Industry researchasia · tw
Official website
Research library12linked papers
Opportunities0open roles
Selected work

Representative Papers

Training-Free Instruction TTS Gender Bias Calibration Using Model-Adaptive Steering

Oct 07, 2026

Instruction-based text-to-speech (ITTS) systems exhibit implicit gender biases that are difficult to calibrate without retraining. This work proposes a model-adaptive steering method that directly adjusts post-encoder representations for training-free bias correction through group bias vectors, coarse-to-fine intensity search, and deterministic lexical gating. To our knowledge, this mechanism is the first to operate without retraining while generalizing across heterogeneous architectures, eliminating implicit biases while strictly preserving explicit prompt semantics. Experimental results demonstrate that the proposed approach reduces aggregate calibration error by 0.8 to 5.9 percentage points across four mainstream ITTS models, with negligible degradation in audio quality and intelligibility.

0 citationsRead paper

TSMD: Temporal-Stream Modality Dropout for Robust Video Highlight Detection

Sep 30, 2026

This study addresses the limited robustness caused by missing modalities in multimodal video highlight detection and the misalignment between mean squared error (MSE) loss and evaluation metrics. To tackle these issues, we propose a temporal-stream-level mixed modality dropout strategy that enhances model resilience to interference by simulating structured modality absence. Furthermore, we design a joint loss function integrating MSE, Pearson correlation, and RankNet to precisely align with the peak localization characteristics of highlights. Experimental results demonstrate that the proposed method achieves significant improvements in mAP@15 on the MoSu and Mr. HiSum datasets. Notably, it substantially outperforms baseline models under severe modality-missing conditions, validating its effectiveness and robustness for real-world multimodal video highlight detection scenarios.

0 citationsRead paper

DuoAD: Leveraging [CLS] Dual Characteristics for Training-Free Few-Shot Anomaly Detection

Jul 26, 2026

This work proposes a fully automatic, training-free few-shot anomaly detection framework that addresses the limitation of existing methods, which predominantly rely on local image patch features while overlooking the global contextual information embedded in the [CLS] token of Vision Transformers. The study is the first to reveal and exploit the dual nature of the [CLS] token: its global semantic invariance and its attention map’s ability to indicate spatial anomalies. By integrating a semantic consistency-driven automatic augmentation strategy with an attention-guided dynamic feature reweighting mechanism, the method achieves precise anomaly localization and scoring without manual hyperparameter tuning. Under single-sample settings on MVTec-AD, VisA, and Real-IAD, it attains Image-AUC scores of 97.7%, 93.2%, and 84.5%, respectively, demonstrating plug-and-play state-of-the-art performance across categories, backbone architectures, and datasets.

0 citationsRead paper

The Binding Effect: Analyzing How Multi-Dimensional Cues Form Gender Bias in Instruction TTS

Mar 21, 2026

This study addresses the limitations of existing gender bias evaluations in instructional text-to-speech (ITTS) systems, which often rely on univariate tests and fail to capture the combinatorial effects of social cues. The authors propose a multidimensional prompting framework that systematically integrates social status, occupational stereotypes, and role descriptors, revealing for the first time a binding effect among these dimensions in ITTS outputs. This binding effect indicates that bias arises from a deep coupling between semantic priors embedded in pretrained text encoders and the distributional properties of training data. Through analyses of open-source models, semantic probing, and diversity intervention experiments, the work demonstrates that generic diversity prompts are insufficient to mitigate such entrenched biases, underscoring the necessity of compositional analysis for diagnosing latent risks in synthetic speech and establishing a critical link between semantic priors in pretrained encoders and biased voice generation.

0 citationsRead paper

Feasibility-Guided Planning over Multi-Specialized Locomotion Policies

Feb 08, 2026

This work proposes a feasibility-guided path planning framework to address the challenge of coordinating multiple specialized locomotion strategies over unstructured terrain. The approach equips each terrain-specific strategy with a lightweight Feasibility-Net that predicts a feasibility tensor from local elevation maps and task vectors, thereby guiding classical planning algorithms to generate optimal paths consistent with the capabilities of the selected strategy. The framework supports plug-and-play integration of new strategies without retraining, while preserving both interpretability and strategy consistency. Experimental results in both simulation and real-world environments demonstrate that the method efficiently produces reliable paths and significantly enhances adaptability to complex and diverse terrains.

0 citationsRead paper
Recent publications

Latest Papers

Training-Free Instruction TTS Gender Bias Calibration Using Model-Adaptive Steering

Oct 07, 2026

Instruction-based text-to-speech (ITTS) systems exhibit implicit gender biases that are difficult to calibrate without retraining. This work proposes a model-adaptive steering method that directly adjusts post-encoder representations for training-free bias correction through group bias vectors, coarse-to-fine intensity search, and deterministic lexical gating. To our knowledge, this mechanism is the first to operate without retraining while generalizing across heterogeneous architectures, eliminating implicit biases while strictly preserving explicit prompt semantics. Experimental results demonstrate that the proposed approach reduces aggregate calibration error by 0.8 to 5.9 percentage points across four mainstream ITTS models, with negligible degradation in audio quality and intelligibility.

0 citationsRead paper

TSMD: Temporal-Stream Modality Dropout for Robust Video Highlight Detection

Sep 30, 2026

This study addresses the limited robustness caused by missing modalities in multimodal video highlight detection and the misalignment between mean squared error (MSE) loss and evaluation metrics. To tackle these issues, we propose a temporal-stream-level mixed modality dropout strategy that enhances model resilience to interference by simulating structured modality absence. Furthermore, we design a joint loss function integrating MSE, Pearson correlation, and RankNet to precisely align with the peak localization characteristics of highlights. Experimental results demonstrate that the proposed method achieves significant improvements in mAP@15 on the MoSu and Mr. HiSum datasets. Notably, it substantially outperforms baseline models under severe modality-missing conditions, validating its effectiveness and robustness for real-world multimodal video highlight detection scenarios.

0 citationsRead paper

DuoAD: Leveraging [CLS] Dual Characteristics for Training-Free Few-Shot Anomaly Detection

Jul 26, 2026

This work proposes a fully automatic, training-free few-shot anomaly detection framework that addresses the limitation of existing methods, which predominantly rely on local image patch features while overlooking the global contextual information embedded in the [CLS] token of Vision Transformers. The study is the first to reveal and exploit the dual nature of the [CLS] token: its global semantic invariance and its attention map’s ability to indicate spatial anomalies. By integrating a semantic consistency-driven automatic augmentation strategy with an attention-guided dynamic feature reweighting mechanism, the method achieves precise anomaly localization and scoring without manual hyperparameter tuning. Under single-sample settings on MVTec-AD, VisA, and Real-IAD, it attains Image-AUC scores of 97.7%, 93.2%, and 84.5%, respectively, demonstrating plug-and-play state-of-the-art performance across categories, backbone architectures, and datasets.

0 citationsRead paper

The Binding Effect: Analyzing How Multi-Dimensional Cues Form Gender Bias in Instruction TTS

Mar 21, 2026

This study addresses the limitations of existing gender bias evaluations in instructional text-to-speech (ITTS) systems, which often rely on univariate tests and fail to capture the combinatorial effects of social cues. The authors propose a multidimensional prompting framework that systematically integrates social status, occupational stereotypes, and role descriptors, revealing for the first time a binding effect among these dimensions in ITTS outputs. This binding effect indicates that bias arises from a deep coupling between semantic priors embedded in pretrained text encoders and the distributional properties of training data. Through analyses of open-source models, semantic probing, and diversity intervention experiments, the work demonstrates that generic diversity prompts are insufficient to mitigate such entrenched biases, underscoring the necessity of compositional analysis for diagnosing latent risks in synthetic speech and establishing a critical link between semantic priors in pretrained encoders and biased voice generation.

0 citationsRead paper

Feasibility-Guided Planning over Multi-Specialized Locomotion Policies

Feb 08, 2026

This work proposes a feasibility-guided path planning framework to address the challenge of coordinating multiple specialized locomotion strategies over unstructured terrain. The approach equips each terrain-specific strategy with a lightweight Feasibility-Net that predicts a feasibility tensor from local elevation maps and task vectors, thereby guiding classical planning algorithms to generate optimal paths consistent with the capabilities of the selected strategy. The framework supports plug-and-play integration of new strategies without retraining, while preserving both interpretability and strategy consistency. Experimental results in both simulation and real-world environments demonstrate that the method efficiently produces reliable paths and significantly enhances adaptability to complex and diverse terrains.

0 citationsRead paper