CircleMatch: Prototype Matching with Circular Temporal Statistics for Tiny Keyword Spotting
为解决在参数受限条件下提高关键词识别准确性的问题,提出CircleMatch框架,通过原型匹配和无参数的循环聚合方法处理时间信息。
为解决在参数受限条件下提高关键词识别准确性的问题,提出CircleMatch框架,通过原型匹配和无参数的循环聚合方法处理时间信息。
本文通过结合算子学习和流方法,提出了一种神经采样器,用于从随机微分方程的不变测度中高效采样,显著提高了高维问题的处理速度。
This work addresses the challenge of linking anonymized trajectories to user identities, a task hindered by the difficulty of effectively integrating multi-source heterogeneous mobility semantics and leveraging trajectory structural knowledge. To this end, the study introduces knowledge graph representation learning into the trajectory-user linkage problem for the first time, constructing a multi-relational mobility knowledge graph that models visit time, POI category, and transition speed as typed relations. These relations jointly constrain POI embeddings and incorporate higher-order co-occurrence patterns to enrich representations. A dual-branch classifier is further designed to fuse structural and sequential evidence. The proposed method achieves significant improvements in linkage accuracy, demonstrating particularly strong performance in scenarios involving sparse and overlapping trajectories.
This work addresses the performance limitations of unified models in multi-task affective behavior analysis, where discrepancies among valence-arousal regression, expression classification, and action unit detection hinder joint optimization. To overcome this, the authors propose a task-adaptive feature fusion strategy that leverages complementary frame-level features extracted from frozen DINOv2 ViT-L and DINOv3 ConvNeXt-base backbones. Each task is equipped with a dedicated prediction head and tailored fusion mechanism—such as gated or residual fusion—to avoid enforced architectural sharing. The framework further integrates temporal convolutions, LightGBM-based post-processing, and threshold calibration. Evaluated on the ABAW11 validation set, the method achieves an EXPR macro F1 of 0.4222, AU macro F1 of 0.5402, and VA mean CCC of 0.6717, yielding a total score of 1.6341 and demonstrating substantial improvement in multi-task collaborative performance.
This work addresses the stability-plasticity dilemma in multilingual speech recognition caused by data imbalance: fully shared parameters subject low-resource languages to negative interference, while entirely independent parameters impede cross-lingual knowledge transfer. To resolve this, the authors propose Zipper-LoRA, a novel framework featuring rank-level dynamic decoupling that enables fine-grained parameter sharing and disentanglement within the LoRA subspace via lightweight language-conditioned routers. The approach incorporates Static, Hard, and Soft variants alongside a two-stage training strategy—including an Initial-B warm start—to share parameters when compatible and decouple them when conflicting. Experiments across 12 languages with mixed resource levels demonstrate that Zipper-LoRA significantly outperforms both fully shared and fully independent baselines, with especially pronounced gains under extremely low-resource conditions, and maintains robust performance across both chunked and non-chunked encoder configurations.
为解决在参数受限条件下提高关键词识别准确性的问题,提出CircleMatch框架,通过原型匹配和无参数的循环聚合方法处理时间信息。
本文通过结合算子学习和流方法,提出了一种神经采样器,用于从随机微分方程的不变测度中高效采样,显著提高了高维问题的处理速度。
This work addresses the challenge of linking anonymized trajectories to user identities, a task hindered by the difficulty of effectively integrating multi-source heterogeneous mobility semantics and leveraging trajectory structural knowledge. To this end, the study introduces knowledge graph representation learning into the trajectory-user linkage problem for the first time, constructing a multi-relational mobility knowledge graph that models visit time, POI category, and transition speed as typed relations. These relations jointly constrain POI embeddings and incorporate higher-order co-occurrence patterns to enrich representations. A dual-branch classifier is further designed to fuse structural and sequential evidence. The proposed method achieves significant improvements in linkage accuracy, demonstrating particularly strong performance in scenarios involving sparse and overlapping trajectories.
This work addresses the performance limitations of unified models in multi-task affective behavior analysis, where discrepancies among valence-arousal regression, expression classification, and action unit detection hinder joint optimization. To overcome this, the authors propose a task-adaptive feature fusion strategy that leverages complementary frame-level features extracted from frozen DINOv2 ViT-L and DINOv3 ConvNeXt-base backbones. Each task is equipped with a dedicated prediction head and tailored fusion mechanism—such as gated or residual fusion—to avoid enforced architectural sharing. The framework further integrates temporal convolutions, LightGBM-based post-processing, and threshold calibration. Evaluated on the ABAW11 validation set, the method achieves an EXPR macro F1 of 0.4222, AU macro F1 of 0.5402, and VA mean CCC of 0.6717, yielding a total score of 1.6341 and demonstrating substantial improvement in multi-task collaborative performance.
This work addresses the stability-plasticity dilemma in multilingual speech recognition caused by data imbalance: fully shared parameters subject low-resource languages to negative interference, while entirely independent parameters impede cross-lingual knowledge transfer. To resolve this, the authors propose Zipper-LoRA, a novel framework featuring rank-level dynamic decoupling that enables fine-grained parameter sharing and disentanglement within the LoRA subspace via lightweight language-conditioned routers. The approach incorporates Static, Hard, and Soft variants alongside a two-stage training strategy—including an Initial-B warm start—to share parameters when compatible and decouple them when conflicting. Experiments across 12 languages with mixed resource levels demonstrate that Zipper-LoRA significantly outperforms both fully shared and fully independent baselines, with especially pronounced gains under extremely low-resource conditions, and maintains robust performance across both chunked and non-chunked encoder configurations.