MixiMotion: One-Step Text-to-Motion Generation via Asymmetric Set Distillation
MixiMotion通过离线集合蒸馏和不对称双向匹配方法,实现了一步生成文本到动作的转换,提高了效率同时保持高质量。
MixiMotion通过离线集合蒸馏和不对称双向匹配方法,实现了一步生成文本到动作的转换,提高了效率同时保持高质量。
本文针对越南2025年高考非线性评分系统下部分正确答案得分问题,提出THPT-Ladder基准来准确评估语言模型表现。
This work addresses the challenge of computing fixed points of contractive mappings without requiring prior knowledge of the contraction factor or manual hyperparameter tuning. We propose a fully adaptive Halpern-type algorithm that operates without line searches, bisection procedures, or any user-specified parameters, automatically exploiting the inherent contractiveness of the mapping to achieve explicit linear convergence even in the absence of a priori estimates of the contraction constant. Theoretical analysis establishes linear convergence rates both in terms of fixed-point residuals and distance to the solution, with an iteration complexity of 𝒪(ε⁻¹ln(ε⁻¹)) for solving cocoercive equations. By integrating Tikhonov regularization and Nesterov acceleration, we further extend the algorithm’s applicability. Numerical experiments confirm its superiority over existing adaptive methods, offering strong theoretical guarantees alongside low computational overhead.
This study addresses the significant limitations of large language models in understanding Vietnamese figurative expressions—such as idioms and proverbs—that are deeply rooted in cultural context. To this end, the authors introduce VIVID, the first systematic evaluation benchmark for Vietnamese figurative language, comprising 1,636 annotated expressions labeled for complexity and semantic themes. They propose an evaluation framework integrating generative and discriminative tasks, augmented with a human-validated, aspect-level LLM-as-a-Judge mechanism (Cohen’s κ = 0.792). Experiments on eight state-of-the-art models reveal that Vietnamese-specific models substantially underperform multilingual counterparts (e.g., VinaLLaMA-7B scores 0.13 versus GPT-4o’s 2.46 on a 5-point scale), with all models scoring below 50% of the maximum. Notably, few-shot prompting degrades GPT-4o’s performance due to stylistic overfitting, underscoring a systemic deficiency in cultural pragmatic comprehension.
This work addresses the challenge of balancing cross-domain adaptation and zero-shot generalization in zero-shot sketch-based image retrieval (ZS-SBIR). To this end, the authors propose a semantic-consistent prompt learning framework that effectively adapts the CLIP model through a text-guided intermediate representation injection mechanism and a perturbation-based asymmetric consistency constraint. This approach preserves CLIP’s inherent generalization capability while significantly enhancing retrieval performance. The framework incorporates learnable prompt vectors, a cross-modal coupling function, and a lightweight adapter module, optimized jointly via a multi-objective loss combining triplet, NT-Xent, and classification losses. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance across all three standard ZS-SBIR benchmarks—within-category, generalized, and cross-dataset settings.
MixiMotion通过离线集合蒸馏和不对称双向匹配方法,实现了一步生成文本到动作的转换,提高了效率同时保持高质量。
本文针对越南2025年高考非线性评分系统下部分正确答案得分问题,提出THPT-Ladder基准来准确评估语言模型表现。
This work addresses the challenge of computing fixed points of contractive mappings without requiring prior knowledge of the contraction factor or manual hyperparameter tuning. We propose a fully adaptive Halpern-type algorithm that operates without line searches, bisection procedures, or any user-specified parameters, automatically exploiting the inherent contractiveness of the mapping to achieve explicit linear convergence even in the absence of a priori estimates of the contraction constant. Theoretical analysis establishes linear convergence rates both in terms of fixed-point residuals and distance to the solution, with an iteration complexity of 𝒪(ε⁻¹ln(ε⁻¹)) for solving cocoercive equations. By integrating Tikhonov regularization and Nesterov acceleration, we further extend the algorithm’s applicability. Numerical experiments confirm its superiority over existing adaptive methods, offering strong theoretical guarantees alongside low computational overhead.
This study addresses the significant limitations of large language models in understanding Vietnamese figurative expressions—such as idioms and proverbs—that are deeply rooted in cultural context. To this end, the authors introduce VIVID, the first systematic evaluation benchmark for Vietnamese figurative language, comprising 1,636 annotated expressions labeled for complexity and semantic themes. They propose an evaluation framework integrating generative and discriminative tasks, augmented with a human-validated, aspect-level LLM-as-a-Judge mechanism (Cohen’s κ = 0.792). Experiments on eight state-of-the-art models reveal that Vietnamese-specific models substantially underperform multilingual counterparts (e.g., VinaLLaMA-7B scores 0.13 versus GPT-4o’s 2.46 on a 5-point scale), with all models scoring below 50% of the maximum. Notably, few-shot prompting degrades GPT-4o’s performance due to stylistic overfitting, underscoring a systemic deficiency in cultural pragmatic comprehension.
This work addresses the challenge of balancing cross-domain adaptation and zero-shot generalization in zero-shot sketch-based image retrieval (ZS-SBIR). To this end, the authors propose a semantic-consistent prompt learning framework that effectively adapts the CLIP model through a text-guided intermediate representation injection mechanism and a perturbation-based asymmetric consistency constraint. This approach preserves CLIP’s inherent generalization capability while significantly enhancing retrieval performance. The framework incorporates learnable prompt vectors, a cross-modal coupling function, and a lightweight adapter module, optimized jointly via a multi-objective loss combining triplet, NT-Xent, and classification losses. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance across all three standard ZS-SBIR benchmarks—within-category, generalized, and cross-dataset settings.