Institution profile

Appier Inc.

Industry researchasia · tw
Official website
Research library9linked papers
Opportunities0open roles
Selected work

Representative Papers

$α$Transfer: Coefficient Transfer for Efficient Model Merging

Oct 06, 2026

This study addresses the prohibitive computational and memory overhead incurred by coefficient search as model merging scales up. To this end, we propose $\alpha$Transfer, a paradigm grounded in the assumption that coefficient distributions remain consistent within a model family. This method decouples the coefficient search from large models to small proxy models, achieving efficient merging through parameter arithmetic, performance distribution analysis, and cross-model transfer techniques. Experimental results demonstrate that $\alpha$Transfer yields a 6× speedup with 70% memory reduction on Vision Transformers, and a 20× speedup with 85% memory savings on large language models, all while maintaining performance comparable to the original methods.

0 citationsRead paper

A Broader Look at Model Merging: Rethinking Implicit Regularization Induced by Task Arithmetic

Oct 06, 2026

This study addresses the limitation in model merging where implicit regularization introduced during coefficient search constrains weights to a restricted subspace, thereby hindering multi-task performance improvements. To overcome this, we re-examine the implicit regularization mechanism in task arithmetic and propose directly searching the pre-trained weight space via unconstrained optimization, breaking through the subspace constraints inherent in traditional linear combinations. Our work reveals and eliminates these implicit regularization effects, demonstrating that superior solutions reside outside conventional subspaces. Across diverse architectures and extremely data-scarce scenarios, the proposed method significantly outperforms existing model merging techniques, establishing a new optimization paradigm for multi-task learning.

0 citationsRead paper

PRISM: A Geometric Risk Bound that Decomposes Drift into Scale, Shape, and Head

May 12, 2026

Existing methods struggle to diagnose representation drift in post-training variants of large language models—such as quantization or LoRA fine-tuning—as they can only assess performance degradation without identifying root causes or guiding mitigation. This work proposes PRISM, a method that leverages the linear output head and near-isometric backbone structure of LLMs to derive a closed-form upper bound on cross-entropy risk discrepancy. PRISM uniquely decomposes representation drift into three geometrically interpretable and independently measurable axes: scale, shape, and head. Notably, the shape component is differentiable and can be employed as a regularizer to mitigate catastrophic forgetting. Experiments show that PRISM achieves Spearman correlation coefficients of 0.820 and 0.831 in risk ranking for quantized and LoRA-adapted models, respectively, across two model families and five benchmarks; moreover, shape-based regularization outperforms experience replay in alleviating downstream forgetting.

0 citationsRead paper

On Calibration of Large Language Models: From Response To Capability

Feb 14, 2026

This work addresses a critical limitation in existing calibration methods for large language models (LLMs), which focus on the correctness of individual responses and often fail to reflect the model’s overall task-solving capability, leading to a misalignment between confidence and actual performance. To bridge this gap, we propose a novel paradigm—capability calibration—that estimates the expected accuracy of an LLM on a given query, thereby shifting the focus from response-level to task-level reliability assessment. We formally distinguish capability calibration from traditional response calibration, develop a theoretical framework grounded in the stochasticity of LLM decoding, and systematically evaluate various confidence estimation methods under this new paradigm. Experiments demonstrate that capability calibration substantially improves the accuracy of pass@$k$ prediction and enhances the efficiency of reasoning resource allocation, offering a more reliable foundation for downstream applications.

0 citationsRead paper

Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models?

May 23, 2025

This study identifies an implicit language bias in large reasoning models (LRMs): when processing multilingual inputs, LRMs default to high-resource languages—particularly English—for internal reasoning, severely degrading performance on low-resource language tasks. To systematically investigate this, we design a multidimensional, controllable evaluation framework covering MMMLU, MATH-500, CulturalBench, and LMSYS-toxic, integrating reasoning-path tracing and cross-lingual attribution analysis. We empirically demonstrate that enforcing same-language reasoning—though it reduces general reasoning capability (especially for low-resource languages)—significantly improves cultural alignment and language-specific accuracy in safety evaluation. Crucially, this work is the first to empirically establish “reasoning-language–input-language mismatch” as a fundamental bottleneck to multilingual fairness. Our findings provide both theoretical grounding and methodological tools for developing language-neutral LRMs.

0 citationsRead paper
Recent publications

Latest Papers

$α$Transfer: Coefficient Transfer for Efficient Model Merging

Oct 06, 2026

This study addresses the prohibitive computational and memory overhead incurred by coefficient search as model merging scales up. To this end, we propose $\alpha$Transfer, a paradigm grounded in the assumption that coefficient distributions remain consistent within a model family. This method decouples the coefficient search from large models to small proxy models, achieving efficient merging through parameter arithmetic, performance distribution analysis, and cross-model transfer techniques. Experimental results demonstrate that $\alpha$Transfer yields a 6× speedup with 70% memory reduction on Vision Transformers, and a 20× speedup with 85% memory savings on large language models, all while maintaining performance comparable to the original methods.

0 citationsRead paper

A Broader Look at Model Merging: Rethinking Implicit Regularization Induced by Task Arithmetic

Oct 06, 2026

This study addresses the limitation in model merging where implicit regularization introduced during coefficient search constrains weights to a restricted subspace, thereby hindering multi-task performance improvements. To overcome this, we re-examine the implicit regularization mechanism in task arithmetic and propose directly searching the pre-trained weight space via unconstrained optimization, breaking through the subspace constraints inherent in traditional linear combinations. Our work reveals and eliminates these implicit regularization effects, demonstrating that superior solutions reside outside conventional subspaces. Across diverse architectures and extremely data-scarce scenarios, the proposed method significantly outperforms existing model merging techniques, establishing a new optimization paradigm for multi-task learning.

0 citationsRead paper

PRISM: A Geometric Risk Bound that Decomposes Drift into Scale, Shape, and Head

May 12, 2026

Existing methods struggle to diagnose representation drift in post-training variants of large language models—such as quantization or LoRA fine-tuning—as they can only assess performance degradation without identifying root causes or guiding mitigation. This work proposes PRISM, a method that leverages the linear output head and near-isometric backbone structure of LLMs to derive a closed-form upper bound on cross-entropy risk discrepancy. PRISM uniquely decomposes representation drift into three geometrically interpretable and independently measurable axes: scale, shape, and head. Notably, the shape component is differentiable and can be employed as a regularizer to mitigate catastrophic forgetting. Experiments show that PRISM achieves Spearman correlation coefficients of 0.820 and 0.831 in risk ranking for quantized and LoRA-adapted models, respectively, across two model families and five benchmarks; moreover, shape-based regularization outperforms experience replay in alleviating downstream forgetting.

0 citationsRead paper

On Calibration of Large Language Models: From Response To Capability

Feb 14, 2026

This work addresses a critical limitation in existing calibration methods for large language models (LLMs), which focus on the correctness of individual responses and often fail to reflect the model’s overall task-solving capability, leading to a misalignment between confidence and actual performance. To bridge this gap, we propose a novel paradigm—capability calibration—that estimates the expected accuracy of an LLM on a given query, thereby shifting the focus from response-level to task-level reliability assessment. We formally distinguish capability calibration from traditional response calibration, develop a theoretical framework grounded in the stochasticity of LLM decoding, and systematically evaluate various confidence estimation methods under this new paradigm. Experiments demonstrate that capability calibration substantially improves the accuracy of pass@$k$ prediction and enhances the efficiency of reasoning resource allocation, offering a more reliable foundation for downstream applications.

0 citationsRead paper

Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models?

May 23, 2025

This study identifies an implicit language bias in large reasoning models (LRMs): when processing multilingual inputs, LRMs default to high-resource languages—particularly English—for internal reasoning, severely degrading performance on low-resource language tasks. To systematically investigate this, we design a multidimensional, controllable evaluation framework covering MMMLU, MATH-500, CulturalBench, and LMSYS-toxic, integrating reasoning-path tracing and cross-lingual attribution analysis. We empirically demonstrate that enforcing same-language reasoning—though it reduces general reasoning capability (especially for low-resource languages)—significantly improves cultural alignment and language-specific accuracy in safety evaluation. Crucially, this work is the first to empirically establish “reasoning-language–input-language mismatch” as a fundamental bottleneck to multilingual fairness. Our findings provide both theoretical grounding and methodological tools for developing language-neutral LRMs.

0 citationsRead paper