Institution profile

Zendesk

Industry researchnorthamerica · us
Official website
Research library6linked papers
Opportunities0open roles
Selected work

Representative Papers

Multilingual GSM-Symbolic: What determines capability transfer across languages?

Oct 02, 2026

This study addresses the unclear mechanisms of cross-lingual capability transfer and the incomparability of evaluation data by constructing a multilingual symbolic mathematics dataset and proposing a joint estimation framework. Methodologically, it employs symbolic template generation to ensure data diversity, prevent overfitting, and systematically quantify key factors influencing capability transfer. The findings reveal that model scale and resource availability dominate cross-lingual transfer. Notably, the proposed framework explains 92% of the performance variance across languages and achieves a prediction error of merely six percentage points on unseen languages. Overall, this work provides a reliable theoretical foundation and an evaluation paradigm for understanding cross-lingual capability transfer in large language models.

0 citationsRead paper

Agent Reliability Profiles in Financial Services

Oct 02, 2026

This study addresses the challenge of scaling AI agents in financial services due to the absence of a unified evaluation framework by proposing a standardized reliability assessment system. Methodologically, it employs structured schemas to delineate agent operational boundaries and constructs reliability profiles encompassing autonomy and operational domains, alongside a three-tier hierarchical verification mechanism and an automated benchmarking toolkit. The core contribution lies in pioneering a falsifiable chain of evidence that enables cross-institutional benchmark comparisons, transitioning from mere assertions to independent verification. Ultimately, this research establishes an interoperable trust evaluation framework that effectively accelerates the large-scale compliant deployment and regulatory adoption of AI agents within financial institutions.

0 citationsRead paper

How Far Can Prompting Go for Minimal-Edit Ukrainian Grammatical Error Correction?

Jun 08, 2026

This study investigates whether prompt engineering alone can approach the performance of fine-tuned large language models on Ukrainian minimal-edit grammatical error correction (GEC). We systematically evaluate twelve large language models—including eleven commercial systems and one open-source Ukrainian-specific model—using zero-shot and few-shot prompting, minimal-edit constraints, and model-assisted prompt optimization, enhanced with linguistically informed instructions grounded in Ukrainian grammar. Our work provides the first comprehensive validation of prompt engineering’s efficacy for Ukrainian GEC, revealing its strong dependence on prompt language and identifying five distinct overcorrection patterns tied to Ukrainian linguistic characteristics. The best-performing configuration, Gemini 1.5 Pro, achieves an F0.5 score of 69.22 on the UNLP 2023 benchmark, closing over 90% of the performance gap with the current fine-tuned state-of-the-art model.

0 citationsRead paper
Recent publications

Latest Papers

Multilingual GSM-Symbolic: What determines capability transfer across languages?

Oct 02, 2026

This study addresses the unclear mechanisms of cross-lingual capability transfer and the incomparability of evaluation data by constructing a multilingual symbolic mathematics dataset and proposing a joint estimation framework. Methodologically, it employs symbolic template generation to ensure data diversity, prevent overfitting, and systematically quantify key factors influencing capability transfer. The findings reveal that model scale and resource availability dominate cross-lingual transfer. Notably, the proposed framework explains 92% of the performance variance across languages and achieves a prediction error of merely six percentage points on unseen languages. Overall, this work provides a reliable theoretical foundation and an evaluation paradigm for understanding cross-lingual capability transfer in large language models.

0 citationsRead paper

Agent Reliability Profiles in Financial Services

Oct 02, 2026

This study addresses the challenge of scaling AI agents in financial services due to the absence of a unified evaluation framework by proposing a standardized reliability assessment system. Methodologically, it employs structured schemas to delineate agent operational boundaries and constructs reliability profiles encompassing autonomy and operational domains, alongside a three-tier hierarchical verification mechanism and an automated benchmarking toolkit. The core contribution lies in pioneering a falsifiable chain of evidence that enables cross-institutional benchmark comparisons, transitioning from mere assertions to independent verification. Ultimately, this research establishes an interoperable trust evaluation framework that effectively accelerates the large-scale compliant deployment and regulatory adoption of AI agents within financial institutions.

0 citationsRead paper

How Far Can Prompting Go for Minimal-Edit Ukrainian Grammatical Error Correction?

Jun 08, 2026

This study investigates whether prompt engineering alone can approach the performance of fine-tuned large language models on Ukrainian minimal-edit grammatical error correction (GEC). We systematically evaluate twelve large language models—including eleven commercial systems and one open-source Ukrainian-specific model—using zero-shot and few-shot prompting, minimal-edit constraints, and model-assisted prompt optimization, enhanced with linguistically informed instructions grounded in Ukrainian grammar. Our work provides the first comprehensive validation of prompt engineering’s efficacy for Ukrainian GEC, revealing its strong dependence on prompt language and identifying five distinct overcorrection patterns tied to Ukrainian linguistic characteristics. The best-performing configuration, Gemini 1.5 Pro, achieves an F0.5 score of 69.22 on the UNLP 2023 benchmark, closing over 90% of the performance gap with the current fine-tuned state-of-the-art model.

0 citationsRead paper