ORCA: Evaluating LLMs on Data Science Code Translation

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient evaluation of large language models (LLMs) in data science code translation by constructing a comprehensive benchmark encompassing multi-domain tasks and complete projects to systematically assess their cross-library translation capabilities. Furthermore, this work proposes an innovative intent-augmented approach that guides the translation process by inferring the deep semantic intent of source code. Experiments based on functional equivalence verification demonstrate that the success rates of current state-of-the-art models remain limited. Nevertheless, the proposed intent augmentation method yields an average absolute improvement of approximately 5% in success rate and reveals directional preferences inherent in these models, thereby providing an effective paradigm and novel insights for optimizing code translation.
📝 Abstract
Data Science Code Translation (DSCT) is the process of converting code between data science libraries while preserving functional equivalence and enabling interoperability across data science ecosystems. While Large Language Models (LLMs) have demonstrated considerable progress in Data Science Code Generation (DSCG), their performance in DSCT remains insufficiently studied. To address this gap, we introduce ORCA, a comprehensive benchmark with two complementary settings: ORCA-MAIN, which comprises 1,600 carefully curated grounding-level tasks across 3 representative domains: Data Querying, Data Manipulation, and Deep Learning; and ORCA-PROJECT, which contains 200 translation tasks over complete data science projects across 7 data science task types. Each task is accompanied by annotated reference translations and test cases for validating functional equivalence. We further incorporate a multi-stage quality verification process that thoroughly verifies task correctness and test case robustness. Experimental results demonstrate challenges in DSCT, with even frontier LLMs showing limited performance. Specifically, Claude-Opus-4.6 achieves a success rate of 56.92% on ORCA-MAIN and 33.67% on ORCA-PROJECT, indicating considerable room for improvement in DSCT. We also observe a clear directional preference in DSCT, where translation is consistently easier when the source code expresses the task through more explicit, fine-grained operations. Motivated by this, we propose an intent-augmented method, in which the model first infers source-code intent and then uses it as additional context for translation, achieving average absolute success-rate gains of 4.80% and 5.33% on ORCA-MAIN and ORCA-PROJECT, respectively.
Problem

Research questions and friction points this paper is trying to address.

Data Science Code Translation
Large Language Models
Benchmark Evaluation
Functional Equivalence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Data Science Code Translation
Benchmark
Intent-Augmented Method
Large Language Models
Functional Equivalence
🔎 Similar Papers
2024-03-252024 IEEE/ACM First International Conference on AI Foundation Models and Software Engineering (Forge) Conference Acronym:Citations: 22