Institution profile

PayPal

Industry researchnorthamerica · us
Official website
Research library43linked papers
Opportunities0open roles
Selected work

Representative Papers

Structure Tax: How Structured Output affects LLMs Performance

Oct 08, 2026

This study addresses the "structure tax" problem, wherein enforcing structured outputs in large language models degrades accuracy, by conducting a systematic evaluation across multiple models and datasets. Methodologically, it employs Centered Kernel Alignment (CKA) to analyze representation separability in intermediate Transformer layers, alongside confidence calibration and hidden-layer geometric measurements. The findings reveal that accuracy degradation stems from schema design rather than structural constraints per se, proposing a paradigm shift from "whether to structure" to "how to structure." Experiments demonstrate that a reasoning-prioritized field ordering strategy matches or surpasses free-text performance while significantly improving model calibration.

0 citationsRead paper

All Verdicts are Not Equal: Rethinking LLM Judge Reliability

Oct 08, 2026

This study addresses systematic deficiencies in the LLM-as-a-Judge paradigm, including positional bias, irreproducibility, and misalignment with human judgment. Through multi-benchmark stress testing and controlled experiments, we comprehensively audit the reliability of judgments produced by frontier models. We propose a "trustworthy judgment rate" metric to quantify evaluation stability and derive the theoretical upper bound on accuracy imposed by positional bias. Our findings reveal that judgments remain unstable even at zero temperature, and that reliability is sample-specific rather than an intrinsic model-level property. Furthermore, transitioning from pairwise comparison to holistic scoring yields substantially greater improvements in trustworthiness than single-prompt interventions. This work provides both a theoretical foundation and a practical framework for constructing robust NLP evaluation systems.

0 citationsRead paper

CM-DPO: Constraint-Margin Direct Preference Optimization for LLM Planning

Oct 06, 2026

This study addresses the limitations of Direct Preference Optimization (DPO) in large language model (LLM) planning, specifically its neglect of constraint violation severity and inherent biases in preference data. To overcome these issues, we propose CM-DPO, a method that leverages a symbolic verifier to generate continuous constraint margin signals and employs a lexicographic objective function to strictly decouple hard and soft constraints. Furthermore, we introduce the SynPlan-R framework, which integrates programmatic generation with minimal-edit distillation to construct unbiased training data. Experimental results demonstrate that an 8B-parameter model achieves an 89.2% pass rate while reducing inference latency by 13 times. Notably, our approach outperforms GPT-4o by 9.2 percentage points on the Blocksworld benchmark, highlighting its effectiveness for efficient and reliable LLM-based planning.

0 citationsRead paper

Student-Guided Teacher Distillation for Efficient LLM Task Routing: Positioning Against Jev-Style System-1 Classifiers

Oct 01, 2026

This study addresses the computational bottleneck in zero-shot LLM task routing, where scoring all labels causes inference costs to scale linearly with taxonomy size. To mitigate this, we propose a student-guided teacher distillation framework in which a compact model predicts label distributions and retrieves top-k candidates, while a large language model reranks only these candidates to iteratively refine the student. Furthermore, we design an end-to-end objective for taxonomy-aware candidate generation, revealing that directly applying truncated scores for KL divergence distillation corrupts dark knowledge and introduces systematic bias. Built upon architectures such as ModernBERT, the optimal student model achieves 77.5% agreement with the teacher and attains a Coverage@16 of 91–100%, substantially reducing inference overhead while preserving routing effectiveness.

0 citationsRead paper

PDE-OBS: Controlled Evaluation Across Observation Patterns

Sep 28, 2026

This study addresses the limitation that single observation mode evaluations fail to capture performance fluctuations in physical field reconstruction by proposing an integrated benchmark platform. The platform incorporates seven types of partial differential equation data with configurable observation operators, enabling parameterized mode definitions through the decoupling of observation construction from physical records, and establishes a standardized cross-mode evaluation protocol to quantify model sensitivity. Experiments reveal that cross-mode errors consistently exceed matched-mode errors, and dense observations do not necessarily reduce errors. Furthermore, mixed training strategies effectively mitigate transfer errors. This work provides a systematic tool and novel insights for the robustness evaluation of physical field reconstruction.

0 citationsRead paper
Recent publications

Latest Papers

Structure Tax: How Structured Output affects LLMs Performance

Oct 08, 2026

This study addresses the "structure tax" problem, wherein enforcing structured outputs in large language models degrades accuracy, by conducting a systematic evaluation across multiple models and datasets. Methodologically, it employs Centered Kernel Alignment (CKA) to analyze representation separability in intermediate Transformer layers, alongside confidence calibration and hidden-layer geometric measurements. The findings reveal that accuracy degradation stems from schema design rather than structural constraints per se, proposing a paradigm shift from "whether to structure" to "how to structure." Experiments demonstrate that a reasoning-prioritized field ordering strategy matches or surpasses free-text performance while significantly improving model calibration.

0 citationsRead paper

All Verdicts are Not Equal: Rethinking LLM Judge Reliability

Oct 08, 2026

This study addresses systematic deficiencies in the LLM-as-a-Judge paradigm, including positional bias, irreproducibility, and misalignment with human judgment. Through multi-benchmark stress testing and controlled experiments, we comprehensively audit the reliability of judgments produced by frontier models. We propose a "trustworthy judgment rate" metric to quantify evaluation stability and derive the theoretical upper bound on accuracy imposed by positional bias. Our findings reveal that judgments remain unstable even at zero temperature, and that reliability is sample-specific rather than an intrinsic model-level property. Furthermore, transitioning from pairwise comparison to holistic scoring yields substantially greater improvements in trustworthiness than single-prompt interventions. This work provides both a theoretical foundation and a practical framework for constructing robust NLP evaluation systems.

0 citationsRead paper

CM-DPO: Constraint-Margin Direct Preference Optimization for LLM Planning

Oct 06, 2026

This study addresses the limitations of Direct Preference Optimization (DPO) in large language model (LLM) planning, specifically its neglect of constraint violation severity and inherent biases in preference data. To overcome these issues, we propose CM-DPO, a method that leverages a symbolic verifier to generate continuous constraint margin signals and employs a lexicographic objective function to strictly decouple hard and soft constraints. Furthermore, we introduce the SynPlan-R framework, which integrates programmatic generation with minimal-edit distillation to construct unbiased training data. Experimental results demonstrate that an 8B-parameter model achieves an 89.2% pass rate while reducing inference latency by 13 times. Notably, our approach outperforms GPT-4o by 9.2 percentage points on the Blocksworld benchmark, highlighting its effectiveness for efficient and reliable LLM-based planning.

0 citationsRead paper

Student-Guided Teacher Distillation for Efficient LLM Task Routing: Positioning Against Jev-Style System-1 Classifiers

Oct 01, 2026

This study addresses the computational bottleneck in zero-shot LLM task routing, where scoring all labels causes inference costs to scale linearly with taxonomy size. To mitigate this, we propose a student-guided teacher distillation framework in which a compact model predicts label distributions and retrieves top-k candidates, while a large language model reranks only these candidates to iteratively refine the student. Furthermore, we design an end-to-end objective for taxonomy-aware candidate generation, revealing that directly applying truncated scores for KL divergence distillation corrupts dark knowledge and introduces systematic bias. Built upon architectures such as ModernBERT, the optimal student model achieves 77.5% agreement with the teacher and attains a Coverage@16 of 91–100%, substantially reducing inference overhead while preserving routing effectiveness.

0 citationsRead paper

PDE-OBS: Controlled Evaluation Across Observation Patterns

Sep 28, 2026

This study addresses the limitation that single observation mode evaluations fail to capture performance fluctuations in physical field reconstruction by proposing an integrated benchmark platform. The platform incorporates seven types of partial differential equation data with configurable observation operators, enabling parameterized mode definitions through the decoupling of observation construction from physical records, and establishes a standardized cross-mode evaluation protocol to quantify model sensitivity. Experiments reveal that cross-mode errors consistently exceed matched-mode errors, and dense observations do not necessarily reduce errors. Furthermore, mixed training strategies effectively mitigate transfer errors. This work provides a systematic tool and novel insights for the robustness evaluation of physical field reconstruction.

0 citationsRead paper