Institution profile

Optum

Industry researchnorthamerica · us
Official website
Research library18linked papers
Opportunities0open roles
Selected work

Representative Papers

Reasoning Quality Emerges Early: Data Curation for Reasoning Models

Jun 25, 2026

Existing approaches to filtering high-quality reasoning data rely heavily on strong reasoning models, resulting in high costs and limited effectiveness. This work proposes an efficient alternative that reliably identifies challenging samples by analyzing the loss of a pretrained model over the first 100 reasoning tokens. By further incorporating loss patterns and gradient similarity from a small number of perturbed checkpoints, the method enables precise selection of diverse, high-difficulty data without requiring complex inference procedures. Evaluated on Qwen2.5-7B and Llama3.1-8B, this approach achieves up to a 1.7% improvement in fine-tuning performance while reducing token consumption by 91%, substantially lowering the cost of data curation.

0 citationsRead paper

Beyond Visual Forensics: Auditing Multimodal Robustness for Synthetic Medical Image Detection

Jun 24, 2026

This work proposes a novel framework based on adaptive feature fusion and contrastive learning to address the limited generalization of existing methods in complex scenarios. By dynamically integrating multi-scale semantic information and incorporating cross-sample consistency constraints, the approach significantly enhances model robustness under distribution shifts. Extensive experiments demonstrate that the proposed method consistently outperforms state-of-the-art models across multiple benchmark datasets, with particularly notable gains in low-resource and long-tailed settings. Beyond offering a new perspective for improving model generalization, this study also releases the associated code and pre-trained models to facilitate future research.

0 citationsRead paper
Recent publications

Latest Papers

Reasoning Quality Emerges Early: Data Curation for Reasoning Models

Jun 25, 2026

Existing approaches to filtering high-quality reasoning data rely heavily on strong reasoning models, resulting in high costs and limited effectiveness. This work proposes an efficient alternative that reliably identifies challenging samples by analyzing the loss of a pretrained model over the first 100 reasoning tokens. By further incorporating loss patterns and gradient similarity from a small number of perturbed checkpoints, the method enables precise selection of diverse, high-difficulty data without requiring complex inference procedures. Evaluated on Qwen2.5-7B and Llama3.1-8B, this approach achieves up to a 1.7% improvement in fine-tuning performance while reducing token consumption by 91%, substantially lowering the cost of data curation.

0 citationsRead paper

Beyond Visual Forensics: Auditing Multimodal Robustness for Synthetic Medical Image Detection

Jun 24, 2026

This work proposes a novel framework based on adaptive feature fusion and contrastive learning to address the limited generalization of existing methods in complex scenarios. By dynamically integrating multi-scale semantic information and incorporating cross-sample consistency constraints, the approach significantly enhances model robustness under distribution shifts. Extensive experiments demonstrate that the proposed method consistently outperforms state-of-the-art models across multiple benchmark datasets, with particularly notable gains in low-resource and long-tailed settings. Beyond offering a new perspective for improving model generalization, this study also releases the associated code and pre-trained models to facilitate future research.

0 citationsRead paper