Institution profile

Trip.com Group

Industry researchasia · cn
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Bridging the Last Mile of Time Series Forecasting with LLM Agents

Jun 01, 2026

This work addresses the “last-mile” gap between statistical time series forecasting and business decision-making—specifically, the challenge of effectively incorporating weakly structured contextual factors such as holidays and marketing campaigns. To bridge this gap, we propose the first framework that systematically integrates large language model (LLM) agents into the post-forecasting refinement stage. Our approach unifies the forecasting workspace, leverages tool-augmented retrieval of external evidence, and translates LLM reasoning into explicit revision operations under structured safety constraints. It further supports Map-Reduce-style long-horizon divide-and-conquer forecasting and incorporates a memory-based reflection mechanism. Evaluated in real-world business settings, the method significantly enhances the operational relevance and decision utility of forecasts, delivering controllable, auditable, and business-ready predictions.

0 citationsRead paper

AdaTIR: Adaptive Tool-Integrated Reasoning via Difficulty-Aware Policy Optimization

Jan 21, 2026

This work addresses the inefficiency of large language models in tool-integrated reasoning, where redundant external tool invocations often occur due to an inability to assess task difficulty. To mitigate this, we propose AdaTIR, a framework that employs a reinforcement learning–driven, difficulty-aware strategy to dynamically allocate tool-call budgets: performing internal reasoning for simple tasks and selectively invoking tools for complex ones. AdaTIR introduces a novel difficulty-aware efficiency reward and Clipped Advantage Shaping (CAS) to alleviate the suppression of correctness rewards by tool penalties, thereby achieving an adaptive balance between internal reasoning and tool usage. Experiments show that AdaTIR reduces tool calls by 97.6% on simple tasks and 28.2% on complex tasks while maintaining or improving accuracy; notably, it outperforms baselines by 4.8% on the AIME 2024 no-tool evaluation.

0 citationsRead paper

PASS-FC: Progressive and Adaptive Search Scheme for Fact Checking of Comprehensive Claims

Apr 14, 2025

To address low accuracy and poor cross-domain generalization in automated fact-checking of complex real-world claims, this paper proposes the Progressive Adaptive Search Framework (PASS). PASS introduces three core innovations: (1) temporal- and entity-aware contextual enhancement for claim rewriting; (2) dynamic question generation triggered by reflective labeling; and (3) multi-granularity evidence aggregation with iterative reflective verification. Integrated with retrieval-augmented generation (RAG), language adaptation, and multilingual support modules, PASS enables end-to-end adaptive verification. Evaluated across six heterogeneous datasets—including general knowledge, scientific claims, real-world scenarios, and multilingual tasks—PASS consistently outperforms state-of-the-art baselines, achieving significant gains in accuracy. All code and experimental artifacts are publicly released.

0 citationsRead paper
Recent publications

Latest Papers

Bridging the Last Mile of Time Series Forecasting with LLM Agents

Jun 01, 2026

This work addresses the “last-mile” gap between statistical time series forecasting and business decision-making—specifically, the challenge of effectively incorporating weakly structured contextual factors such as holidays and marketing campaigns. To bridge this gap, we propose the first framework that systematically integrates large language model (LLM) agents into the post-forecasting refinement stage. Our approach unifies the forecasting workspace, leverages tool-augmented retrieval of external evidence, and translates LLM reasoning into explicit revision operations under structured safety constraints. It further supports Map-Reduce-style long-horizon divide-and-conquer forecasting and incorporates a memory-based reflection mechanism. Evaluated in real-world business settings, the method significantly enhances the operational relevance and decision utility of forecasts, delivering controllable, auditable, and business-ready predictions.

0 citationsRead paper

AdaTIR: Adaptive Tool-Integrated Reasoning via Difficulty-Aware Policy Optimization

Jan 21, 2026

This work addresses the inefficiency of large language models in tool-integrated reasoning, where redundant external tool invocations often occur due to an inability to assess task difficulty. To mitigate this, we propose AdaTIR, a framework that employs a reinforcement learning–driven, difficulty-aware strategy to dynamically allocate tool-call budgets: performing internal reasoning for simple tasks and selectively invoking tools for complex ones. AdaTIR introduces a novel difficulty-aware efficiency reward and Clipped Advantage Shaping (CAS) to alleviate the suppression of correctness rewards by tool penalties, thereby achieving an adaptive balance between internal reasoning and tool usage. Experiments show that AdaTIR reduces tool calls by 97.6% on simple tasks and 28.2% on complex tasks while maintaining or improving accuracy; notably, it outperforms baselines by 4.8% on the AIME 2024 no-tool evaluation.

0 citationsRead paper

PASS-FC: Progressive and Adaptive Search Scheme for Fact Checking of Comprehensive Claims

Apr 14, 2025

To address low accuracy and poor cross-domain generalization in automated fact-checking of complex real-world claims, this paper proposes the Progressive Adaptive Search Framework (PASS). PASS introduces three core innovations: (1) temporal- and entity-aware contextual enhancement for claim rewriting; (2) dynamic question generation triggered by reflective labeling; and (3) multi-granularity evidence aggregation with iterative reflective verification. Integrated with retrieval-augmented generation (RAG), language adaptation, and multilingual support modules, PASS enables end-to-end adaptive verification. Evaluated across six heterogeneous datasets—including general knowledge, scientific claims, real-world scenarios, and multilingual tasks—PASS consistently outperforms state-of-the-art baselines, achieving significant gains in accuracy. All code and experimental artifacts are publicly released.

0 citationsRead paper