Institution profile

RayNeo

Industry research
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Probability is Not Enough: Exploring and Counting Divergent Tokens for Reasoning Uncertainty Quantification in LLMs

Sep 29, 2026

This study addresses the unclear mechanisms underlying confidence estimation in large language model (LLM) reasoning and the limited effectiveness of existing probability calibration methods. To this end, it proposes Divergent Token Counting (DTC), a training-free framework that leverages Jensen-Shannon divergence to identify and count strongly divergent tokens between two models along identical decoding trajectories. The analysis reveals a significant negative correlation between divergence counts and predictive accuracy, enabling efficient uncertainty quantification in both black-box and white-box settings. Evaluated on mathematical reasoning benchmarks, DTC reduces calibration error to approximately 13%, substantially outperforming conventional baselines. Overall, this work presents a lightweight yet effective solution for assessing LLM reliability.

0 citationsRead paper

GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning

Aug 08, 2026

This work addresses the challenges of physical hallucination, poor generalization, and inadequate environmental perception in large language model (LLM)-driven embodied agents when performing long-horizon tasks. To overcome these limitations, the authors propose a synergistic planning framework that integrates task graphs and scene graphs. The task graph encodes structured prior knowledge to guide high-level planning, while the scene graph enables event-driven dynamic replanning, establishing a closed-loop perception and error-correction mechanism. This approach is the first to incorporate dual-graph structures into LLM-based planning pipelines, combining contextual prompting, GRPO-based reward design, and model fine-tuning. Evaluated on the ALFRED benchmark, the method achieves state-of-the-art performance, significantly improving task success rates under both zero-shot and few-shot settings and demonstrating strong out-of-distribution generalization on unseen long-horizon tasks.

0 citationsRead paper
Recent publications

Latest Papers

Probability is Not Enough: Exploring and Counting Divergent Tokens for Reasoning Uncertainty Quantification in LLMs

Sep 29, 2026

This study addresses the unclear mechanisms underlying confidence estimation in large language model (LLM) reasoning and the limited effectiveness of existing probability calibration methods. To this end, it proposes Divergent Token Counting (DTC), a training-free framework that leverages Jensen-Shannon divergence to identify and count strongly divergent tokens between two models along identical decoding trajectories. The analysis reveals a significant negative correlation between divergence counts and predictive accuracy, enabling efficient uncertainty quantification in both black-box and white-box settings. Evaluated on mathematical reasoning benchmarks, DTC reduces calibration error to approximately 13%, substantially outperforming conventional baselines. Overall, this work presents a lightweight yet effective solution for assessing LLM reliability.

0 citationsRead paper

GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning

Aug 08, 2026

This work addresses the challenges of physical hallucination, poor generalization, and inadequate environmental perception in large language model (LLM)-driven embodied agents when performing long-horizon tasks. To overcome these limitations, the authors propose a synergistic planning framework that integrates task graphs and scene graphs. The task graph encodes structured prior knowledge to guide high-level planning, while the scene graph enables event-driven dynamic replanning, establishing a closed-loop perception and error-correction mechanism. This approach is the first to incorporate dual-graph structures into LLM-based planning pipelines, combining contextual prompting, GRPO-based reward design, and model fine-tuning. Evaluated on the ALFRED benchmark, the method achieves state-of-the-art performance, significantly improving task success rates under both zero-shot and few-shot settings and demonstrating strong out-of-distribution generalization on unseen long-horizon tasks.

0 citationsRead paper