Institution profile

California State University

Academic institutionnorthamerica · us
Official website
Research library17linked papers
Opportunities0open roles
Selected work

Representative Papers

CARET: Training-Free Test-Time Scaling for Repository-Level Code Completion

Oct 04, 2026

This work addresses the disconnect between retrieval and generation stages in repository-level code completion, where existing methods often overlook generative potential. We propose CARET, a training-free framework that pioneers a synergistic optimization paradigm integrating retrieval and generation. Specifically, CARET introduces a sample-consistency-based cascaded retrieval routing mechanism to dynamically filter contexts, leverages KV cache reuse to reduce multi-sampling overhead, and employs reverse context likelihood for self-scoring candidate code. Evaluated across six models and multiple benchmarks, CARET improves average exact match rates by 10.8 and 5.3 percentage points over greedy decoding and self-consistency, respectively, while incurring a computational cost of only approximately 1.45× that of single-pass generation.

0 citationsRead paper

Merged, Not Measured: An Empirical Study of Performance Issues Fixed by Coding Agents

Sep 29, 2026

This study investigates the practical effectiveness of coding agents in resolving performance issues and the mechanisms governing maintainer acceptance. By analyzing over 70,000 agent-generated pull requests and identifying 1,262 performance-related fixes, the research employs a combination of text filtering, LLM-assisted classification, and empirical reproduction testing. Results indicate that while 57% of these fixes are merged, maintainer decisions rely predominantly on repository history rather than code quality. Furthermore, verification reveals that most merged patches fail to achieve the anticipated performance improvements. The analysis identifies redundant computation as the primary source of performance defects and demonstrates that such fixes frequently introduce substantial behavioral change risks. These findings provide critical empirical evidence for guiding agent-based performance repair practices in software maintenance.

0 citationsRead paper

NAQD Env: A benchmark for selective withdrawal in language agents

Sep 29, 2026

This study investigates the selective retraction capability of language agents when evidence changes or instructions are revoked, specifically their ability to precisely suspend affected actions while preserving valid work. We construct a synthetic environment, NAQD-Env, alongside a selective retraction benchmark grounded in deterministic reference policies to jointly evaluate policy consistency, task value, and recovery performance. Experiments reveal that existing models exhibit extremely low retraction recall and lack recovery capabilities. While exploratory fine-tuning significantly improves decision accuracy, it induces over-retraction and the loss of event reports. This work exposes critical deficiencies in the dynamic adaptability of current agents and offers new directions for designing trustworthy AI systems.

0 citationsRead paper
Recent publications

Latest Papers

CARET: Training-Free Test-Time Scaling for Repository-Level Code Completion

Oct 04, 2026

This work addresses the disconnect between retrieval and generation stages in repository-level code completion, where existing methods often overlook generative potential. We propose CARET, a training-free framework that pioneers a synergistic optimization paradigm integrating retrieval and generation. Specifically, CARET introduces a sample-consistency-based cascaded retrieval routing mechanism to dynamically filter contexts, leverages KV cache reuse to reduce multi-sampling overhead, and employs reverse context likelihood for self-scoring candidate code. Evaluated across six models and multiple benchmarks, CARET improves average exact match rates by 10.8 and 5.3 percentage points over greedy decoding and self-consistency, respectively, while incurring a computational cost of only approximately 1.45× that of single-pass generation.

0 citationsRead paper

Merged, Not Measured: An Empirical Study of Performance Issues Fixed by Coding Agents

Sep 29, 2026

This study investigates the practical effectiveness of coding agents in resolving performance issues and the mechanisms governing maintainer acceptance. By analyzing over 70,000 agent-generated pull requests and identifying 1,262 performance-related fixes, the research employs a combination of text filtering, LLM-assisted classification, and empirical reproduction testing. Results indicate that while 57% of these fixes are merged, maintainer decisions rely predominantly on repository history rather than code quality. Furthermore, verification reveals that most merged patches fail to achieve the anticipated performance improvements. The analysis identifies redundant computation as the primary source of performance defects and demonstrates that such fixes frequently introduce substantial behavioral change risks. These findings provide critical empirical evidence for guiding agent-based performance repair practices in software maintenance.

0 citationsRead paper

NAQD Env: A benchmark for selective withdrawal in language agents

Sep 29, 2026

This study investigates the selective retraction capability of language agents when evidence changes or instructions are revoked, specifically their ability to precisely suspend affected actions while preserving valid work. We construct a synthetic environment, NAQD-Env, alongside a selective retraction benchmark grounded in deterministic reference policies to jointly evaluate policy consistency, task value, and recovery performance. Experiments reveal that existing models exhibit extremely low retraction recall and lack recovery capabilities. While exploratory fine-tuning significantly improves decision accuracy, it induces over-retraction and the loss of event reports. This work exposes critical deficiencies in the dynamic adaptability of current agents and offers new directions for designing trustworthy AI systems.

0 citationsRead paper