Score
Designs and applies analyses, observational protocols, and measurement instruments to characterize how people detect, diagnose, and recover from faults in programs or interactive systems. Work includes building taxonomies of recovery strategies, coding debugging behaviors (e.g., intent clarification, code tracing), quantifying strategy preferences, and relating error types to chosen recovery actions.
A significant gap exists between academic research and industrial practice in debugging machine learning (ML) systems. Method: We propose the first comprehensive, lifecycle-spanning taxonomy of ML debugging faults and corresponding mitigation methods, derived from a systematic literature review (SLR), in-depth interviews with 28 ML practitioners, and empirical analysis of 1,247 GitHub issues. Contribution/Results: Our study identifies 13 core debugging challenges; only 48% are addressed by existing academic work, while 52.6% of GitHub issues and 70.3% of interview-elicited problems lack corresponding methodological support. Critically, we quantitatively demonstrate that over half of real-world ML debugging difficulties remain unaddressed by current research—revealing a substantial knowledge gap. This work establishes a foundational classification framework, provides empirical evidence of methodological coverage gaps, and delivers a prioritized roadmap to bridge the theory-practice divide in ML debugging.
Existing debugging tools excel at verifying hypotheses but struggle to support hypothesis generation, as programmers must manually reconstruct the program’s state evolution. This work proposes a novel debugging paradigm centered on complete execution traces, leveraging program tracing techniques to record and temporally visualize the actual code paths executed, rather than relying on the static structure of the source code. By presenting runtime behavior in a chronological and contextualized manner, this approach significantly enhances the comprehensibility of program execution, thereby facilitating more efficient hypothesis generation during debugging. We implement a prototype system and conduct preliminary experiments that demonstrate its effectiveness in improving program understanding efficiency, while also uncovering key challenges and promising directions for future research.
This work addresses the limited interpretability of large language model–driven automated program repair, which hinders diagnosis and reproducibility of failure cases. To this end, we present TraceView, the first interactive visualization tool that structures repair trajectories into Thought-Action-Result triplets and supports semantic relationship annotation. By integrating trajectory parsing, relational modeling, and graph-based visualization techniques, TraceView enables traceable analysis from high-level overviews to fine-grained details. A user study demonstrates that TraceView significantly enhances developers’ comprehension of the repair process and improves navigation efficiency. The implementation and a demonstration video are publicly available.
This work addresses the inefficiency of general-purpose language agents in self-repair, which often stems from a lack of fine-grained failure diagnosis, leading to blind context expansion and conflation of distinct error types. To overcome this, the authors propose DARC, a novel framework that prioritizes diagnosis before repair: it first analyzes failure patterns across a task family using a development set, selects appropriate repair interventions, and employs a validator to freeze the optimal success-cost strategy, thereby enforcing a causal “diagnose-then-repair” workflow. By designing recovery-oriented interfaces that integrate failure mode analysis, pruning of a shared repair library, and strategy freezing, DARC significantly improves task success rates while reducing interaction steps or retrieval overhead across diverse environments—including ALFWorld, AppWorld, and XBRL Finance—outperforming both standard foundation agents and existing general-purpose repair methods.
This study addresses the high cognitive load and attentional depletion developers experience during debugging, which threaten project sustainability. Grounded in cognitive science, it conceptualizes debugging as a cognitive process of constructing mental models that explain anomalous system behavior. Through semi-structured interviews with 27 professional developers and employing a constructivist grounded theory approach, the research systematically uncovers, for the first time, the neural and attentional mechanisms underlying the experience of confusion during debugging. The resulting empirically grounded cognitive theory of debugging elucidates the intrinsic relationships among debugging difficulty, cognitive resource consumption, and development sustainability. This work provides a theoretical foundation for developer experience research and offers practical guidance for industry efforts to optimize debugging tools and workflows.