Score
Designs and implements abstraction layers and APIs that expose a common problem-solving interface and adapter implementations so the same problem code can be mapped to and executed by multiple concrete solver backends. Builds solver integration features such as unified/common solver APIs, dependency injection and lifecycle management, backend adapters that isolate solver-specific glue code, and validation and testing hooks to ensure swapping solvers preserves correctness.
This study addresses the lack of systematic empirical investigation into the technical and integration challenges of deploying the Matter standard in practice. By analyzing over 13,000 developer issues from the official Project CHIP GitHub repository, this work employs topic modeling and qualitative content analysis to uncover key obstacles encountered during Matter implementation at scale. The research identifies four core challenge categories: testing, interoperability, development support, and platform networking. These findings provide the first community-driven empirical evidence and actionable recommendations for improving Matter’s test infrastructure, cross-vendor documentation, and developer toolchains.
研究通过对比分析三种SBOM生成工具在JavaScript和Rust项目中的表现,揭示了因SBOM规范模糊导致的系统性差异问题,并建议未来需明确标准化规则以提高互操作性和合规性。
Verifying the equivalence of implementations of the same large model across different frameworks is highly challenging due to significant discrepancies in operator decomposition, tensor layouts, and fusion strategies. This work proposes Emerge, a framework that unifies two implementations into a single e-graph representation, infers candidate equivalences guided by runtime values, and automatically synthesizes rewrite rules on demand without manual intervention. By integrating symbolic SMT-based verification with constraint-aware randomized testing, Emerge supports scenarios involving opaque operators. Experimental results demonstrate that Emerge successfully verifies equivalence for correct implementation pairs, detects 10 out of 13 known bugs, and uncovers 8 previously unknown issues confirmed by developers. The automatically generated block-level rewrite rules achieve effectiveness comparable to handcrafted ones.
To address the dual challenge of multi-task generalization and computational efficiency in Automated Program Repair (APR) with Code Large Language Models (Code LLMs), this paper proposes a novel adapter-oriented continual fusion mechanism. Unlike conventional uniform or static-weighted adapter merging, our approach introduces task ordering and dynamic weighting into adapter composition—a first in Code LLM adaptation. Leveraging parameter-efficient fine-tuning, we empirically validate the method on CodeLlama: under optimal task sequences, continual fusion improves repair success rates by up to 12.7% over single-task adapters, significantly enhancing cross-task generalization while reducing inference overhead. Our core contribution lies in establishing the critical roles of task sequence and dynamic weight assignment in adapter fusion—thereby overcoming fundamental limitations of existing fusion paradigms.
This work addresses the limited cross-file contextual awareness of code large language models (CodeLLMs) in repository-level code generation. Methodologically, we introduce RepoExec—the first executable and functionally correct repository-level benchmark—and propose Dependency Invocation Rate (DIR), a novel metric quantifying the accuracy of cross-file dependency invocation. We further design an instruction-tuning dataset integrating test-driven validation and context-aware dependency modeling. Our contributions include the first comprehensive evaluation framework encompassing context-awareness, execution-driven assessment, and cross-file dependency modeling. Experimental results demonstrate that instruction tuning significantly improves contextual utilization and debugging capability, whereas pre-trained models exhibit stronger functional correctness. RepoExec has since become the de facto standard benchmark for repository-level code generation research.
This study addresses the challenge of characterizing the propagation and evolutionary dynamics of interface variants within package dependency graphs in distributed software ecosystems—a limitation of traditional interface compatibility research. The authors introduce, for the first time, the concept of selection coefficients from population genetics into software ecosystem analysis. By integrating package graph mining, 2,100 parser probe experiments, and conflict probability modeling, they develop a non-neutral evolutionary model that delineates parser rules as diagnostic signals versus independent selective pressures. Leveraging absorbing Markov processes, directional permutation null models, and time-series predictive validation (via Brier score and AUC), the work presents the first reproducible framework for dynamically assessing interface variants. Empirical results demonstrate that parser-derived directionality significantly influences variant adoption (MAE: 0.07 vs. 0.43, p=0.002), whereas purely temporal features fail to surpass baseline performance.
This study addresses the engineering challenges encountered when large language models construct complete code repositories, including module coupling, interface conflicts, and complex dependency interactions. To overcome these limitations, this work proposes LEGO, a framework that introduces an agent-native reusable code primitive system. It employs persistent LLMs for relevance assessment and adaptive modification, combined with dependency closure resolution and cross-component constraint solving to activate, integrate, and debug primitives. Furthermore, the authors present LEGO-REPO, an end-to-end refactoring benchmark. Experimental results demonstrate an average performance improvement of 0.1474 across various models, with GPT-5.6-terra achieving a 61.4% gain. Notably, the framework maintains 95.1% homogenized evaluation scores at reduced computational cost, significantly enhancing the controllability of automated large-scale repository construction.
Existing large code models struggle to efficiently capture repository-level context—such as imports, APIs, and project conventions—and current approaches are limited either by inference overhead or adaptability to code evolution. This work proposes Code2LoRA, a novel framework that leverages a hypernetwork to dynamically generate lightweight LoRA adapters for each code repository, injecting repository-specific knowledge at zero inference token cost while supporting both static snapshots and continuous evolution scenarios. Specifically, Code2LoRA-Evo incorporates GRU hidden states to track code changes over time. Evaluation on our newly introduced RepoPeftBench benchmark shows that Code2LoRA-Static achieves 66.2% and 63.8% exact match accuracy on in-repo and cross-repo tasks, respectively, matching per-repo fine-tuning performance; meanwhile, Code2LoRA-Evo attains 60.3% cross-repo accuracy on evolution tasks, substantially outperforming shared LoRA baselines.
Existing bug-fixing frameworks struggle to address the unique challenges posed by AI/ML systems, including non-deterministic behavior, experiment-driven workflows, and the need for coordinated changes across multiple artifacts. Through a qualitative analysis of 100 issue reports and pull requests from TensorFlow, scikit-learn, MLflow, and AutoGPT, this study systematically uncovers core characteristics of AI/ML debugging and repair—namely cross-phase activities, iterative validation, and multi-artifact coordination. The research identifies key obstacles such as reproducibility issues, behavioral non-determinism, and artifact misalignment. Building on these findings, the paper articulates a vision for a tailored bug-fixing framework specifically designed for AI/ML systems, offering an empirical foundation to guide future toolchain development and research in this emerging domain.
AI-assisted development tools enable rapid prototyping of services but often lack awareness of architectural constraints, infrastructure dependencies, and organizational standards required in production environments. Consequently, generated artifacts may exhibit brittle behavior and limited deployability. We propose a retrieval-augmented scaffolding approach that combines platform-based code generation with agentic clarification loops to expose and resolve architectural constraint ambiguities. By combining template retrieval with structured interaction, the method embeds production-relevant considerations during service scaffolding. Evaluation indicates improved architectural consistency and deployability compared to general-purpose AI code generation workflows, suggesting that constraint-aware retrieval is essential for aligning AI-assisted service development with production software engineering practices.