Score
Design and implement reproducible, open-source reference implementations of attribution methods for models, including clean APIs, integration hooks for large models, automated tests, and performance and correctness checks. Produce documentation, examples, and tooling that let researchers and practitioners run, compare, and extend attribution methods reliably.
This study addresses the lack of systematic understanding regarding the implementation and maintenance of the Model Context Protocol (MCP) in real-world open-source projects. To bridge this gap, we introduce a transparent, reproducible multi-stage validation pipeline that integrates GitHub REST/GraphQL APIs with custom Python scripts to systematically annotate structural evidence, classify repository roles, and filter out non-functional examples from 3,238 candidate repositories. This process yields a high-quality dataset of 2,297 verified MCP projects, achieving a validation precision of 83% at 95% confidence. Our analysis reveals Python and TypeScript as the dominant implementation languages and identifies hybrid architecture as the most prevalent design pattern, thereby establishing the first large-scale empirical benchmark for MCP ecosystem research.
This study addresses the limitation of existing research in effectively measuring whether open-source models are genuinely translated into publicly accessible applications, as metrics based solely on release, visibility, or technical reuse inadequately capture real-world impact. To bridge this gap, the paper introduces “public application translation” as a distinct dimension of model influence and constructs a large-scale dataset leveraging Model-Space links from the Hugging Face platform. Through systematic analysis combining metadata readiness assessment with heterogeneous space configuration, the work reveals that only a small fraction of models are linked to Spaces—and these are highly concentrated. Models successfully translated into public applications exhibit higher metadata quality and are deeply embedded within a diverse ecosystem encompassing datasets, SDKs, and task categories, thereby extending the evaluative framework for open-source model impact.
Reproducibility remains a critical challenge in large language model (LLM)-driven software engineering (SE) research, undermining credibility and cumulative scientific progress. Method: We conducted a systematic literature review of 640 papers, integrating structured metadata extraction, manual annotation, and cross-platform analysis to diagnose reproducibility deficiencies across code, data, execution environments, and version control. Contribution/Results: We propose a novel taxonomy of seven reproducibility defect categories and introduce the Reproducibility Maturity Model (RMM), shifting evaluation from binary “reproducible/not reproducible” to a multi-dimensional, incremental framework. Our findings reveal that even top-tier conferences’ artifact evaluation badges exhibit low enforcement fidelity and poor long-term reproducibility; publication venue transparency practices vary substantially. This work provides both a theoretical framework and empirical evidence to enhance the rigor and trustworthiness of LLM-SE research.
This paper addresses the systemic absence of responsible practices in foundational model development by introducing the first comprehensive, multimodal resource guide covering text, vision, and speech modalities. Through systematic literature review, cross-modal taxonomy construction, and tool-to-capability mapping, it identifies four critical structural gaps: (1) scarcity of multimodal and multilingual tooling; (2) weak capabilities in data curation and safety evaluation; (3) insufficient system-level monitoring and reproducibility infrastructure; and (4) lack of environmental impact assessment and release governance frameworks. The project delivers a curated practice inventory comprising 250+ open-source tools and resources spanning data governance, training optimization, safety auditing, carbon footprint analysis, and responsible deployment. Empirically grounded, the findings inform policy formulation, tool development, and standardization efforts—advancing AI development from heuristic practice toward a verifiable, auditable, and sustainable engineering paradigm.
Although top-tier conferences such as ICSE now commonly require authors to submit replication packages, the actual executability and reproducibility of these packages remain largely unassessed. This study presents a large-scale empirical investigation of 100 replication packages from ICSE papers published between 2015 and 2024, involving approximately 650 person-hours of manual execution, debugging, and root-cause analysis. The findings reveal that only 40% of the packages are executable, with just 32.5% running without modification; 82.5% require moderate to substantial changes. Among the executable packages, merely 35% successfully reproduce the original results. This work is the first to expose a significant gap between executability and reproducibility in software engineering replication packages and proposes three actionable guidelines to improve their reliability and utility.
LumiXAI解决了解释模型时软件碎片化问题,通过提供一个模块化的全栈框架,整合了特征归因分析,并支持多种用户访问。
Automatically reproducing executable bug-fix code pairs from unstructured developer Q&A posts is hindered by ambiguous descriptions and missing dependencies. This work proposes Reprodgen, the first end-to-end automated framework that leverages large language models to jointly model code intent (CI), functional requirements (FR), and structured chains of thought (SCoT) to generate semantically consistent and executable bug-fix code pairs. The approach incorporates an LLM-based iterative review mechanism coupled with real execution validation to ensure correctness. Evaluated on Stack Overflow and GitHub Issues across seven widely used data science libraries, the study introduces the first expert-validated, runnable benchmark of bug-fix pairs. Experimental results demonstrate that Reprodgen reliably reproduces code pairs exhibiting clear behavioral differences between buggy and fixed versions.
This study addresses the lack of systematic understanding regarding creators’ usage patterns of base and fine-tuned models—such as LoRA adapters—in the current open-source image generation ecosystem. The authors construct a large-scale dataset comprising six million generated images along with their associated metadata, enabling the first empirical analysis of how 22.4K base models and 154K LoRA models are combined and utilized in real-world creative workflows. Through data mining, metadata analysis, and log correlation, the research uncovers distinctive strengths and inherent challenges within this ecosystem. These findings provide empirical grounding for enhancing its sustainability and innovation potential, while the publicly released high-quality dataset supports further community engagement and academic inquiry.
研究使用MLReproMutate软件通过四种变异类别对机器学习研究仓库进行测试,发现现有验证流程常未能检测到影响可复现性的变化。
研究针对Maven生态系统中构建可再现性问题,通过开发AROMA+工具自动化查找库源代码及恢复原始发布环境信息,实现高达99.8%的准确率。