Score
Designs and implements pipelines that identify and derive observable indicators (features or signals) from raw data, including filtering noisy or spurious signals. Formats and maps those extracted indicators into signature- or rule-compatible representations for use by detection or rule-based systems.
This work addresses the challenges of high early-stage uncertainty in manufacturing monitoring system development—leading to redundant modeling and substantial training costs—and the limited transferability of filtering pipelines in cross-domain image segmentation tasks. To tackle these issues, the authors propose a problem-centric design paradigm that constructs an abstract system model to continuously accumulate and retrieve historical segmentation tasks along with their associated filtering pipelines, enabling solution reuse and incremental optimization. The approach integrates similarity-based problem retrieval, abstract modeling, pipeline reuse, and a retrieval-augmented evolutionary learning mechanism. Experimental results demonstrate that the method significantly reduces training costs and late-stage revision risks, provides the first systematic validation of filtering pipeline transferability across similar segmentation tasks, and achieves a favorable balance among complexity, technical requirements, and reliability under lightweight model constraints.
This work addresses a critical limitation in existing chart-to-code generation methods, which rely on reference code containing unobservable latent variables for supervision, often leading to model hallucination and over-specification. The study systematically identifies this issue and introduces an observation-aligned supervision framework that restricts training objectives to quantities directly inferable from chart images—such as boxplot statistics, pie chart proportions, and histogram bin weights—ensuring alignment between supervision signals and visual observations. By integrating chart understanding from vision-language models with supervised fine-tuning and data rewriting techniques, the proposed approach significantly improves both the accuracy of observable attribute recovery and code executability on ChartMimic and ChartX benchmarks, demonstrating the pivotal role of observation-aligned supervision in enhancing model performance.
Small-object detection performance is hindered by fragmented optimization across stages in conventional pipeline-based detectors. To address this, we propose PLUSNet, an end-to-end co-optimization framework introducing the novel “Purify–Label–Utilize” paradigm. Specifically, we design a hierarchical feature purifier to suppress noise; develop a multi-criterion dynamic label assignment mechanism to improve positive/negative sample quality; and introduce a frequency-domain decoupled detection head for fine-grained feature modeling. All modules are lightweight, modular, and seamlessly integrate with mainstream detectors. Extensive experiments on MS COCO, VisDrone, and other benchmarks demonstrate consistent and significant gains in small-object AP (+3.2–5.8 points), validating the effectiveness of joint upstream-downstream optimization and strong generalizability across diverse scenarios and architectures.
Slug flow in oil and gas pipelines poses significant safety risks, yet conventional detection methods rely on offline analysis and expert knowledge, lacking real-time capability and interpretability. Method: This paper proposes an end-to-end, interactive, data-driven system supporting a full closed-loop workflow—from CSV data ingestion and interactive visualization-based labeling to snapshot-persistent model training and real-time inference. It innovatively integrates time-series superposition visualization, persistent alerting mechanisms, and configurable multi-classifiers to enable human-in-the-loop modeling and transparent, explainable diagnostics. Contribution/Results: The lightweight, plug-and-play system demonstrates high detection accuracy and robustness in real industrial deployments. Its modular architecture ensures seamless adaptability to other time-series fault diagnosis tasks, offering strong generalizability and practical scalability.
Existing methods for digitizing engineering diagrams—particularly Piping and Instrumentation Diagrams (P&IDs)—suffer from incomplete structural information extraction and weak global topological modeling. Method: This paper proposes an end-to-end graph-structure extraction framework featuring (i) Relationformer, a novel architecture jointly modeling symbol detection and relational reasoning; (ii) an image tiling-and-stitching strategy to enhance accuracy on large-scale diagrams; and (iii) PID2Graph, the first open-source, graph-structured annotation dataset for P&IDs, accompanied by a unified evaluation framework. Results: Experiments on real-world P&ID data demonstrate that our approach improves edge detection accuracy by over 25% compared to state-of-the-art modular pipelines, significantly enhancing topological completeness and practical utility. This work establishes a new paradigm and foundational infrastructure for intelligent parsing of industrial engineering drawings.
In supervised learning, modular decomposition lacks identifiability when models exhibit equivalent local linear responses under input perturbations—termed the “mirage regime”—raising the question of whether internal module assignments are uniquely recoverable. Method: We propose Modular Jets (MoJet), a differential-geometric framework that estimates first-order module-wise responses to infinitesimal input perturbations (empirical jets), integrating task manifold geometry and module-level representations to formulate a rank-based criterion distinguishing the mirage from the identifiable regime. Contribution/Results: We theoretically prove that, in two-module linear regression, jets uniquely recover the ground-truth decomposition. Algorithmically, MoJet implements jet estimation and mirage diagnosis. Experiments demonstrate its effectiveness across linear and deep regression and classification pipelines. Crucially, MoJet enables the first *identifiability diagnosis*—inferring modular structure directly from input-output behavior—moving beyond conventional reliance on predictive risk alone.
This study addresses the compliance challenges faced by data practitioners in machine learning systems under regulations such as the GDPR and the AI Act, particularly concerning data quality. Through semi-structured interviews with practitioners in the European Union, combined with thematic analysis of regulatory texts and engineering workflows, the research systematically uncovers a structural disconnect between regulation-driven data quality requirements and ML engineering practices. It identifies five core challenges: misalignment between legal principles and engineering implementation, fragmented data pipelines, lack of purpose-built compliance tools, ambiguous accountability, and reactive responses to audits. Building on these findings, the work proposes directions for designing compliance-oriented tooling, establishing effective governance mechanisms, and fostering cultural transformation to bridge the gap between regulatory mandates and practical ML development.
In highly regulated domains such as finance, data quality control (QC) is often fragmented into isolated preprocessing steps, undermining end-to-end trustworthy AI pipelines. To address this, we propose the first AI-driven DataOps framework that embeds QC as a system-level core component. Our framework deeply integrates rule-based engines, statistical analysis, and custom AI-powered anomaly detection across the entire data lifecycle—from ingestion and transformation to model deployment—enabling dynamic remediation, policy-configurable workflows, and end-to-end auditability. Technically, it unifies data profiling, stream processing, cloud-native storage interfaces, and a proprietary AI detection module. Evaluated in a real-world financial production environment, the framework achieves significantly improved anomaly recall, reduces manual intervention by 42%, and ensures audit completeness and full data traceability under high-throughput conditions, fully satisfying regulatory compliance requirements.