large-scale machine learning

Designs and implements machine learning models, training algorithms, and data pipelines that operate efficiently at very large dataset sizes or model scales. Work includes distributed and parallel training and inference, scalable feature engineering and data sharding, resource- and communication-aware optimization, fault tolerance, and system-level engineering to meet performance and cost constraints.

large-scalemachinelearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.6
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$226K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Data-Driven Breakthroughs and Future Directions in AI Infrastructure: A Comprehensive Review

May 22, 2025
BB
Beyazit Bestami Yuksel
🏛️ Istanbul Technical University

This paper addresses critical challenges facing AI infrastructure under tightening data privacy regulations and heightened oversight—specifically, difficulties in acquiring real-world data and the lack of rigorous utility evaluation for synthetic data. It systematically analyzes the co-evolution of computing (GPU proliferation), data (ImageNet-driven centralization), and algorithms (Transformer/GPT breakthroughs) from 2009 to 2024. Methodologically, it innovatively integrates statistical learning theory—particularly sample complexity and data efficiency—into a novel “data site” paradigm unifying federated learning, privacy-enhancing technologies (PETs), and synthetic data generation. The contribution is a unified theoretical framework that rigorously characterizes trade-offs among security, efficiency, and scalability. This framework provides both foundational principles for next-generation AI systems and actionable insights for evidence-based policy design. (138 words)

Analyzing AI breakthroughs via computational, data, and algorithmic convergenceAssessing synthetic data utility amid real-world data accessibility constraintsEvaluating data-centric solutions for privacy, security, and scalability challenges

Must-Read Papers

Most classic and influential ideas
View more

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training

Nov 20, 2024
JF
Jared Fernandez
🏛️ Meta | Carnegie Mellon University

Modern large-scale distributed training faces sharply diminishing returns in hardware scaling: as GPU counts reach thousands, communication overhead dominates performance bottlenecks, rendering conventional parallelism strategies—data, tensor, and pipeline parallelism—suboptimal. Method: Leveraging real-world LLM training workloads, this project establishes an empirical analytical framework spanning diverse model scales, hardware configurations, and parallelization strategies. It quantifies the nonlinear relationship between accelerator count and performance gain, precisely identifying critical inflection points across model, data, and compute scaling dimensions. Contribution/Results: We discover that low-communication “suboptimal” strategies become optimal at extreme scale; we empirically determine hardware selection criteria, cluster topology requirements, and optimal parallelism combinations for training billion-parameter models. Our findings provide actionable, deployment-ready optimization guidelines for trillion-parameter LLM training infrastructures.

Assessing diminishing returns in scaling accelerators for large model trainingEvaluating parallelization strategies to minimize distributed communication overheadOptimizing hardware configuration for efficient large-scale model training

High-Dimensional Data Processing: Benchmarking Machine Learning and Deep Learning Architectures in Local and Distributed Environments

Dec 11, 2025
JJ
José Julián Rodríguez Gutiérrez
🏛️ División de Ingenierías Campus Irapuato-Salamanca

To address the lack of unified benchmarks for model performance evaluation on high-dimensional big data in both local and distributed environments, this work designs an end-to-end evaluation framework covering three representative tasks—Epsilon (numerical regression), RestMex (text classification), and IMDb (movie feature analysis). Leveraging Apache Spark (Scala), we establish a reproducible heterogeneous computing experimental infrastructure to systematically compare traditional machine learning and deep learning models across accuracy, training efficiency, and resource consumption. This study presents the first pedagogically implemented standardized benchmark supporting multiple models, multimodal data, and diverse deployment scenarios, empirically uncovering performance bottlenecks and architectural trade-offs inherent in distributed scaling. The outcomes include an open-source evaluation pipeline, a standardized reporting template, and a reusable teaching paradigm—providing empirical foundations for AI system selection and optimization in big data contexts.

Benchmark machine learning architectures for high-dimensional data processingCompare local and distributed computing environments for big dataImplement workflows for text analysis and classification tasks

Is Your Training Pipeline Production-Ready? A Case Study in the Healthcare Domain

Jun 07, 2025
DL
Daniel Lawand
🏛️ University of São Paulo | Tilburg University | Technical University of Eindhoven

Medical AI deployment is hindered by insufficient production readiness of machine learning (ML) training pipelines. Method: This paper presents a progressive architectural evolution path—monolithic (chaotic) → modular monolithic → microservices—using SPIRA, a voice-based pre-diagnostic system for respiratory insufficiency, as a case study. It systematically introduces continuous training (CT) and a software-quality-attribute-driven MLOps governance framework tailored to healthcare, integrating modular design, microservice decomposition, and engineered CI/CD pipelines. Contribution/Results: The approach significantly improves pipeline maintainability, fault tolerance, and scalability, enabling stable, iterative evolution of SPIRA. It establishes an “agile ML + robust software engineering” co-design paradigm, delivering a reusable methodology and practical benchmark for engineering medical AI in highly regulated environments.

Ensuring ML training pipelines are production-ready in healthcareEvolving architecture for better maintainability and robustnessImproving software quality in MLES for respiratory pre-diagnosis

This study addresses the inherent tension between scalability and maintainability in machine learning (ML) systems—a critical challenge impeding robust, production-grade deployment. Method: We conduct a systematic literature review (SLR) grounded in 124 high-quality publications, developing the first six-dimensional analytical framework spanning data engineering, model engineering, and system deployment. Contribution/Results: The work identifies 41 categories of maintainability issues and 13 categories of scalability issues, uncovering their stage-crossing trade-offs and synergies. It introduces the first taxonomy of scalability–maintainability challenges in ML systems, accompanied by a problem distribution map and an evidence-based repository quantifying solution effectiveness. Collectively, these findings deliver empirically grounded, cross-stage design principles and actionable optimization pathways for industrial ML system development.

Analyzes solutions across ML lifecycle from data to deploymentExplores interdependencies between scalability and maintainability in MLIdentifies scalability and maintainability challenges in ML systems

A Large-Scale Study of Model Integration in ML-Enabled Software Systems

Aug 12, 2024
YS
Yorick Sens
🏛️ Ruhr University Bochum | Chalmers University of Gothenburg

This study addresses the challenges of integrating machine learning (ML) models into software systems—namely, poor integration practices, low reusability, and unclear architectural boundaries. It presents the first large-scale empirical investigation across 2,928 open-source ML-enabled systems. Leveraging GitHub code mining, static analysis, topic modeling, and architectural pattern identification, the work systematically characterizes ML integration topologies, code/model reuse practices, and maintenance bottlenecks. Key contributions include: (1) the first comprehensive classification framework and architectural pattern atlas for ML-enabled systems; (2) identification of seven prevalent integration topologies and four model reuse patterns; and (3) uncovering critical interdisciplinary collaboration barriers in ML-software co-development. The findings bridge the methodological gap between data science and software engineering at the model embedding stage, providing industry-practical architectural guidelines that significantly enhance the maintainability and reusability of ML systems.

ML and code reuse practicesML model integration challengesML-enabled system characteristics

Latest Papers

What's happening recently
View more

This work addresses the inefficiencies in large-scale recommendation systems caused by maintaining separate models for different scenarios and objectives, which hinders development velocity and delays technology adoption. To overcome this, the authors propose the Standardized Model Template (SMT) framework, which leverages composable, standardized machine learning components to enable “design once, deploy everywhere,” uniformly accommodating diverse data distributions and optimization objectives. By decoupling model architecture from scenario-specific configurations, SMT reduces the complexity of technology deployment from O(n·2ᵏ) to O(n+k), breaking away from the conventional “one objective, one model” paradigm. Empirical evaluation on Meta’s ad ranking system demonstrates that SMT improves average cross-entropy by 0.63%, reduces engineering time per model iteration by 92%, and increases the throughput of technology-model pair adoption by 6.3×.

computational advertisinglarge-scale ML ecosystemsML technique propagation

This study systematically investigates the impact of network topology on collective communication performance in large-scale machine learning training. Focusing on Clos (fat-tree) and torus topologies, the work proposes an analytical model to quantitatively evaluate their completion times for operations such as AllReduce, while jointly accounting for network failures and task placement strategies. It establishes, for the first time, a quantitative relationship between topology structure and collective communication efficiency, demonstrating that Clos topologies consistently outperform torus networks across most scenarios—particularly in terms of communication latency, fault tolerance, and scheduling flexibility. These findings provide a rigorous theoretical foundation for designing network architectures tailored to ML training clusters.

Clos networkcollective communicationmachine learning training

This work addresses the challenges of SLO violations and resource inefficiency in machine learning model serving caused by inadequate capacity planning. To this end, the authors propose an adaptive, feedback-driven load testing framework that formalizes the ML serving load testing process for the first time. The framework incorporates real-traffic-based workload calibration and a warm-up mechanism, combined with adaptive search, performance signal feedback control, convergence detection, and GPU monitoring to efficiently estimate the maximum sustainable throughput under SLO constraints. Evaluation across 14 industrial cases demonstrates that the approach reduces capacity estimation error from approximately 30% to 2–6%, with the warm-up mechanism improving accuracy by 22.2%. This significantly mitigates deployment incidents and enhances GPU resource utilization efficiency.

capacity planningload testingML model serving

This work addresses the challenges of low communication efficiency and system complexity in large-scale AI training across geographically distributed data centers. It presents the first systematic characterization and joint optimization of three critical dimensions in “scale-across” training: parallelism strategy deployment, job scheduling, and network transport. By co-designing parallel placement, scheduling policies, and advanced networking techniques—and validating the approach through both real-world testbeds and large-scale simulations—the study achieves end-to-end global optimization of computation and communication. Experimental results demonstrate that the proposed solution improves training throughput by up to 64.62% over current production configurations and by 37.59% compared to state-of-the-art baselines.

communication optimizationdata centerdistributed training

This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.

compression and accelerationconstraint-drivendeployment constraints

Hot Scholars

SB

Subhankar Bhadra

The Pennsylvania State University
Network scienceCausal inference
AA

Abhishta Abhishta

Associate Professor, Finance and Cyber Risk Management, University of Twente
Security EconomicsInternet MeasurementsIT EconomicsCyber Crime
YL

Yixun Liang

PhD student @ Hong Kong University of Science and Technology, HKUST
3D VISION
SŠ

Sanja Šćepanović

Nokia Bell Labs
Ethics in AIResponsible AIAI for HealthSocial Computing