ilp partitioning

Formulates and solves integer linear programming (ILP) models to partition computations, data, or hardware resources into groups or stages subject to architectural, resource, and parallelism constraints. Builds ILP encodings that capture latency, allocation, and low-level effects and integrates empirical profiling or analytical models to compute latency- or resource-optimal partitions.

ilppartitioning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Integer Linear Programming (ILP) suffers from low computational efficiency on conventional CPUs and GPUs due to its inherent sparsity and strong branching behavior, making it ill-suited for real-time decision-making. This work proposes SPARK, a near-cache accelerator integrated alongside the CPU L1 cache, which for the first time co-designs sparsity awareness, computation reuse, and near-cache architecture, incurring only 1.4% area overhead. By leveraging sparse pattern detection, reuse-aware scheduling, and a low-power data path, SPARK efficiently supports both sparse and dense ILP as well as Linear Programming (LP) problems. Evaluated on the MIPLIB 2017 benchmark suite, SPARK achieves 15–20× speedup over a Zen3 CPU and an NVIDIA V100 GPU for sparse ILP instances, with energy efficiency improvements ranging from 152× to 740×.

branch-intensiveenergy efficiencyInteger Linear Programming

This work addresses key challenges in integer linear programming (ILP) solvers—limited generalization, reliance on external solvers, and poor efficiency in multimodal energy landscapes—by proposing a training-free, solver-free sampling-based optimization framework that directly explores the discrete feasible region. Leveraging the linear structure of ILP, the method designs a proposal distribution satisfying detailed balance and incorporates a dual tempering mechanism that jointly modulates temperature and penalty parameters to dynamically adjust constraint barriers while preserving the original objective function, thereby enhancing global exploration. Experiments demonstrate that the approach consistently outperforms SCIP across four benchmarks, matches or exceeds Gurobi’s performance on two tasks within 200 seconds, exhibits superior robustness on out-of-distribution instances, and competes effectively with classical solvers on MIPLIB 2017 without any parameter tuning.

Combinatorial OptimizationConstraint SatisfactionGeneralization

This work addresses the challenge of solving large-scale combinatorial scheduling problems, which are NP-hard and often lead to slow convergence and large optimality gaps when tackled by traditional integer linear programming (ILP) solvers on industrial instances. The paper introduces the first CPU-GPU hybrid framework that leverages differentiable optimization to warm-start state-of-the-art ILP solvers such as CPLEX, Gurobi, and HiGHS. By rapidly generating high-quality partial solutions through differentiable presolving, the approach provides strong initial feasible solutions that significantly enhance early pruning efficiency within branch-and-bound search. This integration of machine learning with exact optimization yields substantial performance gains: on industrial benchmarks, it achieves up to a 10× speedup over existing baselines while reducing the optimality gap to below 0.1%.

combinatorial schedulingInteger Linear ProgrammingNP-hard

This study addresses the problem of minimizing makespan on a fixed number of parallel machines. It introduces, for the first time, an efficient approach based on short integer linear programming (Short ILP) to design a quasipolynomial-time algorithm. When the maximum processing time $p_{\text{max}}$ is moderate, the algorithm achieves a running time of either $\widetilde{O}(p_{\text{max}}^{O(1)} + n)$ or $\widetilde{O}(p_{\text{max}}^{O(1)} \cdot n)$, significantly outperforming existing methods. This work not only extends the applicability of Short ILP techniques to scheduling problems but also provides a more efficient solution pathway for instances of moderate scale.

makespan minimizationparallel machinespseudo-polynomial time

LLMs for Cold-Start Cutting Plane Separator Configuration

Dec 16, 2024
CL
Connor Lawless
🏛️ Stanford University

Configuring Mixed-Integer Linear Programming (MILP) solver parameters—particularly for cutting-plane separation—is challenging due to high-dimensional, problem-dependent search spaces; existing machine learning approaches suffer from poor generalization, heavy reliance on large-scale labeled data, and difficulty integrating into solver pipelines. Method: We propose the first LLM-driven zero-shot cutting-plane separator configuration framework. It leverages large language models to jointly parse natural-language problem descriptions and LaTeX-based MILP formulations, augmented by literature-informed prompt engineering and semantic modeling of separators—requiring no custom interfaces or extensive retraining. A lightweight, performance-driven clustering ensemble strategy ensures both robustness and real-time responsiveness. Results: On benchmark combinatorial optimization instances and real-world datasets, our method matches state-of-the-art performance while reducing training data requirements by over 90% and generating configurations in under one second.

Configuring MILP solver parameters is difficult for non-expert usersCurrent approaches are hard to integrate into existing solver workflowsExisting ML methods require extensive training data and generalize poorly

Latest Papers

What's happening recently
View more

This study addresses the NP-hard problem of hardware/software partitioning in computing architectures. Leveraging the directed pathwidth of task graphs, this work proposes a novel family of problem formulations that subsumes existing models, along with exact fixed-parameter tractable (FPT) algorithms. Methodologically, by integrating directed pathwidth analysis, FPT theory, and integer linear programming (ILP), the proposed approach achieves exact and efficient solutions for this problem family. The primary theoretical contribution lies in extending the modeling framework for hardware/software partitioning and establishing its fixed-parameter tractability. Empirically, experiments on real-world application scenarios demonstrate that the proposed method achieves up to a 200-fold speedup over general-purpose ILP solvers such as Gurobi, highlighting its practical efficacy and computational advantage.

computational cost optimizationHardware-Software Partitioningmakespan minimization

Mixed-integer programming (MIP) is a cornerstone in applied optimization, both in industry and academia. Recently, there has been increased attention to finding strong primal solutions quickly. This is reflected, for example, in the development of the NVIDIA cuOpt solver and, most recently, in the new MIPFEAS benchmark, which has a tight time limit of 600 seconds and evaluates solvers based on how quickly they find high-quality primal solutions. This article introduces a MIP portfolio parallelization scheme, focusing on efficiently exchanging information between its workers. We present two implementations of this scheme: one built directly into the open-source MIP solver SCIP, and an external one, which we call ReXi. ReXi is currently the fastest non-commercial solver in the MIPFEAS benchmark, followed by the SCIP-integrated implementation. Moreover, we present new versions of both implementations that considerably outperform their predecessors on the MIPFEAS benchmark.

MIPFEAS benchmarkMixed-integer programmingprimal solutions

This study investigates the role of task replication in graph partitioning and DAG scheduling, aiming to substantially reduce or even eliminate inter-processor communication with minimal computational overhead. It presents the first systematic analysis of how replication affects the computational complexity of these two problems and introduces an optimal replication model based on integer linear programming (ILP) alongside an efficient heuristic algorithm tailored for large-scale DAGs. Experimental results demonstrate that, in hypergraph partitioning, communication cost is reduced by 17%–65% on average, with complete elimination in certain scenarios; in DAG scheduling, reductions range from 11.61% to 23.13% on average, reaching as high as 58.17%. These findings confirm the effectiveness and practical value of replication strategies.

communication costDAG schedulinggraph partitioning

This work addresses the limitations of existing machine learning workload partitioning approaches for compute-in-memory (CIM) systems, which often overlook critical RRAM constraints—including storage capacity, high write latency, and endurance—and fail to exploit the full potential of CPU–CIM协同 computation. To overcome these challenges, the paper introduces the first unified integer linear programming (ILP) framework that jointly models RRAM physical constraints, parallelism, and heterogeneous resource scheduling to minimize end-to-end inference latency while respecting hardware limitations. By integrating empirical performance profiling with analytical modeling, the proposed method achieves globally optimal workload partitioning and enables design space exploration for CIM accelerators. Experimental results demonstrate significant speedups of 30.9× and 7.3× over CPU-only execution on edge and high-performance CPU platforms, respectively, substantially enhancing inference efficiency in heterogeneous systems.

Computing-in-MemoryHeterogeneous ComputingMachine Learning

This work addresses the significant degradation in inference throughput caused by GPU memory constraints when concurrently deploying multiple large language models on shared heterogeneous hardware, where resource scheduling, model offloading, and preemption become critical bottlenecks. Through empirical methodologies—including cross-platform performance profiling, layer-wise offloading experiments, and fine-grained decomposition of preemption overhead—the study systematically uncovers, for the first time, the nonlinear relationship between offloading and throughput decline. It further identifies model state reloading as the primary source of preemption overhead. The findings reveal that smaller models are more sensitive to reduced GPU residency, and that such overhead is jointly influenced by model architecture and hardware characteristics. These insights motivate a scheduler design that integrates model-specific sensitivity with data migration costs, offering crucial guidance for building efficient multi-model serving systems.

CPU-GPU offloadingGPU memory constraintsheterogeneous hardware

Hot Scholars

SF

Simon Flügel

University of Osnabrück
neuro-symbolic integrationcheminformatics
TM

Till Mossakowski

Professor of Computer Science, University of Osnabrück
Logicformal ontologyknowledge representationneuro-symbolic AI
AC

Anupam Chattopadhyay

Associate Professor, CCDS, NTU, Singapore
EDACPS SecurityAI SecurityQuantum Computing