FAIR-Compute: A Roadmap for Fair and Efficient Allocation of Federated Digital Research Infrastructure

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses inequities and inefficiencies in computational resource allocation within federated digital research infrastructures (DRIs) by framing resource scheduling as an intersection of technical and policy considerations. Integrating policy analysis, mechanism design, and simulation experiments, the work reveals a significant disconnect between resource allocation and actual utilization. It innovatively applies Braess’s paradox to demonstrate how uncoordinated federation can exacerbate load imbalances. Leveraging algorithmic game theory, transport economics models, and simulations based on both real-world (Fresco/Anvil) and synthetic workloads—complemented by empirical investigation of international DRI allocation mechanisms—the research finds that carefully tuned, transparent scheduling heuristics can closely approximate ideal performance. While federated approaches hold promise, they require coordinated governance, and current systems notably lack effective evaluation of scientific value generated per unit of allocated resources.
📝 Abstract
As demand for high-performance computing (HPC), high-throughput computing and data storage grows, the way scarce compute is allocated -- not just how much exists -- has become a decisive factor in the productivity of UK research. FAIR-Compute studies allocation in a federated Digital Research Infrastructure (DRI) as a problem at the intersection of algorithmic game theory, transport economics and HPC scheduling. Our central observation is simple: once a shared system must decide who runs, when and under what evidence, its scheduler settings cease to be a purely technical matter and become policy. The project combined three strands of evidence: a landscape review of UK and international allocation practice (including EuroHPC, WLCG, ACCESS, JASMIN and DiRAC) supported by a stakeholder survey; a mechanism-design model of allocation under strategic and uncertain user reports; and a simulation study built on the public Fresco/Anvil workload trace and controlled synthetic stress tests. Three results recur across all three strands. First, allocation records measure occupancy (resources reserved) rather than utilisation (useful work done), so the system cannot currently answer the question UKRI and DSIT most want answered -- whether resources deliver productive value. Second, simple, transparent scheduling heuristics perform close to an offline full-information benchmark once tuned, so the near-term opportunity is to tune and instrument existing schedulers rather than replace them. Third, federation is genuinely valuable but behaves like uncoordinated routing when left unmanaged: it can minimise average delay while quietly concentrating load -- and potentially harm -- on the receiving system. This mirrors a classic effect from road-traffic economics (Braess's paradox), where letting everyone independently pick the fastest route can leave the whole network worse off.
Problem

Research questions and friction points this paper is trying to address.

FAIR-Compute
federated Digital Research Infrastructure
resource allocation
HPC scheduling
algorithmic fairness
Innovation

Methods, ideas, or system contributions that make the work stand out.

federated computing
resource allocation
algorithmic game theory
scheduler optimization
Braess's paradox