centrality-based data mixing

Designs and implements data sampling and mixture strategies for model training or pretraining that weight or combine examples according to centrality scores computed from an underlying graph or interaction network. Builds pipelines that integrate centrality with other quality or metadata signals and evaluates how centrality-weighted mixtures affect learned representations and downstream task performance.

centrality-baseddatamixing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.2
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the high computational cost of existing pretraining data selection methods, which often rely on auxiliary models or labeled data. The authors propose WebGraphMix, a novel framework that leverages the host-level web graph topology derived from Common Crawl to guide data mixing without requiring additional training or annotations. By employing unsupervised centrality scores, WebGraphMix adaptively balances the proportion of core and peripheral documents, revealing their complementary learning value in the web graph—a signal orthogonal to existing content-quality metrics. Evaluated on the DataComp-LM benchmark, a simple 1:1 mixing strategy achieves an average score of 41.4% on 400M and 1B language models, outperforming uniform sampling (39.8%). Further gains are realized by combining this approach with content-quality scoring, yielding a performance of 43.8%.

Common Crawldata curationlanguage models

This work addresses the limitations of existing meta-learning approaches for predicting machine learning pipeline performance (PPE) and estimating dataset similarity (DPSE), which predominantly rely on dataset meta-features while overlooking rich historical experiments and pipeline metadata, thereby failing to effectively model interactions between datasets and pipelines. To overcome this, the study introduces knowledge graph embedding into meta-learning for the first time, constructing a unified knowledge graph that integrates datasets, pipelines, and large-scale experimental results from 144,177 OpenML experiments. By jointly leveraging meta-features and empirical performance records, the proposed method—KGmetaSP—explicitly captures the complex interactions between datasets and pipelines. Remarkably, KGmetaSP achieves substantial improvements in both PPE prediction accuracy and DPSE retrieval effectiveness using a single, general-purpose meta-model, establishing a novel paradigm for cross-dataset meta-learning.

dataset similarityknowledge graph embeddingsmeta-features

The Limits of Graph Samplers for Training Inductive Recommender Systems: Extended results

May 20, 2025
TE
Theis E. Jendal
🏛️ Aalborg University | University of Verona | TU Wien

This work systematically investigates the applicability boundaries of graph sampling in inductive recommendation training on heterogeneous temporal user-item graphs. Addressing the challenge of scalability in inductive learning, we comparatively evaluate six sampling strategies—including RandomNode and LayeredSampling—across three real-world datasets, providing the first empirical analysis of their impact on inductive performance. We identify a critical sampling threshold: retaining only 50% of the original graph preserves full model accuracy while accelerating training by up to 86%; below this threshold, performance degrades sharply. Crucially, we reveal that temporal dynamics constitute a fundamental constraint on sampling quality, rendering conventional static sampling methods inadequate for inductive settings. Consequently, we argue for the urgent development of temporally aware graph sampling mechanisms jointly optimized with inductive recommendation models. Our findings establish theoretical foundations and practical guidelines for scalable, efficient graph-based recommendation training.

Highlights need for temporal-aware sampling in recommendation graphsInvestigates graph sampling for inductive recommender systems' efficiencyTests performance impact of reduced training data on accuracy

This study addresses the underutilization of knowledge distillation (KD) for pretrained models in distributed and federated learning settings, systematically evaluating multiple KD variants under heterogeneous data distributions. We propose the first lightweight, practical KD framework tailored for federated learning, incorporating multi-strategy data partitioning and hyperparameter sensitivity analysis to uncover adaptation patterns for critical hyperparameters—such as temperature scaling and loss weighting. Innovatively, we integrate tuned KD, deep mutual learning, and data-partitioned KD, jointly optimized via grid search. Experiments demonstrate that our approach significantly reduces communication rounds and accelerates model convergence in federated learning. Moreover, it delivers reusable, scenario-specific optimal KD configurations across diverse data partitioning schemes, boosting student model accuracy by 3.2–5.7% on average.

Enhancing federated learning through reduced communication demandsEvaluating KD techniques across data distribution strategiesOptimizing knowledge distillation in pre-trained models

Latest Papers

What's happening recently
View more

This study addresses the challenge of unified network modeling for heterogeneous variables—encompassing continuous, count, and categorical types—within arbitrarily structured multilayer topologies. Building upon mixed graphical models (MGM), the work proposes a flexible framework for both single-layer and multilayer network construction. Edge weights and centrality measures are accompanied by confidence intervals derived via bootstrap resampling, while community stability is assessed through network clustering. An integrated Shiny-based interface enables interactive visualization of results. Notably, this project delivers the first comprehensive R toolkit that simultaneously supports heterogeneous data, customizable interlayer architectures, uncertainty quantification, and community stability analysis, thereby substantially enhancing the interpretability and robustness of multilayer network modeling.

heterogeneous variablesMixed Graphical Modelsmultilayer networks

Existing Mixture-of-Experts (MoE) models lack a unified, modular framework for systematic construction and analysis. Method: This paper introduces MixtureKit, an open-source, modular framework enabling non-intrusive integration of three MoE paradigms—classical MoE, fine-grained routing via BTX (Branch-Train-Mix), and expert freezing with learnable stitching via BTS (Branch-Train-Stitch)—into arbitrary pre-trained or fine-tuned models (e.g., Hugging Face Transformers). MixtureKit pioneers BTX and BTS architectures, supporting cross-model and cross-layer dynamic expert routing, differentiable routing mechanisms, token-level expert assignment visualization, and contribution analysis based on attention weights, complemented by an interactive web interface. Results: On Arabic–Latin code-mixing tasks, BTX achieves 2.3–4.7 BLEU/ACC gains over same-scale dense baselines. The framework is publicly released to advance MoE research and deployment across domains.

Develops a framework for constructing and training Mixture-of-Experts models from existing modelsOffers visualization tools to analyze routing decisions and expert contributions in modelsProvides methods for fine-grained token routing and controlled information exchange between experts

This work addresses the limited interpretability of Generative Flow Networks (GFlowNets) during training, particularly the lack of intuitive understanding regarding exploration of the sample space, trajectory construction, and the evolution of sampling probabilities. To bridge this gap, we introduce GFlowState, the first interactive visual analytics system enabling fine-grained, dynamic inspection of GFlowNets training processes. GFlowState integrates multiple coordinated views—including candidate ranking graphs, state projections, trajectory node-link diagrams, and transition heatmaps—to support comprehensive trajectory analysis, comparative exploration of sample spaces, and diagnosis of training failures. Case studies demonstrate that GFlowState substantially enhances the interpretability of GFlowNets, accelerates debugging workflows, and improves the assessment of model quality.

Generative Flow Networksinterpretabilitysample space exploration

This work addresses the lack of theoretical guidance in data mixing strategies for large language model training—particularly challenges arising from ambiguous domain definitions, discrepancies between human and model cognition, and the impact of weighting on generalization. The study is the first to theoretically establish a connection between domain distributions and gradient dynamics, formulating data scheduling as a graph-constrained optimization problem. It introduces DoGraph, a gradient-dynamics-based data reweighting framework that explicitly models inter-domain relationships to derive improved mixing strategies. Experiments across multiple GPT-2 model scales demonstrate that DoGraph significantly enhances model generalization and achieves competitive performance.

data mixingdomain definitiondomain weighting

This work addresses the lack of a unified evaluation framework for knowledge graph integration pipelines, which hinders systematic comparison and selection of methods. To bridge this gap, the paper introduces KGI-Bench, the first comprehensive benchmark specifically designed for evaluating knowledge graph data integration. KGI-Bench assesses integration performance across three key dimensions—coverage, correctness, and consistency—when incorporating heterogeneous input data (structured, semi-structured, and unstructured) into a target knowledge graph. Using a curated dataset in the movie domain, the benchmark evaluates twelve representative integration pipelines, revealing significant performance variations attributable to input data types and architectural choices. The results demonstrate the effectiveness and practical utility of KGI-Bench in enabling rigorous, reproducible evaluation of knowledge graph integration approaches.

benchmarkdata integrationknowledge graph

Hot Scholars

BP

Becky P.Y. Loo

Professor of Geography, The University of Hong Kong
Sustainable TransportSmart CitiesRoad SafetyWalking
FA

Fady Alajaji

Queen's University
Information TheoryCommunicationsJoint source-channel codingPolya contagion networks
BG

Bahman Gharesifard

Professor of Mathematics at Queen's University
Control TheoryOptimizationReinforcement LearningNeural Networks
EM

Esteban Moro

Professor of Physics & Network Science Institute, Northeastern University
Computational social scienceNetwork Sciencecomplex systemsfinancial markets