fusion efficiency mapping

Designs and computes efficiency maps (fusion gain maps) that quantify how fusing information sources or estimators changes statistical efficiency relative to baselines; builds functions and visualizations over parameters such as bias magnitude, sample size, and noise that report gain ratios, break‑even bias thresholds, and regions where fusion helps or hurts, including finite‑sample behavior.

fusionefficiencymapping

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Is External Information Useful for Data Fusion? An Evaluation before Acquisition

Jul 29, 2025
GD
Guorong Dai
🏛️ Fudan University | University of Pennsylvania

This paper addresses the fundamental problem of quantifying, *a priori* and using only internal data, the potential efficiency gain conferred by external information on statistical estimation—without access to or reliance on that external information itself. We propose a model- and method-agnostic utility metric—the *estimation efficiency ratio*—grounded in the semiparametric efficiency bound, which characterizes the maximal theoretical improvement attainable from external information. Leveraging the efficient influence function, we construct a computationally feasible estimator solely from internal data and establish its asymptotic normality, enabling both point estimation and valid confidence interval inference. Simulation studies and empirical applications demonstrate robust finite-sample performance. The framework provides a theoretically optimal, broadly applicable pre-evaluation tool for cost-sensitive decisions such as data acquisition and collaborative modeling.

Evaluates utility of external data before acquisitionMeasures efficiency gain from external informationProvides decision framework for cost-effective data fusion

This study addresses the challenge of quantifying the benefit and reliably assessing uncertainty when integrating real-world data with randomized controlled trials. Building on adaptive targeted maximum likelihood estimation (A-TMLE), the authors propose three reproducible tools: a bias model audit report card, a fusion gain efficiency plot, and a selection-aware inference method for data-adaptive gain estimators. Innovatively introducing an auditable bias assessment framework, the work demonstrates that fusion efficiency is governed primarily by the magnitude of bias rather than functional complexity, and establishes a calibrated inference framework. Empirical validation across HIV treatment, public health, and vocational training applications reveals that only the modular jackknife yields calibrated—albeit conservative—confidence intervals, substantially enhancing the reliability of real-world evidence.

data fusionefficiency gainrandomized trial

Data fusion using weakly aligned sources

Aug 28, 2023
SL
Sijia Li
🏛️ University of Washington | Fred Hutchinson Cancer Center | Harvard T.H. Chan School of Public Health

Addressing the challenge of smooth finite-dimensional parameter estimation under weak alignment—where multi-source data exhibit partial,而非 perfect, correspondence and fully aligned samples are scarce—this paper proposes a novel semiparametric data fusion method. We establish, for the first time, the semiparametric efficiency bound under weak alignment and develop a theoretically grounded, robust estimator that jointly models alignment uncertainty and leverages auxiliary information, thereby substantially reducing reliance on fully aligned samples. Our approach relaxes the stringent strong-alignment assumption inherent in conventional fusion frameworks. Applied to an HIV monoclonal antibody prevention trial, it successfully quantifies the association between neutralizing antibodies and viral genotypes, demonstrating improved statistical efficiency and practical applicability. Key contributions include: (i) derivation of the semiparametric efficiency bound under weak alignment; (ii) a computationally feasible, robust fusion algorithm with provable efficiency; and (iii) interpretable, real-world validation in a clinical setting.

Addresses scarcity of fully aligned sources in data fusionEstimates smooth parameters using weakly aligned data sourcesQuantifies efficiency gains from integrating misaligned sources

Semiparametric Efficient Fusion of Individual Data and Summary Statistics

Oct 01, 2022
WH
Wenjie Hu
🏛️ Peking University | Harvard T.H. Chan School of Public Health | Renmin University of China

This study addresses the problem of efficiently integrating internal individual-level data with external aggregate statistics to improve estimation accuracy for population-level functional parameters under weak transportability assumptions. To overcome bias and efficiency loss arising from model misspecification in conventional approaches, we first derive the semiparametric efficiency bound for this fusion setting. Methodologically, we propose two novel estimators: (i) an efficient semiparametric estimator achieving the derived bound, and (ii) an adaptive fusion estimator with oracle properties—ensuring double robustness and automatic bias correction. We establish its asymptotic efficiency and unbiasedness theoretically. Simulation studies and real-data analysis of *Helicobacter pylori* infection demonstrate that our methods significantly enhance estimation precision and statistical power compared to using internal data alone or simple weighted aggregation.

Addressing potential bias and efficiency loss in data integrationEfficient fusion of individual data with external summary statisticsSemiparametric framework for integrating internal and external studies

Harnessing The Collective Wisdom: Fusion Learning Using Decision Sequences From Diverse Sources

Aug 21, 2023
TB
Trambak Banerjee
🏛️ University of Kansas | Fudan University | Yale University

Integrating hypothesis testing results across heterogeneous multi-source studies—some reporting only binary significance decisions, others only FDR control levels—poses a fundamental challenge for rigorous, unified FDR control. Method: We propose the Integrated Ranking and Thresholding (IRT) framework, which operates solely on binary rejection decisions, a prespecified global FDR level, and the set of hypotheses—requiring neither raw data, p-values, nor effect sizes. IRT employs nonparametric evidence aggregation and a ranking-driven thresholding mechanism, circumventing traditional meta-analysis assumptions of statistical homogeneity and reliance on shared summary statistics. Contribution/Results: IRT is the first method to achieve theoretically guaranteed strong FDR control under non-shared statistical summaries. We prove its FDR control property rigorously; simulations demonstrate superior performance over state-of-the-art integration methods; and real-world application to multi-center genome-wide association studies confirms its practical utility and robustness.

Combining findings across diverse data sourcesEnsuring overall false discovery rate controlFusing evidence from multiple testing procedures

Latest Papers

What's happening recently
View more

This study addresses the fundamental trade-off in large language model evaluation among evaluator coupling (γ), policy diversity (measured by entropy H), and few-shot reliability (quantified by the coefficient of variation CV). Extending empirical conditions from five to eleven, the work systematically quantifies the interplay among these three factors and introduces the first standardized benchmark dataset for evaluation. Results reveal a strong negative correlation between γ and H (r = −0.989), indicating that low coupling is accompanied by high measurement noise. Notably, no experimental setting simultaneously achieves γ < 0.2 and CV(N=5) < 0.3, highlighting an inherent tension among these desiderata. The analysis also uncovers anomalous patterns linked to version drift in GPT-4o, offering empirical grounding for the design of more robust and reliable evaluation frameworks.

bias-reliability tradeoffevaluator couplingLLM evaluation

This study addresses the sensitivity of average production inefficiency estimates in stochastic frontier models to benchmark assumptions by introducing, for the first time, the breakdown frontier approach into this framework. By relaxing key identifying assumptions, the paper characterizes the identified set of parameters and derives the breakdown frontier for the parameter of interest to quantify the boundary of misspecification bias. Integrating identified set analysis with sensitivity analysis techniques, the proposed method is validated on classical empirical datasets, demonstrating its effectiveness in assessing robustness. To facilitate reproducibility and practical application, the authors publicly release their implementation code, providing researchers and practitioners with a transparent and accessible tool for evaluating the robustness of inefficiency estimates under potential model misspecification.

average inefficiencybreakdown frontieridentified set

Estimating high-dimensional channel gain maps is highly challenging due to the complexity of the six-dimensional input space and the lack of effective spatial structure modeling. This work proposes a meta-learning-based, cross-environment Transformer estimator that, for the first time, integrates physical laws and environmental characteristics into the modeling process. By employing a physics-informed feature mapping and enforcing invariance constraints such as reciprocity, the method implicitly learns spatial channel gain patterns common across multiple environments. Consequently, it achieves high-fidelity map reconstruction with only a few measurements from a new environment. Experimental results demonstrate that the proposed approach reduces the required number of measurements to one-fifth of those needed by existing methods while significantly improving estimation accuracy.

channel-gain map estimationmetalearningradio map estimation

This study addresses the inherent ambiguity in discrete fuzzy measures (FMs), which cannot be uniquely determined by density parameters alone, thereby complicating the quantification of uncertainty in information fusion. To resolve this issue, the paper proposes—for the first time—an interval-valued fuzzy measure that is uniquely identifiable. The approach constructs initial intervals from density information and further refines the FM by integrating Choquet integrals with empirical data, establishing a confidence relationship between the derived FM and an ideal FM. This framework enables prior characterization of uncertainty in fuzzy integral–based fusion outcomes. Experimental results demonstrate that the Choquet integral computed with the proposed interval-valued FM effectively bounds the ideal fusion result within a well-calibrated confidence interval, significantly enhancing both the reliability and interpretability of the fusion process.

Choquet IntegralDensity-based parametrizationFuzzy Measure

Hot Scholars

GD

Gijs Dubbelman

Associate professor - Eindhoven University of Technology
computer visionrobotics
RB

Rafael Benítez

Dpt. Business Mathematics. University of Valencia
WaveletsDEANumerical analysisIntegral equations
FK

Firuz Kamalov

Canadian University Dubai
operator algebrasmachine learningmathematical financenumerical analysis
TK

Tommie Kerssies

PhD Candidate, Eindhoven University of Technology; Applied Scientist Intern, Amazon
Artificial IntelligenceMachine LearningDeep Learning
MH

Mingyi Hong

Associate Professor, University of Minnesota; Amazon AGI
Machine LearningOptimizationGenerative AISignal processing