Federated Targeted Maximum Likelihood Estimation

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation that Targeted Maximum Likelihood Estimation (TMLE) is restricted to centralized computation, which hinders privacy-preserving inference across institutions. To overcome this, we propose the first federated TMLE algorithm. Built upon a cross-silo framework with a gradient-aggregated local update mechanism, our method incorporates a finite-precision communication protocol and provides a description length analysis, revealing the novel risk that data localization does not inherently guarantee privacy. Theoretically, we establish non-convex convergence bounds. Empirically, the proposed approach achieves inference accuracy comparable to centralized TMLE with negligible communication error, effectively balancing communication efficiency with rigorous statistical inference capabilities.
📝 Abstract
The evidence behind a scientific or operational decision is often held by hospitals, banks, or registries that cannot pool individual observations. Cross-silo federated learning moves computation to the data and exchanges agreed summaries. Targeted maximum likelihood estimation (TMLE) refines a flexible initial fit, yielding plug-in estimators that respect the model and support efficient inference. TMLE itself, however, has remained a fully centralized procedure. To fill this gap, our paper introduces the first federated TMLE algorithm. We federate targeting itself, for an arbitrary target, loss, and fluctuation family, through two complementary frameworks. FedTMLE-G aggregates local gradients and reproduces centralized targeting step for step. FedTMLE-L lets each institution complete its own fluctuation fit before a single exchange of fitted updates, trading synchronized fidelity for local autonomy. For gradient aggregation, we develop a finite-precision protocol that transmits changes rather than values and certifies targeting accuracy within explicit bounds on exchanges and bits. A description-length analysis of the accepted updates then shows that this finite communication leaves numerical targeting error negligible against sampling uncertainty. The cost of computing an estimator is thus distinct from the complexity of selecting it. Our analysis also indicates that keeping data local is not itself a privacy guarantee of TMLE, since instability of full-record reconstruction need not prevent recovery of a specified sensitive attribute. For a personalized version of local averaging, institutions retain their own estimates and leave once local targeting is complete. A nonconvex convergence bound charges the improvement forfeited through averaging to disagreement among local fits and exposes a tradeoff between equal institutional influence and the sampling variability of small silos.
Problem

Research questions and friction points this paper is trying to address.

Federated Learning
Targeted Maximum Likelihood Estimation
Cross-silo
Privacy
Distributed Inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Federated Learning
Targeted Maximum Likelihood Estimation
Finite-precision Protocol
Description-length Analysis
Nonconvex Convergence
🔎 Similar Papers
2024-10-04IEEE International Symposium on Network Computing and ApplicationsCitations: 3