Institution profile

Optimized Markets, Inc.

Industry researchnorthamerica · us
Research library24linked papers
Opportunities0open roles
Selected work

Representative Papers

Position: Machine Learning for Heart Transplant Allocation Policy Optimization Should Account for Incentives

Feb 04, 2026

This study addresses the misaligned incentives among transplant centers, clinicians, and regulatory agencies—a critical yet overlooked factor that undermines the effectiveness of current heart allocation policies. For the first time, organ allocation is modeled as a multi-agent strategic process, and an incentive-aware allocation framework is proposed that integrates mechanism design with data-driven methodologies. By synthesizing insights from mechanism design, strategic classification, causal inference, and social choice theory, the work systematically quantifies the real-world adverse impacts of incentive misalignment in adult heart transplantation in the United States. The resulting framework offers a novel paradigm for designing allocation policies that are not only more equitable and efficient but also robust to strategic behavior, thereby advancing interdisciplinary innovation at the intersection of machine learning and social science.

1 citationsRead paper

Near-Optimal Dynamic Matching via Coarsening with Application to Heart Transplantation

Feb 04, 2026

This work addresses the lack of theoretical guarantees in online matching algorithms commonly used in critical applications such as organ allocation. We propose a coarse-grained dynamic matching approach based on clustering offline nodes into capacity-constrained groups, which substantially reduces problem complexity while preserving essential structural information. For the first time, we rigorously prove that this coarse-graining strategy incurs no significant performance loss and, in fact, achieves near-optimal theoretical guarantees—thereby providing a formal foundation for clustering-based practices in organ allocation. In simulations of heart transplant allocation, our method performs nearly as well as an omniscient benchmark and significantly outperforms conventional heuristic strategies.

1 citationsRead paper

Learning and Computation of $Phi$-Equilibria at the Frontier of Tractability

Feb 25, 2025

This paper studies $Phi$-equilibrium computation and low-$Phi$-regret online learning under a family $Phi$ of $k$-dimensional polynomial mappings, breaking the classical linear-bias restriction. We propose a nested algorithmic framework based on Ellipsoid-based Adversarial Hopes (EAH), integrating a convex-set oracle model with polynomial expectation fixed-point techniques to achieve efficient $Phi$-equilibrium approximation under polynomial-dimensional bias—first such result. Our contributions include: (i) an $varepsilon$-approximate $Phi$-equilibrium algorithm with time complexity $mathrm{poly}(n,d,k,log(1/varepsilon))$; (ii) an online learning algorithm achieving average $Phi$-regret $leq varepsilon$ within $mathrm{poly}(d,k)/varepsilon^2$ rounds; and (iii) matching upper and lower bounds establishing the first tight learnability characterization for $Phi$-equilibria under polynomial mappings.

1 citationsRead paper

Regularized policy gradient with learned mixtures of Gaussians for games with continuous actions

Sep 29, 2026

This study addresses the reliance on discretization and poor sample efficiency of existing algorithms in continuous-action games. It proposes a scalable policy gradient method that, for the first time, integrates magnetic mirror descent with Gaussian mixture model reparameterization. Through self-play reinforcement learning, this approach effectively approximates Nash equilibria in continuous and mixed-action sequential games, even under gradient failure conditions. The proposed method substantially enhances solution capabilities in continuous action spaces. Compared to neural fictitious self-play and Policy-Space Response Oracles (PSRO), it achieves a 3.5- to 5.5-fold improvement in sample efficiency. Furthermore, its performance in Texas Hold’em poker is comparable to that of Slumbot, demonstrating strong practical efficacy in complex game-theoretic settings.

0 citationsRead paper
Recent publications

Latest Papers

Regularized policy gradient with learned mixtures of Gaussians for games with continuous actions

Sep 29, 2026

This study addresses the reliance on discretization and poor sample efficiency of existing algorithms in continuous-action games. It proposes a scalable policy gradient method that, for the first time, integrates magnetic mirror descent with Gaussian mixture model reparameterization. Through self-play reinforcement learning, this approach effectively approximates Nash equilibria in continuous and mixed-action sequential games, even under gradient failure conditions. The proposed method substantially enhances solution capabilities in continuous action spaces. Compared to neural fictitious self-play and Policy-Space Response Oracles (PSRO), it achieves a 3.5- to 5.5-fold improvement in sample efficiency. Furthermore, its performance in Texas Hold’em poker is comparable to that of Slumbot, demonstrating strong practical efficacy in complex game-theoretic settings.

0 citationsRead paper

Watermarked Game Solving via Perturbed Regret Minimization

Aug 14, 2026

This study addresses the vulnerability of existing game agent watermarking schemes to incomplete information and misuse risks by proposing a novel embedding method based on perturbed regret minimization. For the first time, this approach integrates watermarking mechanisms directly into the learning process of incomplete-information games, enabling both embedding and detection through utility perturbation. The core contribution lies in guaranteeing bounded and negligible exploitability loss while maintaining high detection reliability. Notably, verification can be completed within hours at human play speed. This work effectively bridges the technical gap in copyright protection for agents operating under incomplete information, offering a robust solution that balances security with performance integrity in strategic learning environments.

0 citationsRead paper

Distributed Constraint Optimization via Online Learning and Iterative Pricing with Application to Large-Scale Satellite Scheduling

Jul 28, 2026

This work addresses the challenge of globally solving distributed constraint optimization problems (DCOPs) in large-scale decentralized satellite scheduling under stringent communication constraints. To overcome this, the authors propose a novel framework that integrates online learning with problem decomposition. The approach decouples the DCOP into a high-level task assignment meta-problem and local scheduling subproblems, coordinated through a feedback-driven iterative pricing mechanism. By innovatively combining modern online learning algorithms with potential game modeling, the method efficiently converges to an equilibrium solution. Evaluated on realistic satellite scheduling scenarios, the proposed framework achieves over 99% satisfaction rate for observation requests, substantially outperforming existing baselines, which attain only 87%.

0 citationsRead paper

Aligning Data-Driven Predictors with Allocation: A Decision-Focused Approach to Survival Analysis

Jun 01, 2026

Although traditional survival prediction models achieve strong performance on metrics such as the concordance index (C-index), they often perform poorly—sometimes worse than random selection—in downstream decision-making tasks like organ allocation. This work proposes a decision-oriented approach to survival analysis by introducing normalized discounted cumulative gain (NDCG), a metric from information retrieval, directly aligning prediction objectives with allocation policies. We develop a ranking-based evaluation method tailored for right-censored data and construct a bootstrap-based NDCG optimization framework to fine-tune existing survival models. Evaluated on real-world U.S. heart transplant data, our method improves baseline NDCG by 50–100%, potentially saving tens of thousands of life-years annually while providing theoretical guarantees on allocation performance.

0 citationsRead paper