Asymptotic enumeration of admixed arrays and a different independence heuristic

📅 2026-04-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the asymptotic enumeration of dense binary matrices under dual constraints—fixed row sums and pairwise column sums—in the context of large-scale genetic data. By integrating saddle-point approximation, probabilistic methods, entropy inequalities, and combinatorial enumeration, we derive a refined asymptotic expansion for the logarithm of the count. Our key finding reveals that when the sample size \(N = \Theta(P)\) and marginal sums are uniform, the constraint ensemble satisfies an independence heuristic but with a correction factor of \(1/\sqrt[4]{e}\), deviating from the classical \(e^{\pm 1/2}\). We further provide the first explicit fourth-moment contribution and quantitative control of higher-order remainder terms. The work yields an exact counting formula for single-constraint ensembles, an asymptotic theory for dual constraints, and an open-source extended Miller–Harrison algorithm for numerical validation.

Technology Category

Constraint Satisfaction and Optimization: Other Foundations of Constraint SatisfactionKnowledge Representation and Reasoning: Computational Complexity of ReasoningMachine Learning: Matrix & Tensor Methods

Application Category

Graph Algorithms and Modeling for the Web: Representation, reconstruction, and subgraph or motif discovery in Web-related graphsSecurity and Privacy: Large-scale security measurementsEconomics, Online Markets and Human Computation: Incentives in network design for Web infrastructures and ecosystems
📝 Abstract
We introduce a class of paired binary matrices called admixed arrays, which arise in analyses of large-scale genetic data and can be viewed as weighted edge colorings of complete bipartite graphs. This combinatorial structure gives rise to two natural families of marginal constraints: a row-sum constraint and a paired column-sum constraint, the latter inducing an inequality among entries of the matrix pair. We study the enumeration of admixed arrays under these constraints in dense regimes. First, we obtain exact formulas for the sizes of the families defined by each constraint in isolation and derive a finite-size criterion characterizing when one constraint is more restrictive than the other. In the large-dimension limit, this comparison simplifies to an entropy inequality, yielding an information-theoretic interpretation and a quantifiable error bound in the semi-regular case. We then analyze the asymptotic enumeration of the doubly constrained family in a semi-regular setting. Using saddle-point approximation and probabilistic techniques, we derive a detailed asymptotic expansion for the logarithm of the count, isolating an explicit fourth-moment contribution and establishing quantitative control of the higher-order remainder. A consequence of this analysis is a phenomenon absent from classical binary and integer matrix models: in the regime $N=\Theta(P)$ with uniform margins and density bounded away from zero, the two constraint families obey the independence heuristic with a correction factor $1/\sqrt[4]{e}$ rather than the familiar $e^{\pm1/2}$. Numerical experiments corroborate the analytical approximations, and we implement and extend an algorithm of Miller and Harrison (2013) as open-source software to enumerate constrained admixed arrays.
Problem

Research questions and friction points this paper is trying to address.

admixed arrays
asymptotic enumeration
marginal constraints
binary matrices
independence heuristic
Innovation

Methods, ideas, or system contributions that make the work stand out.

admixed arrays
asymptotic enumeration
saddle-point approximation
independence heuristic
entropy inequality
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Alan J. Aw
Department of Genetics, Perelman School of Medicine, University of Pennsylvania