🤖 AI Summary
This paper addresses sequential experimentation with categorical response variables, where a key challenge arises from outcome-dependent treatment assignment: subsequent interventions are dynamically determined by prior outcomes, inducing dependent observations, missing Cartesian product structure, and random, unpredictable total sample size. To handle this nonstandard data structure, the authors introduce staged trees—a graphical model from algebraic statistics—providing a directed-tree-based parametrization that jointly encodes sequential decision rules and response dependencies. They derive closed-form properties of maximum likelihood estimators, construct valid test statistics, and establish their asymptotic distributions. Extensive simulations and analysis of real sequential clinical trial data demonstrate the method’s validity and robustness. This work extends the algebraic statistical paradigm for sequential experiments and establishes an interpretable, computationally tractable foundation for adaptive design with categorical outcomes.
📝 Abstract
The paper focuses on sequential experiments for categorical responses in which whether or not a further observation is made depends on the outcome of a previous experiment. Examples include subsequent medical interventions being performed or not depending on the result of a previous intervention, data about offsprings, life tables, and repeated educational retraininig until a certain proficiency level is achieved. Such experiments do not lead to data with a full Cartesian product structure and, despite a prespecified initial sample size, the total number of observations, or interventions, made cannot be determined in advance. The paper investigates the distributional assumptions behind such data and describes a parameterization of the distribution that arises and the respective model class to analyze it. Both the data structure resulting from such an experiment and the model class are special examples of staged trees in algebraic statistics. The properties of the resulting parameter estimates and test statistics are obtained and illustrated using hypothetical and real data.