bimodality testing

Statistical techniques and model-based tests to detect, quantify, and model two modes in a distribution (e.g., via mixture models or modality statistics) and to identify separation points and trends in empirical data.

bimodalitytesting

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the detection and localization of modes and antimodes in multimodal data by proposing a novel method based on spacings derived from order statistics. The approach smooths the spacing sequence via low-pass filtering and employs nonparametric inference through bootstrap and permutation tests. It further integrates a change-point detection algorithm that fuses parametric and nonparametric information, systematically leveraging both the stability of spacings and their local growth characteristics to jointly identify modes and antimodes—an innovation not previously explored. The method substantially enhances the robustness and accuracy of multimodal structure detection and has been successfully applied to identifying Kirkwood gaps in the asteroid belt. An open-source R package, Dimodal, along with C/Python interfaces (DimodalCPy), is released to facilitate efficient multimodal data analysis.

Kirkwood gapsmodalitymulti-modality

This study addresses the limitation of conventional global goodness-of-fit tests in multivariate settings, which often fail to pinpoint localized model misspecifications. To overcome this, the authors propose a local calibration test based on adaptive partitioning via Beta-trees. Departing from single-statistic global frameworks, the method evaluates whether predicted probabilities fall within finite-sample confidence intervals across data-driven subregions, enabling precise identification and visualization of model inadequacies. By leveraging k-means clustering to generate null distributions and constructing rigorous confidence intervals, the approach effectively detects local deviations in both simulated and real-world datasets, demonstrating superior performance in tasks such as selecting the number of components in mixture models.

Beta-treesgoodness-of-fitlocal deviations

This work proposes a nonparametric method to assess the statistical significance of signal features—such as peaks and plateaus—in data and to detect multimodal structures in inter-event spacing distributions. The approach leverages run theory, employing a Markov chain recursion to precisely characterize the distribution of the longest runs. It integrates permutation testing with a bootstrap procedure tailored for continuous data, enabling a unified evaluation of both high- and low-intensity signal features. The key innovation lies in the first principled synthesis of run-length analysis, permutation tests, and continuous-data bootstrapping, which collectively facilitate accurate detection and localization of multimodal patterns without requiring parametric distributional assumptions, thereby effectively identifying salient morphological features in complex datasets.

bootstrap testmulti-modalitynon-parametric

Analytic inference with two-way clustering

Jun 25, 2025
LD
Laurent Davezies
🏛️ CREST-ENSAE | INRAE

This paper addresses two critical limitations of conventional inference methods under two-way clustering: (i) variance estimators may fail to be positive definite, and (ii) asymptotic validity breaks down under non-Gaussian errors. We propose an analytical correction that reconstructs the variance estimator to guarantee both positive definiteness and asymptotic exactness, while remaining conservative under both Gaussian and non-Gaussian error distributions. Theoretically, the method is proven uniformly valid under general multi-dimensional clustering structures. Simulation studies demonstrate its robust superiority over existing approaches across diverse data-generating processes: it recovers classical inference performance under Gaussian errors, and substantially improves coverage accuracy and test power under non-Gaussian errors. To our knowledge, this work provides the first unified analytical framework for two-way clustered inference that simultaneously ensures theoretical rigor and practical feasibility.

Addresses non-positive variance estimator in two-way clusteringEnsures valid inference in non-Gaussian asymptotic regimesHighlights issues with multiple testing and nonlinear estimators

This study addresses the lack of flexible and robust modeling approaches for financial data exhibiting bimodality, skewness, and heavy tails. To this end, we develop a location-scale mixture model based on the skewed-t distribution of Fernández and Steel (1998) and implement maximum likelihood estimation via the EM algorithm. The proposed framework naturally accommodates any symmetric distribution as its kernel and introduces an innovative likelihood ratio test for assessing component-wise absence of skewness. Simulation studies demonstrate that the method achieves high estimation accuracy and superior fit regardless of whether the assumed model is correctly specified. Empirical analysis successfully uncovers a bimodal structure in S&P 500 returns, providing statistical support for the market characteristic that U.S. equities tend to persist in either bull or bear states over extended periods.

bimodal distributionsflexible modelingheavy-tailed data

Latest Papers

What's happening recently
View more

Hypothesis testing in singular statistical models is often deemed infeasible due to non-identifiable parameters and degenerate Fisher information. This work circumvents these issues by reframing hypotheses in terms of identifiable functionals of the observable distribution rather than unidentifiable parameter functions, thereby recasting the problem as a classical testing problem in the space of probability distributions. The study introduces the novel concept of an “overlap barrier,” which reveals that hypotheses involving non-identifiable quantities inevitably lead to testing impossibility. A Hellinger-distance-based criterion for testability is established, enabling a structural classification of hypotheses in singular models. By integrating distribution-space analysis, posterior contraction theory, and test-driven Bayesian arguments, the framework is validated in Gaussian mixture models and reduced-rank regression, rigorously delineating the boundary between testable and non-testable hypotheses and clarifying the limits of valid statistical inference in singular settings.

hypothesis testingidentifiabilitynon-identifiable parameters

This study addresses fundamental theoretical and applied challenges at the intersection of Bayesian statistics and nonparametric methods, offering a systematic synthesis and extension of the research trajectory in nonparametric Bayesian inference. By incorporating stochastic process priors and rigorous theoretical analysis, the work develops a modeling framework tailored to highly structured random systems. Beyond providing a comprehensive review of core advances in the field, it highlights several emerging directions that address gaps in earlier literature. The resulting contributions significantly advance both the theoretical foundations and practical applicability of nonparametric Bayesian methodologies, thereby strengthening their role in modern statistical modeling and inference.

Bayesian InferenceNonparametric Bayesian StatisticsResearch Overview

This study addresses the lack of rigorous statistical assessment for the reliability of output structures in complex clustering pipelines that involve multiple data-dependent stages such as anomaly detection, feature selection, and clustering. To bridge this gap, the work systematically applies selective inference to the entire clustering analysis workflow, establishing a statistical framework that enables valid significance testing of final cluster assignments. The proposed method rigorously controls the type I error rate at any pre-specified nominal level and demonstrates strong empirical performance on both synthetic and real-world datasets. By doing so, it provides a principled and reliable foundation for statistical inference in multi-stage, data-driven clustering procedures.

clustering pipelinesdata analysis pipelineselective inference

This study addresses the lack of systematic evaluation of statistical power among multivariate two-sample and goodness-of-fit tests, which hinders the selection of effective methods in practice. Through extensive Monte Carlo simulations, it presents the first comprehensive comparison of numerous nonparametric tests under bivariate settings—encompassing both continuous and discrete data—as well as high-dimensional continuous scenarios. Based on empirical findings, the paper proposes a small yet complementary ensemble of methods that collectively ensure high power against a wide range of alternative hypotheses. This ensemble demonstrates strong robustness and broad coverage, significantly outperforming any single test and offering practitioners a reliable, principled recommendation for real-world applications.

goodness-of-fitmultivariate datanon-parametric

This study addresses the challenge of distinguishing between mild and severe outliers in circular data by proposing a three-component Bayesian mixture model. The model employs a symmetric unimodal circular distribution—such as the von Mises or wrapped normal—as a reference component, incorporates a uniform distribution to capture severe outliers, and introduces a low-concentration component sharing the same mean to represent mild anomalies. This dual-contamination framework uniquely enables automatic identification and quantification of both outlier types within a unified probabilistic structure, without requiring predefined thresholds. It further yields interpretable estimates of outlier proportions and dispersion inflation. Simulation studies and real-data analyses—including applications to animal movement and wind direction—demonstrate that the proposed approach substantially enhances model robustness and effectively uncovers latent structures in directional data.

anomaly detectioncircular datagross anomalies

Hot Scholars

NN

Nassir Navab

Professor of Computer Science, Technische Universität München
ZJ

Zhongliang Jiang

University of Hong Kong
Medical RoboticsUltrasound imagingRobot learningSurgical Robotics
AO

Aydogan Ozcan

Chancellor's Professor at UCLA & HHMI Professor
Computational ImagingHolographyMicroscopySensing
YL

Yuzhu Li

University of California, Los Angeles
Computational imagingOptical imaging and sensingMachine learning
DH

Dianye Huang

Technical University of Munich
robotic ultrasoundmedical robotintelligent controlhuman robot interaction