2609005160
  • Open Access
  • Article

Gap-Entropy Testing (GET): A Framework for Structural and Distributional Differences

  • Vangelis D. Karalis  1,2

Received: 12 May 2026 | Revised: 05 Sep 2026 | Accepted: 14 Sep 2026 | Published: 18 Sep 2026

Abstract

The classical two-sample problem is framed as the detection of differences in marginal distributions, in terms of location, scale, or overall shape. In this paper, Gap–Entropy Testing (GET) is proposed, a novel nonparametric framework that extends the two-sample paradigm by adding a spacing-sensitive component to conventional distributional analysis. The approach is based on order statistics, log1p-transformed adjacent spacings normalized by their median gap, and entropy estimation using a Vasicek-type estimator. Inference is performed by permutation testing under a common continuous i.i.d. null. Three complementary statistics were developed: GET-V compares spacing-entropy estimates; GET-Plus combines spacing and median-location components; and GET-Ω additionally incorporates energy distance. Under the common i.i.d. null, GET-V has a type-I error close to the nominal 0.05 level and is insensitive to a pure location shift. For the lattice and micro-cluster example, GET-V has large rejection rateσ, which is related to the spacing regularity and the organization of the local gap. GET-V is also sensitive to marginal shape, and individual-label permutation is anti-conservative under matched lattice and shifted-lattice nulls. Application to public benchmark datasets allows us to use GET-V for continuous variables without ties. GET is a reproducible and interpretable framework that adds spacing information to standard two-sample comparisons.

 

Graphical Abstract

References 

  • 1.

    Student. The Probable Error of a Mean. Biometrika 1908, 6, 1–25.

  • 2.

    Box, G.E.P. Non-Normality and Tests on Variances. Biometrika 1953, 40, 318–335. https://doi.org/10.1093/biomet/40.3-4.318.

  • 3.

    Welch, B.L. The Generalization of “Student’s” Problem when Several Different Population Variances Are Involved. Biometrika 1947, 34, 28–35. https://doi.org/10.1093/biomet/34.1-2.28.

  • 4.

    Lehmann, E.L. Nonparametrics: Statistical Methods Based on Ranks; Holden-Day: San Francisco, CA, USA, 1975.

  • 5.

    Mann, H.B.; Whitney, D.R. On a Test of Whether One of Two Random Variables Is Stochastically Larger than the Other. Ann. Math. Stat. 1947, 18, 50–60. https://doi.org/10.1214/aoms/1177730491.

  • 6.

    Hájek, J.; Šidák, Z. Theory of Rank Tests; Academic Press: New York, NY, USA, 1967.

  • 7.

    Kolmogorov, A.N. Sulla Determinazione Empirica Di Una Legge Di Distribuzione. G. Ist. Ital. Attuari 1933, 4, 83–91.

  • 8.

    Smirnov, N. Table for Estimating the Goodness of Fit of Empirical Distributions. Ann. Math. Stat. 1948, 19, 279–281. https://doi.org/10.1214/aoms/1177730256.

  • 9.

    Cramér, H. Mathematical Methods of Statistics; Princeton University Press: Princeton, NJ, USA, 1946.

  • 10.

    Székely, G.J.; Rizzo, M.L. Energy Statistics: A Class of Statistics Based on Distances. J. Stat. Plan. Inference 2013, 143, 1249–1272. https://doi.org/10.1016/j.jspi.2013.03.018.

  • 11.

    Székely, G.J.; Rizzo, M.L.; Bakirov, N.K. Measuring and Testing Dependence by Correlation of Distances. Ann. Statist. 2007, 35, 2769–2794. https://doi.org/10.1214/009053607000000505.

  • 12.

    Gretton, A.; Borgwardt, K.; Rasch, M.J.; et al. A Kernel Two-Sample Test. J. Mach. Learn. Res. 2012, 13, 723–773.

  • 13.

    Berlinet, A.; Thomas-Agnan, C. Reproducing Kernel Hilbert Spaces in Probability and Statistics; Springer US: New York, NY, USA, 2004. https://doi.org/10.1007/978-1-4419-9096-9.

  • 14.

    Pyke, R. Spacings. J. R. Stat. Soc. Ser. B Stat. Methodol. 1965, 27, 395–436. https://doi.org/10.1111/j.2517-6161.1965.tb00602.x.

  • 15.

    Vasicek, O. A Test for Normality Based on Sample Entropy. J. R. Stat. Soc. Ser. B Stat. Methodol. 1976, 38, 54–59. https://doi.org/10.1111/j.2517-6161.1976.tb01566.x.

  • 16.

    Good, P. Permutation Tests: A Practical Guide to Resampling Methods for Testing Hypotheses, 2nd ed.; Springer: New York, NY, USA, 2000. https://doi.org/10.1007/978-1-4757-3235-1.

  • 17.

    Edgington, E.S.; Onghena, P. Randomization Tests, 4th ed.; Chapman and Hall/CRC: Boca Raton, FL, USA, 2007. https://doi.org/10.1201/9781420011814.

  • 18.

    Ebrahimi, N.; Pflughoeft, K.; Soofi, E.S. Two Measures of Sample Entropy. Stat. Probab. Lett. 1994, 20, 225–234. https://doi.org/10.1016/0167-7152(94)90046-9.

  • 19.

    Fedesoriano. Stroke Prediction Dataset. Available online: https://www.kaggle.com/datasets/fedesoriano/stroke-prediction-dataset (accessed on 1 May 2026).

  • 20.

    El Kharoua, R. Alzheimer’s Disease Dataset. Available online: https://www.kaggle.com/datasets/rabieelkharoua/alzheimers-disease-dataset (accessed on 1 May 2026).

  • 21.

    Adeyemi, T. Dementia Patient Health Dataset. Available online: https://www.kaggle.com/datasets/timothyadeyemi/dementia-patient-health-dataset (accessed on 1 May 2026).

Share this article:
How to Cite
Karalis , V. D. Gap-Entropy Testing (GET): A Framework for Structural and Distributional Differences. Applied Mathematics and Statistics 2026, 3 (2), 19. https://doi.org/10.53941/ams.2026.100019.
RIS
BibTex
Copyright & License
article copyright Image
Copyright (c) 2026 by the authors.
Article Metrics
15
Article Views
0
Citations