June 2022 Higher criticism to compare two large frequency tables, with sensitivity to possible rare and weak differences
David L. Donoho, Alon Kipnis
Author Affiliations +
Ann. Statist. 50(3): 1447-1472 (June 2022). DOI: 10.1214/21-AOS2158

Abstract

We adapt Higher Criticism (HC) to the comparison of two frequency tables which may—or may not—exhibit moderate differences between the tables in some unknown, relatively small subset out of a large number of categories. Our analysis of the power of the proposed HC test quantifies the rarity and size of assumed differences and applies moderate deviations-analysis to determine the asymptotic powerfulness/powerlessness of our proposed HC procedure.

Our analysis considers the null hypothesis of no difference in underlying generative model against a rare/weak perturbation alternative, in which the frequencies of N1β out of the N categories are perturbed by r(logN)/2n in the Hellinger distance; here, n is the size of each sample. Our proposed Higher Criticism (HC) test for this setting uses P-values obtained from N exact binomial tests. We characterize the asymptotic performance of the HC-based test in terms of the sparsity parameter β and the perturbation intensity parameter r. Specifically, we derive a region in the (β,r)-plane where the test asymptotically has maximal power, while having asymptotically no power outside this region. Our analysis distinguishes between cases in which the counts in both tables are low, versus cases in which counts are high, corresponding to the cases of sparse and dense frequency tables. The phase transition curve of HC in the high-counts regime matches formally the curve delivered by HC in a two-sample normal means model.

Citation

Download Citation

David L. Donoho. Alon Kipnis. "Higher criticism to compare two large frequency tables, with sensitivity to possible rare and weak differences." Ann. Statist. 50 (3) 1447 - 1472, June 2022. https://doi.org/10.1214/21-AOS2158

Information

Received: 1 July 2020; Revised: 1 December 2021; Published: June 2022
First available in Project Euclid: 16 June 2022

MathSciNet: MR4441127
zbMATH: 07547937
Digital Object Identifier: 10.1214/21-AOS2158

Subjects:
Primary: 62H15 , 62H17
Secondary: 62G10

Keywords: higher criticism , sparse mixture , testing multinomials , two-sample test

Rights: Copyright © 2022 Institute of Mathematical Statistics

Vol.50 • No. 3 • June 2022
Back to Top