Open Access
August 2020 Minimax optimal rates for Mondrian trees and forests
Jaouad Mourtada, Stéphane Gaïffas, Erwan Scornet
Ann. Statist. 48(4): 2253-2276 (August 2020). DOI: 10.1214/19-AOS1886

Abstract

Introduced by Breiman (Mach. Learn. 45 (2001) 5–32), Random Forests are widely used classification and regression algorithms. While being initially designed as batch algorithms, several variants have been proposed to handle online learning. One particular instance of such forests is the Mondrian forest (In Adv. Neural Inf. Process. Syst. (2014) 3140–3148; In Proceedings of the 19th International Conference on Artificial Intelligence and Statistics (AISTATS) (2016)), whose trees are built using the so-called Mondrian process, therefore allowing to easily update their construction in a streaming fashion. In this paper we provide a thorough theoretical study of Mondrian forests in a batch learning setting, based on new results about Mondrian partitions. Our results include consistency and convergence rates for Mondrian trees and forests, that turn out to be minimax optimal on the set of $s$-Hölder function with $s\in (0,1]$ (for trees and forests) and $s\in (1,2]$ (for forests only), assuming a proper tuning of their complexity parameter in both cases. Furthermore, we prove that an adaptive procedure (to the unknown $s\in (0,2]$) can be constructed by combining Mondrian forests with a standard model aggregation algorithm. These results are the first demonstrating that some particular random forests achieve minimax rates in arbitrary dimension. Owing to their remarkably simple distributional properties, which lead to minimax rates, Mondrian trees are a promising basis for more sophisticated yet theoretically sound random forests variants.

Citation

Download Citation

Jaouad Mourtada. Stéphane Gaïffas. Erwan Scornet. "Minimax optimal rates for Mondrian trees and forests." Ann. Statist. 48 (4) 2253 - 2276, August 2020. https://doi.org/10.1214/19-AOS1886

Information

Received: 1 March 2018; Revised: 1 July 2019; Published: August 2020
First available in Project Euclid: 14 August 2020

MathSciNet: MR4134794
Digital Object Identifier: 10.1214/19-AOS1886

Subjects:
Primary: 62G05
Secondary: 62C20 , 62G08 , 62H30

Keywords: Minimax rates , nonparametric estimation , random forests , supervised learning

Rights: Copyright © 2020 Institute of Mathematical Statistics

Vol.48 • No. 4 • August 2020
Back to Top