Random forests for big data

Nathalie N. Villa-Vialaneix; Robin Genuer; Jean-Michel Poggi; Christine Tuleau-Malot

Communication Dans Un Congrès Année : 2016

Random forests for big data

(1) , (2) , (3) , (4)

1
2
3
4

Nathalie N. Villa-Vialaneix

Fonction : Auteur
PersonId : 4221
IdHAL : nathalie-vialaneix
ORCID : 0000-0003-1156-0639
IdRef : 101680503

Unité de Mathématiques et Informatique Appliquées de Toulouse

Robin Genuer

Fonction : Auteur
PersonId : 1787
IdHAL : robin-genuer
IdRef : 15657490X

Statistics In System biology and Translational Medicine

Jean-Michel Poggi

Fonction : Auteur

Université Paris-Sud - Paris 11

Christine Tuleau-Malot

Fonction : Auteur
PersonId : 8956
IdHAL : christine-malot
IdRef : 194152928

Université de Nice Sophia-Antipolis

Résumé

Based on decision trees combined with aggregation and bootstrap ideas, random forests were introduced by Breiman in 2001. They are a powerful nonparametric statistical method allowing to consider in a single and versatile framework regression problems, as well as two-class and multi-class classification problems. Focusing on classification problems, this paper reviews available proposals about random forests in parallel environments as well as about online random forests. Then, we formulate various remarks for random forests in the Big Data context. Finally, we experiment three variants involving subsampling, Big Data-bootstrap and MapReduce respectively, on two massive datasets (15 and 120 millions of observations), a simulated one as well as real world data.

Mots clés

Big data Random forests

Domaines

Machine Learning [stat.ML]

Migration ProdInra : Connectez-vous pour contacter le contributeur

https://hal.inrae.fr/hal-02796431

Soumis le : vendredi 5 juin 2020-13:40:45

Dernière modification le : mardi 12 mars 2024-10:44:42

Dates et versions

hal-02796431 , version 1 (05-06-2020)

Identifiants

HAL Id : hal-02796431 , version 1
PRODINRA : 374770

Citer

Nathalie N. Villa-Vialaneix, Robin Genuer, Jean-Michel Poggi, Christine Tuleau-Malot. Random forests for big data. Journées de Statistique de Rennes (JSTAR), Oct 2016, Rennes, France. ⟨hal-02796431⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

INRIA INRA INRIA2 UNIV-PARIS-SACLAY UNIV-COTEDAZUR INRAE U1219 INRAEOCCITANIETOULOUSE MATHNUM MIAT

27 Consultations

0 Téléchargements

Random forests for big data

Résumé

Mots clés

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Partager