A random forest is an ensemble learning method that constructs many decision trees and aggregates their predictions, typically by majority vote for classification or averaging for regression. Each tree is trained on a bootstrap sample of the data and considers a random subset of features at each split, which decorrelates the trees and reduces variance. The resulting model is robust, resistant to overfitting, and provides built-in estimates of feature importance.
Overview
- Random forests combine bootstrap aggregation (bagging) with random feature selection to build a diverse collection of decision trees whose errors are largely uncorrelated. Averaging the trees cancels much of the variance inherent in a single deep tree, yielding strong out-of-the-box accuracy with little tuning. The out-of-bag samples left out of each bootstrap provide an unbiased internal estimate of generalisation error.
Key aspects
- Bootstrap aggregation: each tree trained on a resampled subset of the data
- Random feature subsetting at each split to decorrelate trees
- Majority voting or averaging across the ensemble
- Out-of-bag error estimation without a separate validation set
- Permutation-based and impurity-based feature importance measures
Applications
- Tabular classification and regression across science and industry
- Feature selection and ranking in high-dimensional datasets
- Credit scoring and risk modelling
- Baseline models against which deeper architectures are compared