Algorithmic Bias and Variance
- In machine learning, bias and variance represent a trade-off in a model’s ability to generalize.
The Bias-Variance Trade-off
- High Bias: A model with high bias is too simple and makes strong assumptions about the data. This leads to underfitting, where the model performs poorly on both the training data and new, unseen data.
- High Variance: A model with high variance is overly complex and learns the training data too well. This leads to overfitting, where the model performs well on the training data but poorly on new, unseen data.
- The Goal: The goal of a supervised learning model is to achieve a balance between bias and variance. This is known as the bias-variance trade-off. A model with the optimal balance will generalize well to new, unseen data.
Model Evaluation Techniques
Basic Parameters
- Bias & Variance: Measures to understand if our model is too simplistic (high bias) or too complex (high variance).
- Overfitting & Underfitting: Indicators that our model is either too closely tailored to the training data or too general.
Resampling Methods
- Holdout Method (Train / Test Split): A basic approach to split the dataset into training and testing sets.
- Repeated Holdout: Running the holdout method multiple times to get a better estimate of model performance.
- Cross-validation: A technique for assessing how the results of a statistical analysis will generalize to an independent data set.
Hyperparameter Tuning
- Grid Search: An exhaustive search over a specified parameter grid.
- Random Search: A search over a specified parameter distribution.
Statistical Tests
- Model Comparison: Statistical tests to compare the performance of two models.
- Algorithm Comparison: Comparing different algorithms to find which performs best on the data.
Evaluation Metrics
- Metrics to quantify the performance of the model.