Home Up PDF Prof. Dr. Ingo Claßen
Common Machine Learning Algorithms - DSML

Linear Regression

  • Fits a "best" straight line (or hyperplane) through data points
  • Simple

Logistic Regression

  • Used for classification, despite its name
  • Models probability of classes (softmax)
  • Prone to underfitting on complex data.

Decision Trees

  • Can be used for Classification & Regression
  • Splits data into branches based on feature questions
  • E.g., "Is Age > 30?"
  • Greedy algorithm makes locally optimal choices at each step
  • Performance can be poor with high-dimensional data
  • Prone to overfitting

Random Forests

  • Can be used for Classification & Regression
  • Builds hundreds of slightly different decision trees
  • Averages their predictions for regression
  • Majority vote for classification
  • Highly accurate
  • Requires little data preprocessing

Naive Bayes

How it works

  • Used for Classification
  • Based on Bayes' Theorem
  • Calculates the probability of a data point belonging to a certain class
  • Assumes features are conditionally independent given the class
  • Works remarkably well with high-dimensional text data
  • Independent features is rarely true in the real world

Support Vector Machines (SVM)

  • Can be used for Classification & Regression
  • Separates different classes by finding the optimal hyperplane
  • Support vectors are the data points that are closest to the hyperplane
  • Maximizes the "margin" (distance) between the hyperplane and the support vectors
  • Kernel Trick: Transform data into higher dimensions
  • Uses the Kernel Trick to find linear separators for non-linear data
  • Effective in high-dimensional spaces
  • Computationally heavy, only suitable for relatively small datasets
  • Requires careful tuning of parameters
  • Requires additional work for multiple classes

k-Nearest Neighbors (k-NN)

  • Can be used for Classification & Regression
  • Doesn't build a model
  • Looks at the k nearest neighbors to make predictions
  • Averages the values of the k nearest neighbors for regression
  • Majority voting for classification
  • Very intuitive
  • Simple to implement
  • Naturally handles multi-class problems
  • Very slow during the prediction phase
  • highly sensitive to feature scaling and irrelevant features

Gradient Boosting (e.g., XGBoost, LightGBM)

  • Can be used for Classification & Regression
  • Builds decision trees sequentially
  • Each new tree is trained on errors made by the previous trees
  • Combines the predictions of all the trees to make the final prediction
  • Often the most accurate algorithm for structured/tabular data
  • Prone to overfitting
  • Requires careful hyperparameter tuning
  • Slower to train than Random Forests

Linear vs. Non-Linear Models

Concept Linear Nonlinear
Relationship Straight/simple Curved/complex
Feature effect Generally constant Can be more complex
Decision boundary Hyperplane Can be curved/complex
Can model interactions naturally? Not without adding them Often yes
Flexibility Lower Higher
Interpretability Often easier Often harder

Sensitivity to outliers

How strongly an extreme observation can affect the model.

Model Sensitivity Reason
Linear Regression High Squared error gives extreme observations very large influence
Logistic Regression Moderate High-leverage observations can strongly affect coefficients
Decision Trees Low Splits are based on thresholds; individual extreme values often have limited effect
Random Forests Low Averaging many trees reduces the influence of individual unusual observations
Naive Bayes Moderate–High Depends heavily on the distribution assumption; extreme values can distort estimated distributions
SVM Moderate–High Points near/inside the margin are influential; feature scaling also matters
k-NN High Outliers can affect distances and therefore nearest-neighbor selection
Gradient Boosting Moderate–High Depends on the loss function; extreme observations can receive large residual/loss contributions

Pragmatic summary

  • Most sensitive: Linear Regression, k-NN
  • Moderately sensitive: Logistic Regression, Naive Bayes, SVM, Gradient Boosting
  • More robust: Decision Trees, Random Forests

Comparison of Algorithms

  • The following table provides a high-level comparison
  • The characteristics of each algorithm is very roughly described
  • Just to give a general idea of each algorithm's relative strengths and weaknesses
Algorithm Performance Interpretability Training Speed Shape
Linear Regression Lower Higher Faster Linear
Logistic Regression Lower Higher Faster Linear
Decision Trees Lower Higher Faster Nonlinear
Random Forests Higher Lower Slower Nonlinear
Naive Bayes Lower Higher Fast Nonlinear
SVM Higher Lower Slower Nonlinear
k-NN Lower Higher n/a Nonlinear
Gradient Boosting Higher Lower Slower Nonlinear

Quick Guide on Regression

Algorithm Consideration
Linear Regression data is roughly linear
you need interpretability
Decision Trees you need interpretability
your model should capture non-linear relationships
you don't need top performannce
Random Forests you want strong general-purpose performance
your model should be robust to outliers/noise
you don't want much tuning
SVM small-to-medium dataset
high-dimensional feature space
k-NN small dataset, low dimensionality
the relationship is local/non-parametric (no assumed functional form)
Gradient Boosting you want the best possible pereformance on structured/tabular data
you can afford careful tuning and longer training time

Quick Guide on Classification

Algorithm Consideration
Logistic Regression data is roughly linearly separable
you need interpretability
Decision Trees same considerations as for regression
Random Forests same considerations as for regression
Naive Bayes features are roughly independent
you're working with text
you're working with text high-dimensional sparse data
SVM you're working with text
otherwise same considerations as for regression
k-NN same considerations as for regression
Gradient Boosting same considerations as for regression