How it works
| Concept | Linear | Nonlinear |
|---|---|---|
| Relationship | Straight/simple | Curved/complex |
| Feature effect | Generally constant | Can be more complex |
| Decision boundary | Hyperplane | Can be curved/complex |
| Can model interactions naturally? | Not without adding them | Often yes |
| Flexibility | Lower | Higher |
| Interpretability | Often easier | Often harder |
How strongly an extreme observation can affect the model.
| Model | Sensitivity | Reason |
|---|---|---|
| Linear Regression | High | Squared error gives extreme observations very large influence |
| Logistic Regression | Moderate | High-leverage observations can strongly affect coefficients |
| Decision Trees | Low | Splits are based on thresholds; individual extreme values often have limited effect |
| Random Forests | Low | Averaging many trees reduces the influence of individual unusual observations |
| Naive Bayes | Moderate–High | Depends heavily on the distribution assumption; extreme values can distort estimated distributions |
| SVM | Moderate–High | Points near/inside the margin are influential; feature scaling also matters |
| k-NN | High | Outliers can affect distances and therefore nearest-neighbor selection |
| Gradient Boosting | Moderate–High | Depends on the loss function; extreme observations can receive large residual/loss contributions |
Pragmatic summary
| Algorithm | Performance | Interpretability | Training Speed | Shape |
|---|---|---|---|---|
| Linear Regression | Lower | Higher | Faster | Linear |
| Logistic Regression | Lower | Higher | Faster | Linear |
| Decision Trees | Lower | Higher | Faster | Nonlinear |
| Random Forests | Higher | Lower | Slower | Nonlinear |
| Naive Bayes | Lower | Higher | Fast | Nonlinear |
| SVM | Higher | Lower | Slower | Nonlinear |
| k-NN | Lower | Higher | n/a | Nonlinear |
| Gradient Boosting | Higher | Lower | Slower | Nonlinear |
| Algorithm | Consideration |
|---|---|
| Linear Regression |
data is roughly linear you need interpretability |
| Decision Trees |
you need interpretability your model should capture non-linear relationships you don't need top performannce |
| Random Forests |
you want strong general-purpose performance your model should be robust to outliers/noise you don't want much tuning |
| SVM |
small-to-medium dataset high-dimensional feature space |
| k-NN |
small dataset, low dimensionality the relationship is local/non-parametric (no assumed functional form) |
| Gradient Boosting |
you want the best possible pereformance on structured/tabular data you can afford careful tuning and longer training time |
| Algorithm | Consideration |
|---|---|
| Logistic Regression |
data is roughly linearly separable you need interpretability |
| Decision Trees | same considerations as for regression |
| Random Forests | same considerations as for regression |
| Naive Bayes |
features are roughly independent you're working with text you're working with text high-dimensional sparse data |
| SVM |
you're working with text otherwise same considerations as for regression |
| k-NN | same considerations as for regression |
| Gradient Boosting | same considerations as for regression |