Machine learning essentials
The models I trained in the ML course — decision trees, random forests, gradient boosting, a neural network on Fashion-MNIST, PCA and clustering — plus how to evaluate and explain them.
updated 20 Jun 2026 · level beginner · 1 min read
From Machine Learning (FH JOANNEUM, summer 2026). Python, scikit-learn and Jupyter notebooks, with the environment managed by uv.
Workflow
- Split the data before you look at it: train / validation / test. The test set is used once, at the end.
- Start with a simple baseline, then improve.
- Pick a metric that fits the problem: accuracy only works for balanced classes. Otherwise use precision, recall or F1.
Models
- Decision tree: easy to read, but overfits. Limit
max_depthandmin_samples_leaf. - Random forest: many trees on random subsets of rows and features, averaged. Robust and a good default.
- Gradient boosting: trees built one after another, each fixing the errors of the last. Often the strongest on tables, but needs tuning (learning rate, number of trees).
- Neural network (Fashion-MNIST): dense layers with ReLU and a softmax output. Normalise the pixels and watch the validation loss to stop before it overfits.
- PCA: compresses features into a few components, which helps for plotting and speed.
- k-means: groups unlabeled data. Choose k with the elbow method or the silhouette score.
Explainability
Feature importance, permutation importance and SHAP-style explanations show why a model decides something. That matters as soon as the model affects people.
For practice I generated my own dataset (Tokyo ramen shops 🍜) and trained models on it before the quiz.