← All projects
022026 / Deployed / University coursework

ML Predictive Modeling Pipeline

Built a full machine-learning pipeline covering data exploration, preprocessing, supervised and unsupervised modeling, cross-validation and hyperparameter tuning.

PythonScikit-LearnPandasMatplotlibSeabornJupyter
80%Stroke recall achieved
4Customer clusters identified
0.84AUC score

The problem

Needed to derive actionable insights from complex datasets while demonstrating both supervised and unsupervised ML competency.

What I built

  1. 01

    Performed statistical data exploration, missing-value analysis and visualisations including histograms, box plots, scatter plots and heat maps.

  2. 02

    Pre-processed data through imputation, normalisation, standardisation and one-hot encoding.

  3. 03

    Trained and evaluated supervised algorithms (linear regression, decision trees, SVMs) using accuracy, precision, recall, F1 and RMSE.

  4. 04

    Applied k-means clustering, hierarchical clustering and PCA to identify patterns in unlabeled data.

  5. 05

    Improved robustness through k-fold cross-validation and hyperparameter tuning with grid search and Bayesian optimisation.

Project stages

No screenshots yet — add image URLs to this project’s images list to build the slideshow.