Skip to content

About

Supplementary material for the Machine Learning module final project: Heart Disease Risk Prediction: An End-to-End Machine Learning Analysis

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

Heart Disease Risk Prediction: An End-to-End Machine Learning Analysis

Python XGBoost scikit-learn Jupyter License: MIT Grade

A supervised-learning project that stratifies patients into Low, Medium, and High cardiovascular risk categories from routine health metrics and lifestyle factors. The notebook covers the full workflow, from data cleaning and target-leakage prevention through exploratory data analysis, modelling with XGBoost, and model interpretation, reaching 90.55% test accuracy.

Completed as the final assignment for a Machine Learning module and graded 9.2 / 10. It is a personal and academic project, not intended for professional or clinical use.

Contents

Highlights

  • Runs end to end, from raw CSV to a trained and interpreted classifier in a single documented notebook.
  • Identifies and removes a target-leakage column (heart_disease_risk_score) before modelling, and explains each preprocessing decision.
  • Covers univariate, bivariate, and multivariate EDA across 14 clinical and lifestyle features.
  • Cross-checks XGBoost feature importance against the correlation analysis and against established cardiovascular risk factors.
  • Reports per-class metrics, a confusion matrix, an overfitting check, and a 2D PCA projection of class separability.

Dataset

data/cardiovascular_risk_dataset.csv holds 5,500 patient records across 17 columns.

The data is already clean, with no missing values and no duplicate rows. Two columns are dropped before modelling:

Dropped column Reason
Patient_ID Identifier with no predictive value
heart_disease_risk_score risk_category is derived from this score by thresholding, so including it would leak the target

That leaves 14 predictive features and a 3-class target.

Feature dictionary (click to expand)
Group Features
Clinical metrics age, bmi, systolic_bp, diastolic_bp, cholesterol_mg_dl, resting_heart_rate
Lifestyle factors daily_steps, stress_level, physical_activity_hours_per_week, sleep_hours, diet_quality_score, alcohol_units_per_week
Categorical smoking_status (Never, Former, Current), family_history_heart_disease (No, Yes)
Target risk_category, with classes Low (33.4%), Medium (40.8%), High (25.8%)

Workflow

  1. Data audit: shape, dtypes, summary statistics, and sanity ranges.
  2. Data quality: missing-value and duplicate checks.
  3. Leakage control: drop Patient_ID and heart_disease_risk_score.
  4. Encoding: integer-encode the categorical variables.
  5. EDA: univariate, bivariate, and multivariate.
  6. Dimensionality: a reasoned decision on whether to apply PCA.
  7. Split: 80/20 stratified train/test (random_state=12).
  8. Model: XGBoost multi-class classifier.
  9. Evaluation: accuracy, classification report, confusion matrix, feature importance, and PCA projection.

Figures

A few of the analysis figures are shown below. All are reproducible from the notebook, and full-resolution versions live in figures/.

Feature correlation with the target. Clinical measurements dominate the predictive signal, while lifestyle factors show protective (negative) correlations.

Feature correlation with risk category

Health metrics by risk category. Blood-pressure measurements separate the classes most cleanly.

Health metrics by risk category (violin plots)

XGBoost feature importance. The model relies on medically established risk factors, which supports its validity.

XGBoost feature importance

Confusion matrix. Most errors fall between adjacent risk levels rather than being gross misclassifications.

Confusion matrix, counts and normalised

Two-dimensional PCA projection. The classes form a visible risk gradient but overlap at the boundaries, which is why the model operates in the full 14-dimensional space.

PCA 2D projection of risk categories

Results

The XGBoost classifier (n_estimators=100, max_depth=6, learning_rate=0.1) reached 99.64% training accuracy and 90.55% test accuracy, a 9.09% gap that indicates moderate overfitting while still generalising reasonably.

Per-class performance on the test set (n = 1,100):

Class Precision Recall F1-score Support
Low Risk 0.936 0.916 0.926 367
Medium Risk 0.868 0.907 0.887 449
High Risk 0.930 0.891 0.910 284
Macro avg 0.911 0.904 0.908 1,100
Weighted avg 0.907 0.906 0.906 1,100

The model is most reliable on the clearly low- and high-risk patients. The Medium class is the hardest, reflecting overlap at the boundaries between adjacent risk levels.

Key Insights

  • Systolic blood pressure, cholesterol, and family history are the strongest drivers of predicted risk, consistent with established cardiovascular medicine.
  • Diastolic blood pressure is nearly redundant given systolic BP: it correlates with the target but the model rarely splits on it, so the two could be merged into a single blood-pressure feature.
  • Lifestyle factors are modifiable. Diet quality, physical activity, and daily steps all show protective (negative) associations with risk, unlike age or genetics.
  • Errors are mostly off by one. Misclassifications occur between neighbouring risk levels (Low and Medium, Medium and High), so the model captures the overall risk gradient even when it misses the exact class.

Future Work and Improvements

Modelling next steps, already noted in the notebook:

  • Hyperparameter tuning with grid or randomised search over max_depth, learning_rate, and n_estimators.
  • Cross-validation for more robust performance estimates than a single split.
  • Adjusting the decision threshold where one error type (false negative versus false positive) carries more clinical cost.
  • Ensembling with other algorithms such as Random Forest or Logistic Regression, and external validation on independent patient cohorts.
  • Feature engineering, for example combining systolic_bp and diastolic_bp into a single blood-pressure feature.

Extending the multivariate analysis, based on examiner feedback:

The notebook uses PCA only as a visualisation tool and keeps the original, interpretable features for the XGBoost model. The analysis of relationships between variables could be extended by:

  • Projecting the variables, not just the samples, into a PCA loading plot (biplot) or via MDS, to see which features group together in reduced space.
  • Adding a hierarchical clustering dendrogram of the features to make their relationships and redundancies explicit. The two blood-pressure measures, for instance, would be expected to cluster tightly, confirming the redundancy noted above.
  • Treating PCA as fully applicable here. Even where it is not used in the final modelling pipeline, it is a reasonable way to understand correlation structure and class separability.

Repository Structure

Heart-Disease-Risk-Prediction-Machine-Learning-Analysis/
├── HeartAttackRisk_ML_Cobos.ipynb        # Full analysis notebook
├── data/
│   └── cardiovascular_risk_dataset.csv   # 5,500 patient records
├── figures/                              # Exported analysis figures (PNG, 300 dpi)
├── requirements.txt                      # Python dependencies
├── LICENSE                               # MIT
└── README.md

Getting Started

Clone the repository:

git clone https://github.com/Cobos-Bioinfo/Heart-Disease-Risk-Prediction-Machine-Learning-Analysis.git
cd Heart-Disease-Risk-Prediction-Machine-Learning-Analysis

Create a virtual environment and install the dependencies:

python -m venv .venv
source .venv/bin/activate        # Windows: .venv\Scripts\activate
pip install -r requirements.txt

Launch the notebook and run the cells top to bottom. Figures are written to figures/ as they are generated.

jupyter lab HeartAttackRisk_ML_Cobos.ipynb

Tech Stack

Purpose Tools
Data manipulation pandas, numpy
Visualisation matplotlib, seaborn
Modelling xgboost, scikit-learn
Environment Jupyter / JupyterLab, Python 3.11+

Author

Alejandro Cobos. Machine Learning module final assignment, February 2026, graded 9.2 / 10.

This was my first hands-on project in big-data analysis and machine learning. The dataset was chosen for its clean, structured form so the focus could stay on the core principles of EDA and modelling, and on working through each step consciously.

License

Released under the MIT License.

About

Supplementary material for the Machine Learning module final project: Heart Disease Risk Prediction: An End-to-End Machine Learning Analysis

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages