A supervised-learning project that stratifies patients into Low, Medium, and High cardiovascular risk categories from routine health metrics and lifestyle factors. The notebook covers the full workflow, from data cleaning and target-leakage prevention through exploratory data analysis, modelling with XGBoost, and model interpretation, reaching 90.55% test accuracy.
Completed as the final assignment for a Machine Learning module and graded 9.2 / 10. It is a personal and academic project, not intended for professional or clinical use.
- Highlights
- Dataset
- Workflow
- Figures
- Results
- Key Insights
- Future Work and Improvements
- Repository Structure
- Getting Started
- Tech Stack
- Author
- License
- Runs end to end, from raw CSV to a trained and interpreted classifier in a single documented notebook.
- Identifies and removes a target-leakage column (
heart_disease_risk_score) before modelling, and explains each preprocessing decision. - Covers univariate, bivariate, and multivariate EDA across 14 clinical and lifestyle features.
- Cross-checks XGBoost feature importance against the correlation analysis and against established cardiovascular risk factors.
- Reports per-class metrics, a confusion matrix, an overfitting check, and a 2D PCA projection of class separability.
data/cardiovascular_risk_dataset.csv holds 5,500 patient records across 17 columns.
The data is already clean, with no missing values and no duplicate rows. Two columns are dropped before modelling:
| Dropped column | Reason |
|---|---|
Patient_ID |
Identifier with no predictive value |
heart_disease_risk_score |
risk_category is derived from this score by thresholding, so including it would leak the target |
That leaves 14 predictive features and a 3-class target.
Feature dictionary (click to expand)
| Group | Features |
|---|---|
| Clinical metrics | age, bmi, systolic_bp, diastolic_bp, cholesterol_mg_dl, resting_heart_rate |
| Lifestyle factors | daily_steps, stress_level, physical_activity_hours_per_week, sleep_hours, diet_quality_score, alcohol_units_per_week |
| Categorical | smoking_status (Never, Former, Current), family_history_heart_disease (No, Yes) |
| Target | risk_category, with classes Low (33.4%), Medium (40.8%), High (25.8%) |
- Data audit: shape, dtypes, summary statistics, and sanity ranges.
- Data quality: missing-value and duplicate checks.
- Leakage control: drop
Patient_IDandheart_disease_risk_score. - Encoding: integer-encode the categorical variables.
- EDA: univariate, bivariate, and multivariate.
- Dimensionality: a reasoned decision on whether to apply PCA.
- Split: 80/20 stratified train/test (
random_state=12). - Model: XGBoost multi-class classifier.
- Evaluation: accuracy, classification report, confusion matrix, feature importance, and PCA projection.
A few of the analysis figures are shown below. All are reproducible from the notebook, and full-resolution versions live in figures/.
Feature correlation with the target. Clinical measurements dominate the predictive signal, while lifestyle factors show protective (negative) correlations.
Health metrics by risk category. Blood-pressure measurements separate the classes most cleanly.
XGBoost feature importance. The model relies on medically established risk factors, which supports its validity.
Confusion matrix. Most errors fall between adjacent risk levels rather than being gross misclassifications.
Two-dimensional PCA projection. The classes form a visible risk gradient but overlap at the boundaries, which is why the model operates in the full 14-dimensional space.
The XGBoost classifier (n_estimators=100, max_depth=6, learning_rate=0.1) reached 99.64% training accuracy and 90.55% test accuracy, a 9.09% gap that indicates moderate overfitting while still generalising reasonably.
Per-class performance on the test set (n = 1,100):
| Class | Precision | Recall | F1-score | Support |
|---|---|---|---|---|
| Low Risk | 0.936 | 0.916 | 0.926 | 367 |
| Medium Risk | 0.868 | 0.907 | 0.887 | 449 |
| High Risk | 0.930 | 0.891 | 0.910 | 284 |
| Macro avg | 0.911 | 0.904 | 0.908 | 1,100 |
| Weighted avg | 0.907 | 0.906 | 0.906 | 1,100 |
The model is most reliable on the clearly low- and high-risk patients. The Medium class is the hardest, reflecting overlap at the boundaries between adjacent risk levels.
- Systolic blood pressure, cholesterol, and family history are the strongest drivers of predicted risk, consistent with established cardiovascular medicine.
- Diastolic blood pressure is nearly redundant given systolic BP: it correlates with the target but the model rarely splits on it, so the two could be merged into a single blood-pressure feature.
- Lifestyle factors are modifiable. Diet quality, physical activity, and daily steps all show protective (negative) associations with risk, unlike age or genetics.
- Errors are mostly off by one. Misclassifications occur between neighbouring risk levels (Low and Medium, Medium and High), so the model captures the overall risk gradient even when it misses the exact class.
Modelling next steps, already noted in the notebook:
- Hyperparameter tuning with grid or randomised search over
max_depth,learning_rate, andn_estimators. - Cross-validation for more robust performance estimates than a single split.
- Adjusting the decision threshold where one error type (false negative versus false positive) carries more clinical cost.
- Ensembling with other algorithms such as Random Forest or Logistic Regression, and external validation on independent patient cohorts.
- Feature engineering, for example combining
systolic_bpanddiastolic_bpinto a single blood-pressure feature.
Extending the multivariate analysis, based on examiner feedback:
The notebook uses PCA only as a visualisation tool and keeps the original, interpretable features for the XGBoost model. The analysis of relationships between variables could be extended by:
- Projecting the variables, not just the samples, into a PCA loading plot (biplot) or via MDS, to see which features group together in reduced space.
- Adding a hierarchical clustering dendrogram of the features to make their relationships and redundancies explicit. The two blood-pressure measures, for instance, would be expected to cluster tightly, confirming the redundancy noted above.
- Treating PCA as fully applicable here. Even where it is not used in the final modelling pipeline, it is a reasonable way to understand correlation structure and class separability.
Heart-Disease-Risk-Prediction-Machine-Learning-Analysis/
├── HeartAttackRisk_ML_Cobos.ipynb # Full analysis notebook
├── data/
│ └── cardiovascular_risk_dataset.csv # 5,500 patient records
├── figures/ # Exported analysis figures (PNG, 300 dpi)
├── requirements.txt # Python dependencies
├── LICENSE # MIT
└── README.md
Clone the repository:
git clone https://github.com/Cobos-Bioinfo/Heart-Disease-Risk-Prediction-Machine-Learning-Analysis.git
cd Heart-Disease-Risk-Prediction-Machine-Learning-AnalysisCreate a virtual environment and install the dependencies:
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txtLaunch the notebook and run the cells top to bottom. Figures are written to figures/ as they are generated.
jupyter lab HeartAttackRisk_ML_Cobos.ipynb| Purpose | Tools |
|---|---|
| Data manipulation | pandas, numpy |
| Visualisation | matplotlib, seaborn |
| Modelling | xgboost, scikit-learn |
| Environment | Jupyter / JupyterLab, Python 3.11+ |
Alejandro Cobos. Machine Learning module final assignment, February 2026, graded 9.2 / 10.
This was my first hands-on project in big-data analysis and machine learning. The dataset was chosen for its clean, structured form so the focus could stay on the core principles of EDA and modelling, and on working through each step consciously.
Released under the MIT License.




