← Back to portfolio

AI Usage and Impact

Config-driven data science pipeline for high-fidelity industrial domains.

Generic / Auto-detect (any tabular dataset) Public dataset

At a glance

Task: classification  |  Headline: accuracy = 0.6102

accuracy
0.6102
balanced_accuracy
0.5143
precision_weighted
0.601
recall_weighted
0.6102
f1_weighted
0.6025
precision_macro
0.5417
recall_macro
0.5143
f1_macro
0.5247

Summary

Project

Config-driven data science pipeline for high-fidelity industrial domains.

What was investigated

The engagement modelled the target 'Would_Recommend' as a classification problem in the Generic / Auto-detect (any tabular dataset) domain, using 300 records across 19 columns (data source: AI_Usage_and_Impact_on_Students_and_Professionals.csv). The pipeline audited data quality, engineered features, split the data honestly into train/test, and compared several models by cross-validation.

Outcome

The selected model was XGBoost, with f1_macro = 0.5247 on the held-out test set. Supporting metrics: accuracy=0.610, balanced_accuracy=0.514, precision_weighted=0.601, recall_weighted=0.610, f1_weighted=0.603, precision_macro=0.542, recall_macro=0.514, cohen_kappa=0.272, matthews_. Full results, the deployment gates and the audit trail are in the run's report.

Figures

01_missingness.png
01 missingness
02_distributions.png
02 distributions
03_correlation.png
03 correlation
04_vif.png
04 vif
05_target.png
05 target
10_confusion.png
10 confusion
12_leaderboard.png
12 leaderboard
13_feature_importance.png
13 feature importance

Artifacts

Data source: AI_Usage_and_Impact_on_Students_and_Professionals.csv

Download report (.docx)

Audit SHA-256: f44a8c14dcc8ac0db8b1712ec46e421c88edca23e3873538c22c80b0b8e317ce