← Back to portfolio

Generic / Auto-detect (any tabular dataset)

Config-driven data science pipeline for high-fidelity industrial domains.

Generic / Auto-detect (any tabular dataset) Public dataset

At a glance

Task: regression  |  Headline: r2 = 0.8956

rmse
2.957
mae
2.3375
median_ae
1.933
mape
7.7107
r2
0.8956
adj_r2
0.8927
explained_variance
0.8959
mean_residual
0.1412

Summary

Project

Config-driven data science pipeline for high-fidelity industrial domains.

What was investigated

The engagement modelled the target 'target' as a regression problem in the Generic / Auto-detect (any tabular dataset) domain, using 2000 records across 7 columns (data source: Synthetic sandbox). The pipeline audited data quality, engineered features, split the data honestly into train/test, and compared several models by cross-validation.

Outcome

The selected model was Ridge, with r2 = 0.8956 on the held-out test set. Supporting metrics: rmse=2.957, mae=2.338, median_ae=1.933, mape=7.711, adj_r2=0.893, explained_variance=0.896, mean_residual=0.141. Full results, the deployment gates and the audit trail are in the run's report.

Figures

02_distributions.png
02 distributions
03_correlation.png
03 correlation
04_vif.png
04 vif
05_target.png
05 target
10_pred_vs_actual.png
10 pred vs actual
11_residuals.png
11 residuals
12_leaderboard.png
12 leaderboard
13_feature_importance.png
13 feature importance

Artifacts

Data source: Synthetic sandbox

Download report (.docx)

Audit SHA-256: 5a7e09f768a5ed155dd414aba86f4770fe20feb73975e236e04f08c2b7080d76