At a glance
Task: regression | Headline: r2 = 0.6922
Summary
Project
Config-driven data science pipeline for high-fidelity industrial domains.
What was investigated
The engagement modelled the target 'novelty_index' as a regression problem in the Mechanical Engineering: Patent Novelty Graph (NLP) domain, using 2000 records across 16 columns (data source: Synthetic sandbox). The pipeline audited data quality, engineered features, split the data honestly into train/test, and compared several models by cross-validation.
Outcome
The selected model was Ridge, with r2 = 0.6922 on the held-out test set. Supporting metrics: rmse=0.050, mae=0.040, median_ae=0.035, mape=19.536, adj_r2=0.643, explained_variance=0.693, mean_residual=0.003. Full results, the deployment gates and the audit trail are in the run's report.
Figures










Artifacts
Data source: Synthetic sandbox
Download report (.docx)Audit SHA-256: e965c689fa884721f2a9894a8fd3c36b2e57ec03788b956f32c8801c660c90c8