← Back to portfolio

Mechanical Engineering: Patent Novelty Graph (NLP)

Config-driven data science pipeline for high-fidelity industrial domains.

Mechanical Engineering: Patent Novelty Graph (NLP) Public dataset

At a glance

Task: regression  |  Headline: r2 = 0.6922

rmse
0.0502
mae
0.0403
median_ae
0.0347
mape
19.5361
r2
0.6922
adj_r2
0.6429
explained_variance
0.6933
mean_residual
0.0031

Summary

Project

Config-driven data science pipeline for high-fidelity industrial domains.

What was investigated

The engagement modelled the target 'novelty_index' as a regression problem in the Mechanical Engineering: Patent Novelty Graph (NLP) domain, using 2000 records across 16 columns (data source: Synthetic sandbox). The pipeline audited data quality, engineered features, split the data honestly into train/test, and compared several models by cross-validation.

Outcome

The selected model was Ridge, with r2 = 0.6922 on the held-out test set. Supporting metrics: rmse=0.050, mae=0.040, median_ae=0.035, mape=19.536, adj_r2=0.643, explained_variance=0.693, mean_residual=0.003. Full results, the deployment gates and the audit trail are in the run's report.

Figures

02_distributions.png
02 distributions
03_correlation.png
03 correlation
04_vif.png
04 vif
05_target.png
05 target
10_pred_vs_actual.png
10 pred vs actual
11_residuals.png
11 residuals
12_leaderboard.png
12 leaderboard
13_feature_importance.png
13 feature importance
20_patent_citation_graph.png
20 patent citation graph
21_semantic_density.png
21 semantic density

Artifacts

Data source: Synthetic sandbox

Download report (.docx)

Audit SHA-256: e965c689fa884721f2a9894a8fd3c36b2e57ec03788b956f32c8801c660c90c8