← Back to portfolio

ISM + Security: AI Governance & Data Provenance

Config-driven data science pipeline for high-fidelity industrial domains.

ISM + Security: AI Governance & Data Provenance Public dataset

At a glance

Task: classification  |  Headline: accuracy = 1.0

accuracy
1.0
balanced_accuracy
1.0
precision_weighted
1.0
recall_weighted
1.0
f1_weighted
1.0
precision_macro
1.0
recall_macro
1.0
f1_macro
1.0

Summary

Project

Config-driven data science pipeline for high-fidelity industrial domains.

What was investigated

The engagement modelled the target 'is_compliant' as a classification problem in the ISM + Security: AI Governance & Data Provenance domain, using 2000 records across 23 columns (data source: Synthetic sandbox). The pipeline audited data quality, engineered features, split the data honestly into train/test, and compared several models by cross-validation.

Outcome

The selected model was LogisticRegression, with f1_macro = 1.0000 on the held-out test set. Supporting metrics: accuracy=1.000, balanced_accuracy=1.000, precision_weighted=1.000, recall_weighted=1.000, f1_weighted=1.000, precision_macro=1.000, recall_macro=1.000, cohen_kappa=1.000, matthews_. Full results, the deployment gates and the audit trail are in the run's report.

Figures

02_distributions.png
02 distributions
03_correlation.png
03 correlation
04_vif.png
04 vif
05_target.png
05 target
10_confusion.png
10 confusion
12_leaderboard.png
12 leaderboard
13_feature_importance.png
13 feature importance
20_governance_risk_matrix.png
20 governance risk matrix
21_data_lineage_graph.png
21 data lineage graph

Artifacts

Data source: Synthetic sandbox

Download report (.docx)

Audit SHA-256: 678d40bca5df6feefb4dbb6cc4b35ef30d887b66c619f6cf9606a8265bd6f534