At a glance
Task: classification | Headline: accuracy = 0.9848
Summary
Project
Config-driven data science pipeline for high-fidelity industrial domains.
What was investigated
The engagement modelled the target 'is_anomalous' as a classification problem in the Finance, Audit & Accounting Analytics domain, using 5000 records across 12 columns (data source: Synthetic sandbox). The pipeline audited data quality, engineered features, split the data honestly into train/test, and compared several models by cross-validation.
Outcome
The selected model was RandomForest, with f1_macro = 0.9136 on the held-out test set. Supporting metrics: accuracy=0.985, balanced_accuracy=0.886, precision_weighted=0.984, recall_weighted=0.985, f1_weighted=0.984, precision_macro=0.947, recall_macro=0.886, cohen_kappa=0.827, matthews_. Full results, the deployment gates and the audit trail are in the run's report.
Figures









Artifacts
Data source: Synthetic sandbox
Download report (.docx)Audit SHA-256: 2c5102b8f3993a7534baab8c730fe8a14e7f8baa2835e354fe7d36645d2b86b5