← Back to portfolio

Finance, Audit & Accounting Analytics

Config-driven data science pipeline for high-fidelity industrial domains.

Finance, Audit & Accounting Analytics Public dataset

At a glance

Task: classification  |  Headline: accuracy = 0.9848

accuracy
0.9848
balanced_accuracy
0.8856
precision_weighted
0.9842
recall_weighted
0.9848
f1_weighted
0.9842
precision_macro
0.9466
recall_macro
0.8856
f1_macro
0.9136

Summary

Project

Config-driven data science pipeline for high-fidelity industrial domains.

What was investigated

The engagement modelled the target 'is_anomalous' as a classification problem in the Finance, Audit & Accounting Analytics domain, using 5000 records across 12 columns (data source: Synthetic sandbox). The pipeline audited data quality, engineered features, split the data honestly into train/test, and compared several models by cross-validation.

Outcome

The selected model was RandomForest, with f1_macro = 0.9136 on the held-out test set. Supporting metrics: accuracy=0.985, balanced_accuracy=0.886, precision_weighted=0.984, recall_weighted=0.985, f1_weighted=0.984, precision_macro=0.947, recall_macro=0.886, cohen_kappa=0.827, matthews_. Full results, the deployment gates and the audit trail are in the run's report.

Figures

02_distributions.png
02 distributions
03_correlation.png
03 correlation
04_vif.png
04 vif
05_target.png
05 target
10_confusion.png
10 confusion
12_leaderboard.png
12 leaderboard
13_feature_importance.png
13 feature importance
20_benford_law.png
20 benford law
21_anomaly_map.png
21 anomaly map

Artifacts

Data source: Synthetic sandbox

Download report (.docx)

Audit SHA-256: 2c5102b8f3993a7534baab8c730fe8a14e7f8baa2835e354fe7d36645d2b86b5