Summary
Transaction monitoring sits at the intersection of financial crime and model risk. Whether the system is rule-based or uses machine learning, validation follows the same structural discipline: conceptual soundness, data, methodology, performance and governance.
Conceptual soundness
Each scenario should be traceable to a typology in the EWRA. Scenarios without a documented typology rationale are candidates for retirement.
Data
Completeness and lineage of transaction, customer and reference data are the foundation. A validation that does not test data has not validated the model.
Methodology and thresholds
Threshold-setting should be evidence-based: above-the-line and below-the-line testing, productivity analysis, and disposition quality. Round-number thresholds without a rationale are a common finding.
Performance
Effectiveness is measured by alert quality and by coverage of known typologies. False-positive rates alone are misleading; a low false-positive rate can indicate over-tuning.
Tuning governance
Tuning changes should be documented, tested and approved through a defined governance path, with a clear audit trail.
Machine-learning approaches
ML-based monitoring introduces additional validation requirements: feature stability, drift monitoring, and explainability sufficient to defend alerts to investigators and to supervisors.
Limitations
Validation does not certify that no financial crime passes undetected. It provides assurance that the system is fit for its stated purpose and operates within known limits.
Related expertise
Frequently asked questions
What should risk leaders know about conceptual soundness?
Each scenario should be traceable to a typology in the EWRA. Scenarios without a documented typology rationale are candidates for retirement.
What should risk leaders know about data?
Completeness and lineage of transaction, customer and reference data are the foundation. A validation that does not test data has not validated the model.
What should risk leaders know about methodology and thresholds?
Threshold-setting should be evidence-based: above-the-line and below-the-line testing, productivity analysis, and disposition quality. Round-number thresholds without a rationale are a common finding.
What should risk leaders know about performance?
Effectiveness is measured by alert quality and by coverage of known typologies. False-positive rates alone are misleading; a low false-positive rate can indicate over-tuning.
What should risk leaders know about tuning governance?
Tuning changes should be documented, tested and approved through a defined governance path, with a clear audit trail.