Vendor Monitoring Tool Governance
Banks often buy monitoring or screening technology because building every capability internally would be slow, expensive or difficult to maintain. A vendor may provide a transaction-monitoring engine, sanctions or name-screening software, behavioural analytics, machine-learning models, case-management components, reference data, cloud hosting, managed tuning, or some combination of these. Buying the technology, however, does not transfer the bank's accountability for whether the control works.
That distinction is the foundation of good vendor governance. The supplier can operate software, propose detection logic, release new models and provide service reports. The bank still has to understand what population is covered, which data reach the tool, which risks the logic is intended to detect, where the tool is weak, how changes are controlled, what happens during an outage, and how alert outputs connect to investigation and reporting decisions. A green vendor dashboard is not evidence that the bank is detecting the right activity.
A useful mental model is to treat a vendor monitoring tool as part of the bank's financial-crime control chain, not as a product sitting outside it. The control starts before the vendor platform receives data and ends after its output has been investigated and acted on. Failures can therefore occur in the bank's own source systems, integration layer, vendor engine, reference data, configuration, alert routing, case management or downstream decision process. Governance has to cover the whole chain.
This chapter focuses on transaction monitoring, suspicious-activity monitoring and screening technology used in financial-crime control. It does not assume that every vendor arrangement is legally an "outsourcing" arrangement in every jurisdiction. Regulatory definitions differ. It also does not assume that all vendor models are machine learning. A rules engine with fixed scenarios can create just as much control risk as an advanced model if its data, configuration and change process are poorly governed.
The regulatory idea: responsibility stays with the bank
The regulatory language differs by jurisdiction, but a common theme is that reliance on a third party does not remove the regulated firm's responsibility for its obligations. The Basel Committee's December 2025 Principles for the sound management of third-party risk establish a common baseline for banks and supervisors and use a full third-party life-cycle approach. The Financial Stability Board's 2023 toolkit similarly addresses identification of critical services, life-cycle risk management, concentration, incidents and systemic dependencies.
In the United States, the Federal Reserve, FDIC and OCC issued joint interagency guidance in June 2023 covering planning, due diligence and third-party selection, contract negotiation, ongoing monitoring and termination. The agencies are explicit that engaging a third party does not remove a banking organisation's responsibility to operate safely and soundly and comply with applicable requirements. The 2024 community-bank guide is useful operational reading, although it does not create a safe harbour or replace the interagency guidance.
In the European Union, an AML monitoring tool delivered as an ICT service may also fall within the Digital Operational Resilience Act, which has applied since 17 January 2025. DORA treats ICT third-party risk as part of the financial entity's ICT risk-management framework and requires firms to remain fully responsible for compliance. It also introduces requirements around registers of ICT contractual arrangements, critical or important functions, concentration risk, contractual terms, monitoring and exit. On 18 September 2026, the EBA published final Guidelines on third-party risk for non-ICT services; as of this chapter's review date, those final Guidelines were published but not yet applicable, and the existing EBA outsourcing framework remains relevant until the transition is completed. A bank should therefore check the current legal status rather than assume that one EBA document governs every arrangement.
For UK PRA-regulated firms, SS2/21 remains a key source for outsourcing and third-party risk management. It links third-party governance to operational resilience, data security, business continuity and exit planning. Other UK requirements can apply depending on the entity and service. The point is not to memorise one global checklist. It is to map the arrangement to the legal entities, regulators, service type, criticality and jurisdictions that actually apply.
Financial-crime effectiveness has an additional dimension beyond operational outsourcing rules. The Wolfsberg Group's 2024 and 2025 work on effective monitoring for suspicious activity emphasises outcomes, transition and validation, balancing model risk with financial-crime risk, and explainability. These are especially relevant where a vendor introduces analytics, machine learning or automated prioritisation. The bank needs evidence that innovation improves or at least preserves the control outcome; technology sophistication by itself proves nothing.
What exactly is the bank buying?
Governance starts by defining the service precisely. "AML platform" is too vague to control. Two banks can buy the same product but create very different risk profiles depending on configuration and architecture.
One implementation may send daily account transactions to a vendor-hosted rules engine while the bank owns scenarios, thresholds and case management. Another may use a vendor-managed SaaS platform where the supplier hosts data, maintains models, proposes tuning and provides the case workflow. A third may buy only screening algorithms while the bank operates the infrastructure. A fourth may use vendor reference data and name-matching technology but keep payment interdiction and analyst decisions in-house.
The service inventory should therefore identify at least: what business process the tool supports; which legal entities and products use it; whether it supports a critical or important function under applicable rules; where it is hosted; what customer and transaction data it processes; which subcontractors or cloud services are involved; who owns configuration; who can change production logic; who supplies lists or external data; how alerts enter bank workflow; and what the bank must still do manually.
This matters because responsibility follows the real control design, not the sales description. If the vendor provides a score but the bank defines the threshold that sends a case to investigation, the bank owns that threshold decision. If the vendor runs managed tuning but the bank approves scenario changes, the bank needs enough information to challenge the recommendation. If the vendor hosts the entire service but the bank's upstream integration omits a payment channel, the vendor may meet every contractual uptime measure while the bank's AML control remains incomplete.
Criticality and risk tiering
Not every supplier needs the same oversight. A public-data provider used for analyst research does not create the same dependency as a platform through which all retail payment monitoring runs. A proportionate programme should assess criticality before the contract is signed and refresh that assessment when the service changes.
The assessment should consider the impact of failure on legal or regulatory obligations, the volume and risk of customers or transactions covered, substitutability, data sensitivity, concentration, reliance on subcontractors, recovery time, manual fallback, geographic dependencies, and the difficulty of migrating data and configuration. A tool may be technically replaceable but operationally hard to exit because years of tuned scenarios, case history and analyst workflows are embedded in proprietary formats.
Criticality also needs a financial-crime lens. A screening service used only for a low-volume internal workflow may be non-critical from a business-service perspective but still legally sensitive if an outage could allow a prohibited payment to execute. Conversely, an analytics feature may support prioritisation rather than a legal interdiction point; its failure may be serious without producing the same immediate legal consequence. Governance should record the reasoning instead of forcing every tool into the same risk label.
Due diligence: test the service you intend to run
Vendor due diligence should go beyond security questionnaires and corporate financials. Those are important, but they do not prove that the monitoring capability fits the bank's risk profile.
A sound evaluation begins with a control-use case: which risks are in scope, what data are available, what decisions the tool will support, and what evidence the bank needs. The proof of concept should use data and scenarios representative of the bank's actual environment where legally and operationally possible. Synthetic or masked data can be useful for early testing, but final validation should demonstrate performance on realistic distributions, including difficult edge cases.
For a transaction-monitoring tool, due diligence should explore how rules or models handle customer segmentation, peer groups, velocity, linked accounts, cash, cards, instant payments, cross-border flows, correspondent activity and any specialist products in scope. For screening, the bank should understand name matching, transliteration, aliases, date-of-birth handling, identifiers, addresses, ownership or control data, list updates, payment-message fields and the difference between customer and transaction screening. The exact questions depend on the service; the goal is to expose assumptions before they become production dependencies.
A vendor's published benchmark is not enough. Benchmarks may use different populations, labels and objectives. The bank should ask how the benchmark was built, what constitutes a true positive, which data fields were available, how class imbalance was handled, whether results are stable by segment, and what limitations were observed. For machine learning, the bank should also understand features, training-data relevance, retraining, explainability, drift detection and the extent to which the supplier can reproduce a historical decision.
Commercial and operational diligence belongs in the same decision. A technically strong vendor can still create unacceptable risk if it is financially fragile, heavily dependent on one subprocessor, unable to support required recovery objectives, unwilling to provide audit evidence, or unable to export configuration and historical outputs at exit.
Contracts should make the control governable
A contract cannot make a weak tool effective, but poor contract terms can prevent the bank from governing a good tool. The required clauses depend on applicable law, criticality and service model, so legal and third-party-risk teams should translate regulatory requirements into the actual agreement.
For a material monitoring service, the bank normally needs clear service descriptions, data and security obligations, confidentiality, subprocessor rules, incident notification, business continuity, audit and information rights, access for relevant authorities where required, data-location terms where material, change notification, support and escalation, record retention, termination rights, transition assistance and data return or deletion. Where DORA applies, the contract should be checked against DORA's requirements rather than relying on a generic outsourcing template.
Financial-crime controls need additional specificity. The agreement should state who owns scenario or model configuration, how material changes are classified, what evidence accompanies a release, which test environments are available, how quickly list or reference-data defects are handled, whether the bank can obtain historical configuration, and how performance data are supplied. If the vendor will tune logic, the contract should preserve the bank's approval and challenge rights.
Service levels should not stop at platform availability. A monitoring platform can be 99.9% available while silently rejecting a class of records. Operational measures may therefore include input acceptance, processing completeness, maximum unprocessed queues, latency, list-update completion, rejected-record handling, alert delivery, incident response and recovery. Financial-crime effectiveness measures then sit above those operational SLAs: coverage against the bank's risk assessment, scenario or model performance, alert quality, missed-case analysis, stability by segment and remediation of known limitations. These measures are not all contractual penalties, but the bank needs them to govern the control.
Architecture: understand the control boundary
The vendor engine is only one component. A typical monitoring chain starts in core banking, payment hubs, card platforms, customer masters and external reference sources. Data are extracted, transformed, enriched and mapped before reaching the vendor. The vendor applies rules, models or matching logic. Results then move into an alert or case system where investigators decide what to do.
Every boundary can fail.
A payment may never be extracted. A product code may map to the wrong segment. A country field may be derived incorrectly. A late file may arrive after the daily run. An API may retry and create duplicates. The vendor may reject malformed records without raising a bank-visible incident. A model may process the record but use a stale feature. An alert may be generated but fail to enter the case queue. If the bank monitors only the vendor platform's internal health, all of those failures can remain invisible.
Architecture therefore needs reconciliation across control stages. The bank should be able to answer: how many source events were expected; how many were extracted; how many were accepted by the vendor; how many were evaluated by applicable logic; how many alerts were produced; and whether every alert reached the correct workflow. Counts are not sufficient on their own, because a feed can be numerically complete while important fields are blank or wrong. Field-level data-quality controls are needed for risk-relevant attributes.
Historical traceability also matters. When an investigator, auditor or supervisor asks why a transaction did not alert six months ago, the bank should be able to identify the data presented to the engine, the rule or model version, the threshold, the reference data, any suppression or exclusion, and the resulting decision. A vendor-managed service should not become a black box simply because the infrastructure is external.
Validation: separate technical operation from control effectiveness
A successful implementation test shows that the system runs. Validation asks whether it performs the intended financial-crime control.
The first validation layer is data and population coverage. Are all intended customers, accounts, transactions and relevant fields present? Are exclusions documented and approved? Do interfaces reconcile? Are historical and late-arriving records handled correctly?
The second layer is logic or model behaviour. For rules, this includes scenario conditions, thresholds, segmentation, aggregation windows, suppression logic and edge cases. For statistical or machine-learning models, it also includes feature integrity, training-data relevance, stability, explainability, performance by segment, sensitivity to drift and limitations.
The third layer is outcome effectiveness. Does the control identify activity relevant to the bank's actual risk? This can use labelled historical cases, typology testing, backtesting, targeted sampling, challenger logic and post-investigation outcomes. False positives matter because they consume capacity, but reducing alerts is not automatically an improvement. A model can become "efficient" by missing difficult risk. The trade-off has to be evaluated against financial-crime risk and operational capacity.
Validation independence should be proportionate to the risk and the bank's governance model. The people validating a material tool should have enough independence from vendor selection, implementation and day-to-day ownership to challenge assumptions. That does not mean validators operate without business expertise; it means the same commercial or delivery incentives should not determine the conclusion.
The Wolfsberg Group's 2025 statement on innovation in suspicious-activity monitoring is useful here because it explicitly connects transition and validation, model risk versus financial-crime risk, and explainability. A bank should avoid applying generic model-risk controls in a way that blocks beneficial financial-crime innovation, but it should also avoid using "financial-crime urgency" as a reason to skip evidence.
Go-live should be a controlled decision, not a project deadline
A monitoring tool should not enter production because integration testing finished or a licence renewal date is approaching. Go-live should require explicit evidence that the intended population is connected, data quality is acceptable, configuration is approved, test results meet defined criteria, operational teams are ready, incident and fallback procedures exist, and unresolved limitations are owned.
Parallel running can be valuable during replacement or material model change. Comparing old and new outputs helps reveal population gaps and unexpected shifts, but simple alert-count comparison is insufficient. The bank should explain why differences occur. A new model might legitimately produce fewer alerts because it removes low-value noise, or more alerts because it captures a typology the old control missed. Review should look at the composition and risk of the differences.
Where a material limitation is accepted for go-live, the decision should state the exposure, temporary compensating control, owner and expiry date. "Known limitation" should not become permanent vocabulary for a control gap that nobody is funded to close.
Ongoing oversight: monitor both service and effectiveness
Once live, vendor governance becomes a continuous control. A quarterly service review based only on ticket volumes and uptime is inadequate for a material AML platform.
Operational oversight should monitor availability, processing, rejected records, batch or API delays, capacity, reference-data updates, security events, support performance and unresolved defects. Financial-crime oversight should monitor alert distributions, investigation outcomes, scenario or model performance, segment behaviour, override patterns, suppression, unexplained shifts, typology coverage and evidence from missed cases.
The bank should also monitor the vendor itself. Material changes in ownership, financial condition, key subcontractors, hosting locations, product strategy or support model can change the risk of the arrangement. Concentration is particularly important where the same provider supports multiple controls or where many business units depend on a shared cloud, data or analytics service.
Metrics need interpretation. A falling alert rate may reflect better precision, a changing portfolio, a broken feed, a threshold change or a model that has drifted. A rising conversion rate may be positive, or it may mean analysts are seeing only the most obvious cases while medium-risk networks disappear from the queue. Management information should therefore combine technical, control and investigation evidence rather than rewarding one number.
Vendor changes and model updates
A major governance risk appears when a vendor changes the service after initial validation. SaaS providers may release frequently. Reference data change continuously. Matching algorithms are tuned. Machine-learning models may be retrained. Cloud components and subprocessors can change. If every change is treated as routine, the bank can lose track of the control it originally approved.
The contract and operating model should define which changes require notification, impact assessment, testing or approval. Materiality should consider the effect on customer or transaction coverage, detection logic, prioritisation, explainability, data use, hosting, security, resilience and downstream workflow.
For a significant model or algorithm change, the bank should obtain enough information to understand expected behavioural differences and test them on relevant data before deployment where feasible. Regression testing should include known cases and negative controls. The bank should also compare population and output distributions to identify unexpected movement.
Rollback matters. If a release creates a serious control defect, the team needs a tested way to revert, disable a feature, switch to a fallback or apply a compensating control. A rollback plan that exists only in a procedure is weaker than one that has been exercised.
Incidents: a vendor outage is not the only failure mode
The obvious incident is that the service is unavailable. More difficult incidents occur when the platform appears available but the control is degraded.
Examples include a missing list update, partial data ingestion, incorrect enrichment, duplicated records, a corrupted feature, a queue that stops exporting alerts, a model update that changes scoring unexpectedly, a subprocessor outage, or a cyber incident that compromises confidentiality or integrity. Detection controls should therefore watch the control pipeline, not only the web application's uptime.
Incident procedures should define who declares a financial-crime control incident, who can activate fallback, how Legal and Compliance assess regulatory implications, how Operations manages affected payments or cases, how Technology restores service, and how evidence is preserved. If the defect creates historical exposure, a lookback may be required. If it affects sanctions or another time-critical legal control, the decision path can be different from a transaction-monitoring degradation where retrospective review is possible. The applicable legal framework and payment stage matter.
Customer impact also deserves attention. A false-positive surge can delay payments, block onboarding or overwhelm investigators. A false-negative problem can expose the bank to crime and regulatory failure without immediate customer symptoms. Governance must handle both.
The bank's roles and decision rights
Clear ownership prevents vendor issues from falling between Procurement, Technology and Compliance.
The financial-crime control owner defines the control objective, risk coverage and acceptable residual risk. Operations owns day-to-day alert handling and provides evidence about practical quality. Data and technology teams own interfaces, lineage, resilience and technical monitoring. Model or analytics governance provides independent challenge where relevant. Third-party risk and procurement govern due diligence, contract lifecycle and supplier risk. Information security and privacy assess data access, hosting and cyber controls. Legal interprets contractual and jurisdictional requirements. Business continuity and operational resilience teams test severe but plausible disruption and fallback. Internal audit provides independent assurance rather than operating the control.
The vendor has responsibilities too, but it should not become the final decision-maker on whether its own service is effective. A vendor can explain design, investigate defects and provide evidence. The bank decides whether that evidence is sufficient for its control obligations.
Governance forums should have explicit decisions available: approve, approve with conditions, require remediation, restrict use, activate compensating controls, escalate risk acceptance, suspend a release, invoke contract remedies or exit. A forum that can only "note" vendor performance is not governing the risk.
BA, architecture and testing considerations
Business analysts have an unusually important role because vendor governance often fails at the boundary between policy language and system behaviour.
Requirements should identify the complete source population, data fields and effective-dating rules; define which party owns transformations and configuration; document exception and reject handling; specify reconciliation; define audit evidence; map changes to approvals; and state what happens when the service is unavailable. Non-functional requirements should include recovery, performance, retention, security, observability, portability and support, but they should connect to the financial-crime use case rather than being copied from a generic vendor template.
Architects should document trust boundaries and dependencies. A logical diagram should show source systems, integration, data stores, vendor-hosted components, identity and access management, reference-data feeds, case management, reporting, monitoring and fallback. Data residency and subprocessor chains should be visible where relevant. Architecture decisions should avoid designs in which the bank cannot reproduce critical outputs or extract its own control history.
Testing should include more than happy-path functional scripts. Useful test families include population reconciliation, field completeness, scenario boundary tests, model backtesting, late and duplicate records, rejected inputs, clock and business-date behaviour, large-volume performance, reference-data update failures, upstream outage, vendor outage, degraded mode, rollback, alert-routing failure, access revocation, audit-log integrity and data export at termination. Test evidence should be linked to requirements and retained.
What good evidence looks like
A strong vendor-governance file is not a folder of certificates. It tells a coherent story from risk to control.
It should be possible to trace the business need to the vendor selection; selection to due diligence; due diligence to contract terms; contract to architecture and implementation; implementation to validation; validation to go-live approval; production to ongoing monitoring; change events to testing and approval; incidents to remediation; and the whole arrangement to a credible exit plan.
Evidence should also show challenge. If every vendor review is green, every release is approved and every issue is closed by the supplier's explanation, the governance process may be ceremonial. Healthy governance produces questions, conditions, rejected changes and remediation where evidence demands it.
Exit and substitution
Exit planning is easiest before the bank wants to exit.
The bank should know how it will obtain its data, historical alerts, case links, configuration, tuning history, model or rule documentation, audit logs and evidence needed for future investigations or examinations. Proprietary formats can create serious dependence. Data portability should therefore be tested, not assumed from a contractual clause.
A replacement normally needs controlled migration and, for a material control, some form of comparative or parallel validation. The bank should reconcile populations, explain differences in outputs, preserve historical decision evidence, train users and test fallback before the old service is removed. Terminating access too early can destroy the ability to investigate historical cases.
Exit triggers can include persistent control failure, security weakness, service deterioration, unacceptable changes, concentration, vendor financial distress, strategic misalignment or inability to meet new regulatory requirements. Not every trigger requires immediate termination, but governance should know which conditions move the bank from remediation to replacement.
Key takeaways
Vendor monitoring technology should be governed as part of the bank's financial-crime control environment. The bank remains responsible for understanding coverage and effectiveness even when the supplier hosts, configures or supports the tool. Good governance follows the entire life cycle: define the need and criticality, perform due diligence, contract for evidence and control rights, validate on the bank's risk, reconcile the full data path, govern changes, monitor outcomes, rehearse incidents, and preserve a credible exit.
The most important practical test is simple: could the bank explain, without relying on the vendor's assurance alone, why this tool was selected, what it covers today, how the bank knows it works, what changed since go-live, what happens if it fails, and how the bank would replace it? If any answer depends on trust instead of evidence, governance is incomplete.
Operational deep dive: control lineage, validation evidence and run-time observability
A third-party monitoring platform is easiest to govern when the bank can reconstruct what happened from source event to final case decision. That sounds obvious, but many implementations divide responsibility so aggressively that no single team can prove the complete chain. Payments owns the source event, a data platform transforms it, the vendor ingests it, Compliance owns scenarios, Operations reviews alerts, and Technology monitors infrastructure. Each component can be healthy while the end-to-end financial-crime control is incomplete.
Build a control lineage, not only a data lineage
Traditional data lineage explains where a field came from and how it was transformed. A financial-crime control lineage goes further. It should connect:
- the business event that should be monitored;
- the source system and extraction rule;
- transformations and enrichment;
- the vendor input accepted for processing;
- the rule, model, list or configuration version applied;
- the output produced;
- alert routing and prioritisation;
- analyst disposition;
- case escalation, reporting or other outcome; and
- feedback used to tune or remediate the control.
The practical value is accountability. If an instant payment did not produce an alert, the bank can establish whether the transaction was absent, mapped incorrectly, excluded, scored below threshold, suppressed, generated as an alert that failed to route, or reviewed and closed. Without this chain, teams often argue from system logs that prove only their own component worked.
Lineage should be effective-dated. Customer segment, country risk, product mapping, scenario thresholds, sanctions lists and model versions change over time. Historical investigations need the state that applied when the transaction was processed, not today's state silently projected backwards.
Reconcile populations at control boundaries
Population reconciliation is one of the most useful vendor-governance controls because it detects failures that uptime monitoring cannot see. The reconciliation should establish expected and received populations between source systems, bank integration layers, the vendor service and downstream case management.
A daily count can be useful, but count equality is not proof of completeness. Ten thousand source records and ten thousand vendor records can still represent different populations if duplicates replace missing records. Good reconciliation therefore uses stable identifiers, sequence or control totals where available, duplicate checks, rejected-record reporting and field-level completeness measures for risk-relevant data.
Late data require explicit treatment. Some monitoring operates in real time, some near real time and some in batches. The design should define what happens when a file is late, an API call times out, a payment is repaired after initial submission or a backdated transaction enters the ledger. "Processed eventually" may be operationally acceptable for one use case and useless for another.
Validate configuration as well as code
Banks sometimes treat vendor code as the product and bank configuration as administration. In reality, configuration can determine most of the financial-crime outcome. Segment mappings, scenario enablement, thresholds, peer groups, aggregation windows, suppression rules, risk weights and alert-priority rules can materially change detection without any vendor code release.
Configuration should therefore have version control, maker-checker or equivalent approval, test evidence and effective dates. Emergency changes should be identifiable separately and reviewed after implementation. Where the vendor platform does not provide adequate native versioning, the bank should export controlled snapshots.
For a screening engine, the same principle applies to matching thresholds, transliteration options, field weighting, list selection, suppression and whitelisting. A technically unchanged matching algorithm can behave very differently after configuration changes.
Separate four kinds of testing
Technical integration testing proves connectivity, schema handling, authentication, retries and error handling.
Control logic testing proves that rules, matching or model behaviour is consistent with requirements, including boundary and negative cases.
Effectiveness testing asks whether the control identifies relevant financial-crime risk in realistic populations. Historical labelled cases, typology-based test packs, targeted sampling, challenger approaches and investigation outcomes can all contribute.
Resilience testing proves that failures are detected and handled: upstream data loss, vendor unavailability, stale reference data, slow processing, failed alert export, corrupted configuration, subprocessor disruption and rollback.
Keeping these test objectives separate prevents a common mistake in which a successful end-to-end message proves only that a record can travel through the system but is presented as evidence that the AML control is effective.
Model and analytics validation
Where the vendor supplies machine learning or statistical scoring, validation should consider the model's intended use. A score used only to order a queue is different from a score that suppresses alerts or automatically decides that no review is needed. The greater the decision consequence, the stronger the evidence should be.
Useful evidence includes feature definitions, training and validation populations, performance metrics appropriate to the use case, segment-level results, stability, limitations, explainability, drift monitoring, retraining governance and the bank's own testing. The bank does not necessarily need a vendor's proprietary source code to govern the model, but it needs enough transparency to understand the control, challenge performance and investigate failures.
Performance metrics require care. Precision, recall and similar measures depend on labels, and suspicious activity does not provide perfect ground truth. SAR or STR filing is also an imperfect label because filing standards and analyst judgement differ. A mature programme uses several forms of evidence rather than pretending one metric represents "AML accuracy."
Run-time observability for the control
Observability should be designed around the financial-crime service, not merely infrastructure. Useful signals can include source-to-vendor reconciliation, rejected records, null rates in critical fields, processing latency, unprocessed queues, list or reference-data age, rules or models executed, alert-volume shifts by segment, downstream routing failures, case-system backlog and override patterns.
Thresholds should reflect expected behaviour and known business cycles. Month-end, payroll, seasonal commerce or a major migration can legitimately change volumes. The goal is not to alert on every variation but to detect changes that could indicate a control break.
A run-time dashboard should distinguish availability, completeness and effectiveness. Availability asks whether the service is reachable. Completeness asks whether the intended population and fields were processed. Effectiveness asks whether the control continues to produce credible outcomes. These are different questions and should not share one green status.
Evidence package for independent review
For a material vendor tool, an independent reviewer should be able to obtain a compact evidence package containing the current service description and criticality, architecture, source population, data dictionary, configuration inventory, validation report, known limitations, recent material changes, incident history, key performance and effectiveness metrics, subcontractor dependencies, continuity tests and exit status.
The evidence package should not be produced only for audit. Keeping it current reduces incident response time because teams know which versions, owners and dependencies matter. It also makes vendor renewals evidence-based: the bank can decide whether to continue the relationship using control performance and dependency risk rather than procurement momentum.
Acceptance criteria for the technical control
A technically governable vendor monitoring service should be able to demonstrate that every in-scope source has an accountable interface; source-to-vendor populations reconcile; risk-relevant fields are controlled; rejected and late records are visible; configuration is versioned; material changes are testable; historical decisions can be reconstructed; vendor outputs reach the correct case workflow; operational and effectiveness metrics are monitored; fallback and rollback paths are tested; and data plus decision evidence can be exported for investigations and exit.
Those criteria convert a broad requirement such as "the vendor solution must be compliant" into things architects, engineers, testers and control owners can actually prove.
Advanced practice: change governance, performance challenge and concentration risk
The most difficult vendor relationships are rarely the ones that fail on day one. Risk grows after implementation, when the service changes gradually and the bank's original due-diligence evidence becomes stale. Strong governance therefore treats the live relationship as a sequence of controlled decisions rather than a contract that is revisited only at renewal.
Classify changes by control impact
Not every software release needs a full validation cycle. The bank needs a materiality framework that directs effort to changes capable of altering financial-crime outcomes.
A change is more likely to be material when it affects source populations, data fields, matching or detection logic, model features, thresholds, prioritisation, alert suppression, customer segmentation, reference data, explainability, hosting, subcontractors, security, recovery, retention or integration with case management. A user-interface change may be minor unless it changes what evidence analysts can see or how they record a decision.
The change record should describe the old state, new state, expected behavioural effect, affected populations, test plan, approval route, deployment date, rollback method and post-implementation checks. For vendor-managed SaaS, the bank may not control the release date, which makes contractual notice and a representative test environment particularly valuable.
Use a change decision gate
A practical gate asks four questions.
Does the change alter control coverage or decision behaviour? If yes, financial-crime ownership must be involved.
Does it alter an ICT or resilience dependency? If yes, technology, security and third-party-risk assessment may be required.
Can the effect be tested before production? If yes, define the expected differences and acceptance criteria. If not, strengthen post-deployment monitoring and fallback.
Can the bank reverse or compensate for a bad result? A material change without a credible rollback or compensating control should receive stronger challenge.
Post-implementation review should compare actual and expected effects. A release expected to reduce low-value alerts by 10% but instead cuts a high-risk segment by 45% deserves investigation even if technical tests passed.
Performance reviews should challenge explanations
Vendor business reviews often become presentations prepared by the supplier. The bank should bring its own evidence. A useful review combines service measures, control measures, incidents, change history, open limitations and forward roadmap.
When an indicator moves, governance should ask for causal evidence. A fall in false positives might be good. It might also reflect missing input. A reduction in processing time might result from better performance or from logic not executing. A stable alert count can conceal offsetting shifts across segments. The reviewer should therefore look behind aggregate numbers.
Comparison is useful when it is meaningful. The bank can use internal challenger rules, historical cases, targeted sampling or market alternatives to test whether the incumbent tool remains effective. It is rarely possible to compare vendors using one universal "detection rate" because products, labels and data differ. The comparison should be framed around the bank's specific risks and use cases.
Concentration can exist inside one bank
Third-party concentration is not only an industry-wide cloud issue. A bank can create internal concentration by using one provider for transaction monitoring, screening, case management, data enrichment and analytics. A defect, cyber incident or commercial dispute can then affect several lines of defence at the same time.
The dependency map should show which legal entities, products and controls rely on each vendor and its significant subcontractors. It should also identify common infrastructure. Two different vendors hosted on the same underlying service may not provide the independence that procurement records imply.
Concentration decisions should consider substitutability and recovery, not just the number of vendors. Maintaining two suppliers has limited resilience value if the second is not configured, contracted or tested to take over. A realistic mitigation can be a tested fallback, a retained in-house capability, dual sourcing for critical reference data, or a migration plan that can be activated before the primary relationship collapses.
Financial health, ownership and strategy changes
A monitoring vendor can remain technically stable while its corporate risk deteriorates. Acquisition can change product priorities. Financial stress can reduce support or development. A new owner can alter data or hosting strategy. Key staff departures can weaken specialist capability.
The bank does not need to predict every corporate event, but material vendors should be monitored for changes that could affect service continuity or control quality. Change-of-control clauses, notification obligations and renewal rights can create options. The governance question is whether the event changes the bank's risk enough to require deeper diligence, contractual protection, accelerated exit preparation or another action.
Managed services need stronger decision boundaries
Some vendors offer analysts, tuning teams or investigators as part of the service. This can be efficient, but roles must be explicit. A managed service may execute review steps, yet the bank remains responsible for ensuring that local reporting, confidentiality, escalation and customer-treatment obligations are met.
The bank should define which decisions the supplier may make, which require bank approval, how conflicts or uncertain cases are escalated, what training and quality assurance apply to vendor staff, where data can be accessed, and how work is evidenced. Outsourcing operational activity should not create an opaque second case-management process that the bank cannot supervise.
Artificial intelligence does not remove the need for control ownership
Generative or machine-learning features may summarise alerts, rank cases, propose narratives or identify network relationships. Governance should distinguish assistance from decision. A summarisation tool can still introduce material risk if investigators trust incorrect output, sensitive data are sent to an unapproved service, or the tool changes evidence meaning.
The same basic disciplines apply: defined purpose, permitted data, validation, human oversight where required, logging, version control, output testing, known limitations, monitoring and a way to disable the feature. The risk assessment should be based on the actual use and consequence, not the vendor's marketing category.
Renewal should be a fresh risk decision
Renewal is a useful moment to ask whether the original reasons for selection remain valid. The bank should review performance, control effectiveness, incidents, unresolved issues, regulatory change, concentration, price, exit readiness and credible alternatives before the negotiating window closes.
Waiting until a few weeks before contract expiry destroys leverage. For a hard-to-replace monitoring platform, renewal planning may need to begin many months in advance so that the bank has time to test alternatives and negotiate evidence, change and exit rights.
A renewal decision should therefore read like a risk decision, not an administrative continuation: what value the service provides, what material weaknesses remain, what the bank is dependent on, what conditions attach to renewal, and what would trigger replacement.
Practice close: requirements, testing and review checklist
This section converts the chapter into delivery questions that can be used during procurement, implementation, validation or periodic review. It is not a regulatory checklist; the applicable requirements still depend on the bank, service and jurisdiction.
Requirements a BA should make explicit
For the control scope, identify every legal entity, product, customer population, transaction type and channel that must be processed. Define exclusions and who approves them.
For data, identify source systems, required fields, transformations, effective dates, unique identifiers, rejected-record handling, late data and source-to-vendor reconciliation. State which party owns each mapping.
For logic, define whether rules, models, matching or prioritisation are vendor-owned, bank-owned or jointly managed. Require version history, configuration exports, change notice and enough documentation to test behaviour.
For workflow, define how outputs become alerts or cases, priority rules, duplicate handling, evidence available to analysts, disposition codes, escalation and feedback to tuning.
For resilience, define service hours, recovery expectations, fallback, degraded mode, incident severity, notification, rollback and responsibilities during outage.
For evidence, define audit logs, historical reconstruction, retention, regulatory access where applicable, validation artefacts and export on exit.
Minimum implementation test pack
A useful test pack should include normal and high-risk transactions; boundary values; missing and malformed fields; duplicates; late records; reversed or repaired payments; new and changed customers; segment changes; high-volume peaks; upstream failures; vendor rejection; list or reference-data failure; alert-routing failure; vendor outage; access-control changes; material configuration changes; rollback; and export of historical evidence.
Expected results should be based on requirements, not on whatever the vendor platform happens to produce. If a test fails, the defect record should identify whether the cause lies in bank data, integration, vendor configuration, vendor code or downstream workflow.
Questions for a quarterly vendor review
Can the bank reconcile the complete in-scope population? Have any material fields degraded? Which vendor or bank changes affected detection? Did any incident create historical exposure? Are alert and case outcomes changing by segment? Are known limitations still acceptable? Has the vendor changed subprocessors, hosting, ownership or product strategy? Are validation actions overdue? Has exit readiness been tested recently enough for the service's criticality?
The review should end with decisions and owners. "Continue monitoring" is not an adequate action for a repeated defect unless the forum records why the residual risk is acceptable.
Acceptance criteria for renewal or major release
A major release or renewal should not proceed unless material defects have accountable treatment; population and data controls are operating; validation evidence supports intended use; open changes have been tested; incident and continuity obligations are workable; contract rights remain sufficient; and the bank can extract the evidence it would need for a regulator, investigation or exit.
Where an exception is necessary, record the exposure, compensating control, accountable risk owner, target date and trigger for escalation. Time-limited exceptions are governance tools. Permanent exceptions with rolling dates are usually signs that the control design has not been resolved.
Masterclass: the platform was "green" while one payment channel was missing
The following case is fictional but reflects a realistic class of control failure. The numbers are illustrative.
A regional bank replaced an older transaction-monitoring platform with a vendor-hosted service. The programme migrated retail accounts first and added instant payments three months later. The vendor provided the engine and hosting; the bank owned source extraction, product mapping, thresholds and investigations.
The new service performed well in testing. Uptime was stable, daily files completed, alert volumes were close to forecast and the vendor's monthly service report showed every SLA as green. Six months after the instant-payment migration, an investigator noticed that several mule accounts identified through scam complaints had moved most proceeds through the instant rail but had very little corresponding activity in the monitoring case history.
Finding the break
The first assumption was that thresholds were too high. A data analyst instead reconciled payment-hub events to the vendor input using the bank's end-to-end payment identifier. The result showed that about 7% of instant-payment events in one product variant never reached the vendor.
The root cause was a mapping introduced when a new product code went live. The integration layer routed product codes IP01 and IP02 to monitoring but the new variant used IP03. The file-transfer job succeeded every day because the interface processed the records it had selected. The vendor accepted 100% of the records it received. Its uptime and processing SLAs were therefore genuinely green. The control population was still incomplete.
This distinction mattered commercially and operationally. The vendor had not caused the original extraction defect, but the contract and service design also lacked an end-to-end reconciliation requirement. The bank had monitored vendor availability rather than control completeness.
Immediate control response
The bank treated the issue as a financial-crime control incident. Technology corrected the mapping and reconciled the repaired population. Compliance assessed the affected customers and period. Operations introduced a temporary daily population comparison. The vendor helped replay corrected historical data through the relevant scenarios in a segregated environment.
The replay did not automatically turn every generated alert into a suspicious-activity report. Investigators reviewed the cases under the bank's normal legal and policy framework. Some activity was explained by known customer behaviour. Some customers required enhanced due diligence. A smaller set contained scam-proceeds and rapid-onward-movement patterns that were escalated for reporting assessment. Reporting decisions were made under the applicable local requirements, not because the replay itself proved criminality.
The second issue: a vendor model update
During the review, the bank discovered another governance weakness. Two months earlier the vendor had released a new risk-prioritisation model. The release notes described performance improvements but the bank had classified the update as routine because the underlying transaction-monitoring scenarios were unchanged.
Analysis showed that the model did not suppress alerts, but it changed queue priority materially for young retail accounts. The effect was not necessarily wrong, but it had not been independently tested against the bank's mule-account population. The bank therefore paused further model changes, ran segment-level validation and revised its change-materiality rules.
This second finding shows why vendor governance cannot be reduced to "vendor fault." The data gap was created by the bank's integration. The change-governance gap arose from the bank's classification of a vendor release. The supplier had obligations, but accountability for the functioning control sat across the relationship.
Remediation
The remediation programme made five structural changes.
First, source-to-vendor and vendor-to-case reconciliation became a formal control, including identifier-based checks rather than simple totals.
Second, the contract was amended so material processing or model changes required defined notice, impact information and access to a representative test environment.
Third, the vendor score was explicitly documented as a prioritisation input rather than a suspiciousness decision. Queue logic and validation were brought under the bank's model and financial-crime governance.
Fourth, the bank created an incident playbook separating service availability, population completeness, logic defects and downstream routing failures. Different failures now triggered different fallbacks and lookback decisions.
Fifth, exit evidence was tested. The bank exported configuration, historical alerts and audit records to confirm that it could reconstruct decisions without indefinite dependence on the vendor platform.
Lessons from the case
The most important lesson is that a vendor SLA can be correct and the control can still be wrong. The service provider can process everything it receives while the bank sends an incomplete population. A second lesson is that apparently minor vendor changes can alter operational risk when they affect prioritisation, suppression, matching, segmentation or evidence shown to investigators.
The case also illustrates why governance needs engineering evidence. No committee presentation would have found the IP03 gap. Population reconciliation did. No generic "model updated successfully" statement could determine whether prioritisation remained appropriate. Segment-level validation did.
Finally, the case shows how blame can obstruct remediation. The correct question is not "Was this the bank or the vendor?" It is "Where did the end-to-end control break, which responsibilities applied at that point, what historical exposure resulted, and what evidence will prevent recurrence?" That framing produces a better control and a more defensible vendor relationship.
Prove that an exit export is usable
Extend the fictional exit test beyond downloading a file. Ask an independent bank team, without access to the vendor's application, to reconstruct one historical alert and its disposition from the exported evidence. The sample should include the source transactions, the configuration or model version used, relevant reference-data versions, analyst actions and the recorded rationale. The reviewer should be able to distinguish the event time from the time the platform received or transformed the information.
The test should deliberately include a record corrected after the original decision. A successful export preserves both the earlier evidence and the correction, together with their relationship. An export that contains only today's customer profile may look complete while preventing reconstruction of the historical decision. Missing attachments, inaccessible proprietary formats and identifiers that cannot be joined back to bank records should be recorded as exit defects rather than dismissed as documentation issues.
This is a proposed bank acceptance test, not a universal legal export format. Its purpose is to demonstrate that the bank can continue an investigation and answer a legitimate assurance question after the service ends. Contractual exit rights, technical export capability and a usable evidential record are three separate things; the renewal decision should not treat one as proof of the other.
Knowledge check and glossary
Knowledge check
A vendor monitoring platform reports 100% uptime. Does that prove the AML control operated completely?
No. Uptime proves availability of the vendor service, not completeness of the bank's source population, correctness of fields, execution of all intended logic or delivery of alerts to case management. End-to-end reconciliation and control-effectiveness evidence are still required.
Why should a bank validate a vendor model on its own relevant population?
Because performance depends on the customers, products, data, labels and typologies seen by the model. A vendor benchmark can be informative, but it does not establish that the model works for the bank's intended use or for each material segment.
Who owns the risk if a vendor supplies and hosts the monitoring engine?
Responsibilities can be allocated contractually, but the regulated bank does not transfer its own accountability simply by using a third party. The exact legal duties depend on jurisdiction. The bank therefore needs sufficient oversight, evidence and control rights to meet the obligations that apply to it.
Why is change notification important for SaaS monitoring tools?
Because a vendor can alter algorithms, models, infrastructure, subprocessors or other behaviour after initial validation. Material changes may need impact assessment, testing, approval, enhanced monitoring or rollback planning.
When should a vendor incident trigger a lookback?
Not every incident requires one. A lookback becomes relevant when evidence suggests historical customers, transactions or decisions may have been missed or handled incorrectly. Scope should be based on the defect mechanism, affected period, population and applicable legal or policy requirements.
Why can an exit plan fail even when the contract requires data return?
Because practical exit requires usable data formats, historical configuration, evidence, migration capacity, trained people and time to validate a replacement. A contractual right that has never been tested can still leave the bank operationally trapped.
Glossary
Third-party arrangement: A relationship in which an external provider supplies a product or service. Regulatory definitions and classifications differ by jurisdiction.
Outsourcing: A regulatory and contractual concept used when a third party performs an activity or function for the institution. The exact definition is jurisdiction-specific and should not be assumed from ordinary language.
Critical or important function/service: A function or service whose disruption could materially affect the institution or its obligations under the applicable framework. Different regimes use different terminology and tests.
Control owner: The bank role accountable for defining the objective, coverage, performance expectations and residual risk of the financial-crime control.
Population reconciliation: Evidence that all intended source events reached each required processing stage without unexplained omission or duplication.
Model drift: Change over time in data, relationships or model behaviour that can reduce or alter performance.
Configuration drift: Uncontrolled or poorly understood change in rules, thresholds, mappings, segmentation, suppression or other settings.
Subprocessor / subcontractor: Another provider used by the contracted vendor to deliver part of the service. Its significance depends on the arrangement and applicable regulatory framework.
Degraded mode: A controlled operating state used when the normal service is partly unavailable, with defined limitations, monitoring and decision rights.
Exit readiness: The demonstrated ability to terminate or replace the service while preserving continuity, data, historical evidence and control effectiveness.
References and further reading
The sources below are public, authoritative materials used for this chapter. They are not interchangeable legal authorities. Apply the rules and supervisory expectations relevant to the bank's legal entity, service and jurisdiction.
Global third-party risk and resilience
- Basel Committee on Banking Supervision, Principles for the sound management of third-party risk, 10 December 2025: https://www.bis.org/publications/202512-guidelines-principles-sound-management-third-party-risk
- Financial Stability Board, Enhancing Third-Party Risk Management and Oversight: A toolkit for financial institutions and financial authorities, 4 December 2023: https://www.fsb.org/2023/12/final-report-on-enhancing-third-party-risk-management-and-oversight-a-toolkit-for-financial-institutions-and-financial-authorities/
United States
- Federal Reserve, FDIC and OCC, Interagency Guidance on Third-Party Relationships: Risk Management, 6 June 2023: https://www.federalreserve.gov/newsevents/pressreleases/bcreg20230606a.htm
- Federal Reserve, Third-Party Risk Management: A Guide for Community Banks, updated 3 May 2024: https://www.federalreserve.gov/publications/third-party-risk-management-a-guide-for-community-banks.htm
- FFIEC, BSA/AML Examination Manual: https://bsaaml.ffiec.gov/
European Union
- EUR-Lex, Regulation (EU) 2022/2554 on digital operational resilience for the financial sector (DORA), applicable from 17 January 2025: https://eur-lex.europa.eu/eli/reg/2022/2554/oj
- European Banking Authority, Guidelines on third-party risk management page, including the final 2026 Guidelines for non-ICT services and their application status: https://www.eba.europa.eu/activities/single-rulebook/regulatory-activities/internal-governance/guidelines-third-party-risk-management
- European Banking Authority, Guidelines on outsourcing arrangements (2019 framework): https://www.eba.europa.eu/activities/single-rulebook/regulatory-activities/internal-governance/guidelines-outsourcing-arrangements
United Kingdom
- Prudential Regulation Authority, SS2/21 Outsourcing and third party risk management: https://www.bankofengland.co.uk/prudential-regulation/publication/2021/march/outsourcing-and-third-party-risk-management-ss
Financial-crime monitoring effectiveness
- Wolfsberg Group, Statement on Effective Monitoring for Suspicious Activity, July 2024: https://wolfsberg-group.org/resources/195/202
- Wolfsberg Group, Second Statement on Effective Monitoring for Suspicious Activity: a responsible framework for innovation, 27 August 2025: https://dev.wolfsberg-group.org/news/the-wolfsberg-group-publishes-its-second-statement-on-effective-monitoring-for-suspicious-activity
- FATF, The FATF Recommendations, current consolidated page: https://www.fatf-gafi.org/en/publications/Fatfrecommendations/Fatf-recommendations.html