Behaviour and Peer-Group Pattern Review
A payment is not suspicious merely because it is large, new or unusual. The same transaction can be routine for one customer and highly inconsistent for another. Behavioural monitoring therefore asks two connected questions: is this activity different from what this customer normally does, and is it unusual when compared with genuinely similar customers? The first question uses the customer as their own benchmark. The second uses a relevant peer population to add context that the customer's history cannot provide on its own.
This is an important distinction in financial-crime monitoring. FATF Recommendation 10 requires ongoing due diligence, including scrutiny of transactions so that they are consistent with the institution's knowledge of the customer, the customer's business and risk profile, including source of funds where necessary. FATF's banking-sector risk-based guidance goes further operationally: ongoing monitoring can identify changes in behaviour and may compare a customer's activity with that of a peer group. The guidance also says banks should clearly document the criteria and parameters used for customer segmentation. None of this means that FATF prescribes one peer-group algorithm, one anomaly score or one threshold. Those are implementation choices that a bank must design, validate and govern according to its risks, products, data and applicable law.
The practical mental model is therefore simple. A behavioural signal is evidence for review, not a legal conclusion. A self-history deviation can be innocent because the customer's life or business changed. A peer deviation can be innocent because the peer group was badly constructed. A strong case normally comes from several facts lining up: the deviation is material, the baseline is reliable, the peer comparison is meaningful, the customer's explanation does not resolve the concern, and other customer, payment, counterparty or network evidence supports further investigation.
Where behavioural review sits in the bank
Behavioural review sits between customer understanding and investigation. KYC and KYB establish what the bank knows about the customer: identity, occupation or business, expected products, expected geographies, anticipated volumes, counterparties where relevant and other risk factors. Transaction and event data then show what actually happens. Behavioural logic compares those observations with the profile and with prior activity. Detection may create an alert or another review task. An investigator then has to decide whether the signal can be explained, needs more evidence, should be escalated into a case, or contributes to a suspicious-activity reporting decision under the law that applies to the relevant entity.
That lifecycle involves more than the transaction-monitoring platform. Customer master data, account and product systems, payment hubs, card platforms, digital-channel telemetry, sanctions and screening systems, case management, data warehouses, analytics platforms and regulatory-reporting tools may all contribute evidence. A bank that calculates sophisticated anomaly scores on incomplete payment data has not built a strong control. Likewise, a bank with excellent KYC data but no reliable way to connect customer-profile changes to monitoring will keep comparing new behaviour with stale expectations.
For business analysts and architects, the key design principle is traceability. A reviewer should be able to see which data produced the signal, which baseline period was used, which peer segment applied, which model or rule version generated the result, what data was missing, and what subsequent action was taken. Without that lineage, a score becomes difficult to explain, test or defend.
Self-history: using the customer as their own benchmark
Self-history is usually the most intuitive comparison. A salaried customer may receive one employer credit near month end, pay recurring bills and make occasional discretionary transfers. A wholesaler may receive hundreds of commercial credits and make large supplier payments during working hours. The absolute amounts could overlap, but the behavioural meaning is different.
A useful baseline can include value, count, frequency, transaction type, channel, counterparty, geography, currency, timing, balance movement and sequence. The bank does not need every possible feature. It needs features that are relevant to the risk it is trying to detect and that are supported by reliable data. Features also need business meaning. A model that says "behaviour changed" without allowing the reviewer to understand which behaviour changed creates operational friction and weakens investigation quality.
Baseline windows need care. A very short lookback can overreact to temporary volatility; a very long lookback can preserve history that no longer represents the relationship. A retail customer may change employer, move country, sell a property or receive an inheritance. A company may acquire another business, enter a new market or switch suppliers. These events can change genuine behaviour abruptly. Banks can respond through event-driven profile updates, change-point methods, segmented windows or reviewer-confirmed re-baselining. The appropriate technique depends on the control design. No single reset method is universally required.
New customers create a cold-start problem because there is not enough self-history to establish a stable norm. In that situation, expected activity captured during onboarding and peer information can carry more weight, while the bank builds customer-specific history over time. Dormant relationships create a related problem: activity from years ago may not be a credible current baseline. A substantial reactivation can therefore require contextual review or refreshed customer information, depending on risk and policy, rather than automatic suspicion.
Peer groups: useful only when similarity is real
Peer comparison is powerful because it gives context beyond the individual. It can help with new customers, sparse histories and activity that is unusual for the customer but potentially normal for a legitimate segment. FATF's banking-sector guidance explicitly recognises peer-group comparison as one possible monitoring technique, but it does not define a mandatory segmentation model.
A peer group can use customer type, business sector, product mix, size, geography, legal form, channel usage or other risk-relevant characteristics. The important question is whether those characteristics produce a group whose behaviour is meaningfully comparable. "Small business" may be too broad if it combines restaurants, software firms and importers. "Retail" may be too broad if it combines students, pensioners, gig workers and high-income international professionals. A segment that is easy to build technically may still be useless analytically.
A strong design therefore tests both business logic and observed behaviour. The bank should know why each segmentation variable is included, whether the resulting groups are sufficiently populated, whether the distributions are stable enough to interpret, and what happens when a customer does not fit any group well. Small or highly specialised populations may need broader hierarchical comparison, manual contextual review or no peer benchmark at all rather than a statistically fragile result presented with false precision.
Statistics also need to reflect the data. Transaction values are often skewed, so medians, percentiles or other robust measures may be more informative than a simple average. A population can also be multimodal: two legitimate subgroups may behave very differently even though a coarse customer classification places them together. In that case, a single "normal" value can misrepresent both groups. The right response is not automatically more modelling; it may be better segmentation or a clearer limitation statement.
Peer groups can drift. Economic conditions, product changes, new customer acquisition and changes in payment behaviour can alter a segment. Monitoring should therefore include some mechanism to identify when assumptions no longer hold. Recalibration does not have to happen on a universal calendar. It should happen often enough, and through event-driven review where appropriate, to keep the comparison relevant to the institution's risk profile.
Deviation is a question, not an answer
A behavioural deviation should normally start a question. It should not finish one. A customer sending money to a new country may reflect a new supplier, relocation, family support, investment or criminal activity. A sudden increase in turnover may reflect business growth, a seasonal peak, pass-through activity, scam proceeds or layering. The monitoring result tells the bank that something deserves attention; it does not provide the missing explanation by itself.
A practical review separates several kinds of deviation. Magnitude asks whether value or count changed materially. Pattern asks whether the structure of activity changed, such as many credits followed by rapid dispersal. Velocity asks whether timing changed, for example rapid in-and-out movement. Novelty asks whether new counterparties, channels, devices, products or corridors appeared. These categories can overlap and are analytical aids rather than legal classifications.
Good investigation then reconstructs the context. What did the bank expect? What happened? How reliable was the baseline? Is the peer group appropriate? Did the customer's profile change? Are there connected accounts or counterparties? Do payment references, merchant information, invoices or other available evidence explain the activity? Does the explanation fit the customer's broader behaviour? Are there other red flags? The analyst should test plausible legitimate explanations as well as risk hypotheses rather than treating the alert as a presumption of wrongdoing.
From alert to case and reporting outcome
A behavioural signal may close at alert level if the bank can explain it adequately under its procedures. It may be promoted to a case when several transactions, customers or risk indicators need a joined investigation. It may trigger enhanced monitoring, a KYC refresh, a fraud handoff, payment investigation or another control action depending on the facts. A suspicious transaction or suspicious activity report is a separate legal decision governed by the jurisdiction of the reporting entity. The existence of a peer deviation does not itself establish that the legal reporting threshold has been met.
This separation matters in system design. The detection layer should preserve what it observed without pretending to know the final legal outcome. Case management should allow investigators to add evidence, connect related alerts and record reasoning. Regulatory-reporting systems should capture the reportable facts required by local law. Customer restrictions, exits or payment holds also need their own legal and policy basis; they should not be treated as automatic consequences of an anomaly score.
Feedback can improve monitoring, but labels are imperfect. A filed SAR or STR is not necessarily proof that crime occurred. A closed alert is not necessarily proof that the activity was benign. Law-enforcement feedback may be limited. Customer explanations can later prove incomplete. Model development should therefore distinguish between operational outcomes and high-quality ground truth. Where labels are weak, validation should acknowledge that limitation rather than turning uncertain outcomes into confident training targets.
Behaviour patterns and typology context
Behavioural analysis becomes most useful when it identifies combinations rather than isolated values. A funnel-account pattern, for example, may involve many unrelated incoming credits followed by rapid onward movement. Structuring can involve repeated amounts, timing and channels designed to avoid a control threshold. Round-trip behaviour can involve value leaving and returning through related counterparties. A mule network can show coordination across accounts that individually appear ordinary.
These are not self-proving signatures. Legitimate businesses can aggregate payments. Treasury functions can move funds rapidly. Families can share devices or addresses. Marketplaces can collect and disperse value. The investigation has to compare the pattern with the declared business model, product mechanics, ownership, counterparties and other evidence. Behavioural analytics should therefore help find relationships and sequences that deserve review, not replace commercial understanding.
Network behaviour beyond one account
Criminal activity often spans accounts. Shared devices, addresses, phone numbers, beneficiaries, funding sources or timing can reveal relationships that are invisible when each account is reviewed separately. Network analysis can therefore complement self-history and peer comparison, particularly for mule activity, scam proceeds and organised laundering.
The central risk is guilt by association. A common employer, apartment building, remittance corridor or household device may connect entirely legitimate customers. Linkages should carry type, strength and provenance so that investigators can distinguish a verified relationship from a weak technical overlap. Entity-resolution logic should also be tested for both missed links and false merges. A mistaken merge can create a false network; a missed merge can fragment one.
From an architecture perspective, network features should be reproducible. If an investigator sees that an account is "linked to 14 high-risk entities", the case should expose enough information to understand how those links were constructed and when they existed. Effective dating matters because ownership, devices and counterparties change.
Temporal behaviour: rhythm, velocity and structural change
Timing can be as informative as amount. Salary cycles, merchant opening hours, seasonal trading, holidays and invoice schedules create legitimate rhythms. A sudden change in rhythm can indicate account takeover, coercion, business change or criminal use, but the signal still requires context.
Velocity measures need precise definitions. "Rapid movement" can mean seconds for an instant-payment scam, hours for a mule funnel or days for another laundering pattern. Systems should define the window, event types, currencies and aggregation logic rather than use vague labels that testers cannot reproduce. Time zones and cut-off conventions also matter in cross-border or multi-channel data.
Structural change is harder than a simple spike. A customer's activity can move gradually into a new regime. Statistical change-point techniques may help identify a shift, but they should not automatically rewrite the customer's expected behaviour. Re-baselining without verification can normalise criminal use; refusing to re-baseline can punish legitimate transformation. The safe design is controlled change: detect, contextualise, approve where necessary, effective-date the new baseline and retain enough history for retrospective investigation.
Behavioural scores and explainability
Some institutions combine multiple features into a behavioural score used for alerting or prioritisation. The score may include deviation from self-history, peer-relative position, network indicators, customer risk and other signals. The model can be simple or sophisticated. The governance requirement should be proportionate to how materially it affects control decisions and should reflect the bank's applicable model-risk, technology-risk and financial-crime governance frameworks.
Explainability is operationally important. A reviewer needs more than a number. Useful output may show which features drove the score, the comparison period, peer segment, missing-data flags and relevant change since the previous assessment. Explainability also helps testers identify defects and helps control owners understand whether a model is responding to genuine risk or to a data artefact.
Performance measures should be chosen with label limitations in mind. Precision and recall can be useful where credible labelled outcomes exist. Other controls may rely more on alert review, typology coverage, stability, challenge testing, investigative utility and targeted lookbacks. A bank should not claim that a high "accuracy" number proves financial-crime effectiveness if the outcome labels are mostly analyst dispositions rather than confirmed criminal facts.
Fairness, data protection and legitimate atypical behaviour
Peer analysis can create customer harm if "different" silently becomes "bad". Atypical but legitimate customers include people with irregular income, international family ties, gig work, seasonal occupations or unusual commercial models. A good control protects against both missed crime and unjustified adverse treatment.
The exact legal framework for fairness, equality and use of personal data varies by jurisdiction. AML standards do not create one global fairness metric or require banks everywhere to collect protected-characteristic data for model testing. In some jurisdictions, collecting or using such data may itself be restricted. The bank should therefore work with legal, privacy, compliance and model-governance specialists to decide what testing is lawful and appropriate. Where direct demographic data cannot be used, the institution can still examine false-positive concentrations, complaints, restrictions, segment performance, feature relevance and other indicators of unintended harm.
The strongest protection is analytical discipline: build meaningful peer groups, avoid weak proxies, document limitations, seek legitimate explanations, apply proportionate customer actions and provide appropriate review or appeal mechanisms where policy and law require them. Fairness should not be presented as a reason to ignore risk; it is a reason to make the risk judgement better.
Business-customer behaviour
Corporate customers need different baselines from retail customers. Turnover can be seasonal, accounts can support several legal entities, intercompany flows may be normal, supplier corridors can change and acquisitions can transform activity overnight. A restaurant, software company and commodity trader are poor peers merely because all are SMEs.
Useful corporate context can include sector, turnover band, legal structure, group relationships, payment corridors, products, currencies and known trading cycles. Financial statements, invoices, contracts or trade data may help where available and proportionate. The bank should not assume that every transaction-monitoring platform receives those documents in structured form. Requirements should state what data is actually available at decision time.
Related-party flows deserve contextual review because they can be legitimate treasury activity or part of layering. Beneficial-ownership and group-structure data can help explain them, but ownership data must be current and historically reconstructable. Similarity to a typology is not enough to infer lack of commercial substance.
Data architecture and lineage
A behavioural control is only as reliable as the observations feeding it. Transaction feeds need completeness, timestamps, status, reversals, currencies, parties and channel data appropriate to the use case. Customer profile data needs effective dates. Product and account mappings need consistency. Peer assignment needs versioning. Derived features need lineage from source event to final score.
A practical data model can separate five layers: source events; normalised transaction and customer facts; derived behavioural features; baseline or peer reference data; and detection outcomes. That separation helps testing because a team can isolate whether a bad alert came from source data, transformation logic, peer assignment, feature engineering or final decision logic.
Late-arriving data needs explicit treatment. If a customer-profile update arrives after a transaction, the system should know which profile version was available when the alert was generated. Replays should be controlled so that corrected data does not silently overwrite the audit history. The same principle applies to peer assignment and model versions.
Roles and governance
First-line business and operations teams often own customer knowledge and can explain commercial activity. Financial-crime operations review alerts and cases. AML compliance defines policy, challenges control design and handles or oversees reporting decisions according to the bank's governance. Data and engineering teams build pipelines and analytical services. Model-risk or independent validation functions may apply where the institution's model inventory and governance classify the analytics as models. Legal and privacy teams advise on data use and customer actions. Internal audit provides independent assurance over governance and control operation.
The precise division differs by bank and jurisdiction. What matters is that ownership is explicit. Someone must own the behavioural-control objective, someone must own data quality, someone must approve material logic changes, someone must review control performance, and someone independent enough from development must challenge the design in proportion to its risk.
Governance forums should avoid a common failure: treating alert-volume reduction as success on its own. A lower volume may mean better targeting or may mean lost coverage. Change proposals should therefore show impact on risk coverage, customer burden, operational capacity and known limitations. Significant changes should be traceable to evidence and an approval decision.
BA, architecture and testing considerations
For a BA, "compare the customer with peers" is not a testable requirement. A useful requirement specifies the population, segmentation attributes, exclusions, minimum data sufficiency, observation period, feature definitions, comparison method, trigger logic, explainability output, downstream action and fallback behaviour when required data is missing.
For architecture, the main choices include batch versus event-driven feature generation, central versus product-specific behavioural stores, online versus offline scoring, lineage and versioning, resilience, latency and access control. Instant-payment fraud and AML monitoring may use some of the same data but have different time horizons and legal purposes. Reuse should therefore focus on reliable data and shared services without collapsing distinct decision processes into one generic risk score.
Testing should include positive scenarios, legitimate negative scenarios, boundary conditions, data-quality failures, time-zone effects, new-customer cold starts, dormancy, peer reassignment, structural breaks, backdated profile changes, model-version changes and linked-account patterns. Historical back-testing can be useful but must avoid future-data leakage: the test must use only the information that would have been available at the historical decision time.
Fairness and customer-impact testing should be designed with legal and privacy input. Reviewers should also test whether explanations are intelligible: a technically correct alert that cannot tell an investigator what changed can still be operationally weak.
Failure modes to look for
The most common design failure is treating the peer group as ground truth. Another is allowing the baseline to learn suspicious behaviour so gradually that the abnormal becomes normal. A third is using sparse or stale data while presenting the result with unjustified precision. Others include hidden changes to segmentation, inconsistent customer identifiers, missing reversals, poor time-zone normalisation, outcome labels that confuse SAR filing with confirmed crime, and threshold tuning driven only by queue pressure.
Operational failures matter as much as model failures. A strong signal can still be lost in an ageing queue. A customer can be asked repeatedly for information because cases are not linked. An investigator can see a score but not the transactions behind it. A system outage can prevent peer features from refreshing while scoring continues as if they were current. Requirements should specify degraded-mode behaviour and monitoring so that stale or missing analytical inputs do not masquerade as healthy control operation.
Mini case study: a wholesaler changes corridor and velocity
Consider a fictional mid-sized electronics wholesaler. For two years it paid established suppliers in three countries, usually in two weekly batches, with monthly outbound volume between roughly €1.2 million and €1.8 million. It then begins sending €250,000 to €400,000 transfers several times per day to five newly introduced counterparties in two additional countries. The figures are illustrative and are not regulatory thresholds.
The self-history lens identifies several changes: new countries, new counterparties, higher frequency and shorter time between incoming customer funds and outgoing supplier payments. The peer lens shows that some similar wholesalers expanded into the same markets after a supply-chain disruption, so geography alone is not unusual. The customer-risk profile has not been refreshed since before the expansion.
Triage should not jump from "new corridor" to suspicion. The investigator checks whether the business has genuinely expanded, whether the new suppliers exist, whether invoice values and products are coherent with the customer's trade, whether ownership links create concern, whether incoming funds and outgoing transfers show pass-through behaviour, and whether the new pattern is consistent with known commercial events. Assume the bank receives contracts and shipping evidence supporting two suppliers, but the remaining three counterparties are newly incorporated, share directors and receive funds almost immediately after the wholesaler receives unrelated third-party credits.
At that point, the behavioural signal has become useful because it directed the bank toward a joined set of evidence. The case may require deeper investigation, possible KYC refresh, connected-party analysis and assessment against the reporting threshold applicable to the bank's legal entity. The peer result did not prove innocence; the self-history deviation did not prove suspicion. Their value was to structure the questions and show where corroboration mattered.
For a tester, the case also creates concrete scenarios: verify that the two-year baseline is effective-dated; confirm the new-corridor feature does not double-count returned payments; test peer reassignment after the customer's sector or size changes; verify that linked directors are derived from the correct ownership source; and confirm that the case record preserves the feature values and model version that existed when the alert was created.
Key takeaways
Behavioural monitoring works best when it compares activity with both the customer's own history and an honestly constructed peer context. Neither comparison is proof of wrongdoing. The control becomes effective when deviations are explainable, data lineage is preserved, investigations test legitimate as well as risk hypotheses, and outcomes feed governed improvement without pretending uncertain labels are perfect truth.
Peer grouping is an analytical choice, not a globally prescribed segmentation scheme. Statistical sophistication should be proportional to the problem and data. A simple, transparent comparison built on reliable customer knowledge can be more useful than an opaque model with weak lineage.
For delivery teams, the decisive questions are practical: what data is available, what exactly is compared, how is change handled, how are peers selected, what happens when data is thin or stale, what does the investigator see, and how does the bank prove that changes improve coverage rather than merely reduce alerts?
References and further reading
- Financial Action Task Force (FATF), The FATF Recommendations: https://www.fatf-gafi.org/en/publications/Fatfrecommendations/Fatf-recommendations.html
- Financial Action Task Force (FATF), Guidance for a Risk-Based Approach: Banking Sector: https://www.fatf-gafi.org/content/dam/fatf-gafi/guidance/Risk-Based-Approach-Banking-Sector.pdf.coredownload.pdf
- Basel Committee on Banking Supervision, Anti-money laundering and counter-terrorist financing, Basel Consolidated Guidelines AFS10 (current consolidated module, published 1 January 2026): https://www.bis.org/committees/bcbs/basel-consolidated-guidelines/module/afs/10
- Wolfsberg Group, Statement on Effective Monitoring for Suspicious Activity, Part I: Moving Beyond Automated Transaction Monitoring: https://wolfsberg-group.org/resources/general/168
- Wolfsberg Group, Statement on Effective Monitoring for Suspicious Activity, Part II: Transitioning to Innovation: https://wolfsberg-group.org/resources/195/202
- European Banking Authority, Guidelines on ML/TF risk factors: https://www.eba.europa.eu/legacy/regulation-and-policy/regulatory-activities/anti-money-laundering-and-countering-financing-1
- US Federal Financial Institutions Examination Council, Customer Due Diligence — BSA/AML Examination Manual: https://bsaaml.ffiec.gov/manual/AssessingComplianceWithBSARegulatoryRequirements/02
- US Federal Financial Institutions Examination Council, Suspicious Activity Reporting — BSA/AML Examination Manual: https://bsaaml.ffiec.gov/manual/AssessingComplianceWithBSARegulatoryRequirements/04
- UK Financial Conduct Authority, Financial Crime Guide (FCG): https://handbook.fca.org.uk/handbook/fcg1/fcg1s1
Operational deep dive: change, peer statistics and drift
The base chapter explains why self-history and peer comparison need to work together. This deep dive focuses on the analytical machinery underneath them. The important principle is that statistical sophistication does not create regulatory truth. A change-point, percentile or anomaly score is useful only when the institution understands what the calculation means, what data it depends on and how a reviewer should use it.
Detecting structural change without automatically normalising it
Customer behaviour changes for legitimate and illegitimate reasons. A statistical change-point method can help identify when the distribution of activity appears to move into a new regime, but the output should normally trigger contextual assessment rather than an automatic baseline reset.
The method can be simple. A bank might compare rolling averages, transaction counts or counterparty diversity across windows and flag a material shift. More advanced implementations may use probabilistic or machine-learning techniques. In either case, the requirements should define the observation window, minimum data, features, sensitivity, treatment of missing data and the downstream action. A tester should be able to reproduce why a change was flagged.
A controlled re-baselining process should preserve historical state. If a company acquires another business and its payment volume legitimately doubles, the monitoring system may need a new baseline. Effective dating allows the bank to record when the new baseline became applicable without erasing what was expected before the acquisition. That matters for later investigations and for understanding whether activity before the change was already unusual.
Slow change is particularly important. Criminal use can be introduced gradually, but legitimate businesses also grow gradually. The bank should therefore avoid the simplistic rule that gradual change is suspicious or that gradual change is safe. Change detection is a way to direct attention; evidence still determines the outcome.
Building peer statistics that fit the population
Peer benchmarking starts with segmentation, not mathematics. If the population is not meaningfully comparable, sophisticated statistics will only quantify a bad assumption more precisely.
The bank should first document why a segment exists. Business sector, customer type, size, geography, product mix and channel use can all be relevant, but each extra dimension fragments the population. Very small groups can become unstable and expose individual behaviour through the benchmark itself. The design therefore needs a minimum-data strategy: merge into a broader parent segment, use a different comparison method, suppress the peer output or require manual review when data is too thin.
Distribution shape matters. Transaction values are commonly skewed, so a median and percentile range may be more interpretable than a mean and standard deviation. A multi-modal distribution can indicate that the segment actually contains different legitimate populations. The right response may be sub-segmentation rather than increasingly complex outlier logic.
Outliers should not simply be removed because they make the benchmark inconvenient. Some are data errors, some are legitimate extremes and some may be exactly the risk the control is intended to identify. Any trimming or winsorisation rule should therefore have a clear analytical purpose and should be included in validation and change control.
Drift: knowing when yesterday's norm is no longer today's
A baseline or peer benchmark can become stale even when the code never changes. Product migrations, inflation, market disruption, customer acquisition, new payment rails and changes in customer behaviour can all alter a distribution. Drift monitoring asks whether the data being scored still resembles the data on which the comparison was designed.
Useful drift indicators can include movement in feature distributions, changes in peer-group sizes, changing alert rates, changes in missing-data rates, instability in segment assignment and deterioration in the usefulness of alerts to investigators. No single metric proves that a model has failed. A rise in alert volume, for example, could reflect genuine risk, data duplication or a population shift. Drift review needs both technical and business interpretation.
Recalibration should be governed in proportion to impact. A minor percentile refresh may follow a lighter process than a change that redefines customer segments or materially alters who receives alerts. The bank should preserve versions and effective dates so that a historical case can be reconstructed using the logic that actually operated at the time.
Labels are evidence, not perfect ground truth
Validation often needs labelled outcomes, but financial-crime labels are difficult. A SAR or STR indicates that the institution reached a reporting threshold under applicable law; it is not a court finding that laundering occurred. A closed alert means the bank did not escalate on the evidence available; it does not prove that the activity was legitimate. Law-enforcement outcomes may arrive much later or not at all.
This matters for machine learning and for apparently simple performance measures. If the bank trains a model to reproduce analyst dispositions, the model may learn historical inconsistency and bias as well as useful patterns. If it treats every filed report as a confirmed positive, performance statistics may appear more certain than the underlying reality supports.
A sensible validation approach can therefore use multiple evidence classes: high-confidence known outcomes where available, reviewed alert and case outcomes, typology-based synthetic tests, targeted lookbacks and investigator feedback. The methodology should state the strengths and limitations of each source rather than collapsing them into one binary truth label.
Probabilistic reasoning without false precision
Probability can help organise thinking, but it should be used carefully. In some analytical environments, teams may estimate how strongly particular evidence changes the likelihood of a risk outcome. In many AML settings, however, reliable base rates for "criminal" versus "legitimate" behaviour do not exist because observed labels are incomplete and selective.
For that reason, Bayesian language should not be used to imply that a bank can calculate a scientifically precise probability that a customer is laundering money. A more defensible use is to make assumptions explicit: what was the prior risk context, what new evidence was observed, how reliable is that evidence, and how did it change the investigative hypothesis? Where quantitative probabilities are used, calibration should be tested against the best available outcome data and uncertainty should be visible.
The same caution applies to behavioural risk scores. Ranking cases may be useful even when a score is not a literal probability. Documentation should say what the score represents. Calling a rank score "85% suspicious" without evidence that the number is calibrated as a probability is misleading.
Adversarial and robustness testing
Criminal actors can adapt to controls, so robustness testing should ask whether small changes in behaviour can defeat detection. This does not require trying to infer secret criminal knowledge of the bank's exact thresholds. A safer and more testable approach is to construct controlled scenarios around known design boundaries: activity just below a threshold, distributed activity across related accounts, gradual rather than sudden changes, or data mutations that challenge entity resolution.
Testing should happen in controlled QA or analytical environments, not by injecting uncontrolled fake activity into production customer queues. Synthetic cases, masked historical data and replay environments can test behaviour without contaminating regulatory records or confusing operational reviewers.
Defence in depth can reduce dependence on one comparison method. A self-history lens, peer context, network linkage and typology logic may each see different parts of a pattern. Combining them can improve coverage, but fusion logic should remain explainable enough that analysts and validators understand why a case was prioritised.
Validation and governance in practice
Not every behavioural calculation is necessarily classified as a formal "model" under every bank's framework. The institution should apply its own model-risk, technology-risk and financial-crime governance definitions. Materiality matters: a peer chart used only as analyst context is different from an automated score that determines alert creation or customer restriction.
Whatever label the bank uses, the control still needs ownership, data-quality checks, versioning, testing and change management. Validation should examine conceptual soundness, data fitness, performance, stability, limitations and operational use. Where an external vendor supplies the analytics, outsourcing does not remove the bank's responsibility for understanding how the capability affects its own customers and regulatory obligations.
The practical test is reconstructability. For a material historical case, can the bank show the customer data, peer assignment, baseline period, feature values, logic version, output and downstream decision that existed at the time? If not, the analytics may be impressive but the control is difficult to investigate, challenge or audit.
Advanced practice: corporate baselines, networks and operating reality
Behavioural monitoring becomes harder when the customer is a business, activity spans several accounts or jurisdictions, and analysts must make decisions with incomplete evidence. This supplement focuses on those operating conditions rather than assuming a clean retail account with a stable history.
Corporate baselines need business context
A company can change transaction behaviour because its business genuinely changed. Revenue growth, an acquisition, a new distributor, new treasury arrangements or supply-chain disruption can alter value, frequency, counterparties and corridors at the same time. A corporate baseline that ignores those events can generate large volumes of technically correct but low-value deviations.
The design should therefore separate scale from pattern where possible. A larger company may legitimately move more money while retaining a stable supplier structure. Another company may keep the same monthly value but suddenly replace long-standing counterparties with newly created entities. Those are different behavioural changes and should not be reduced to one percentage deviation.
Business context can come from KYC/KYB records, sector, turnover band, group structure, products, expected corridors and other verified information available to the bank. External commercial data can help where it is lawfully obtained and reliable, but the control should never assume that every commercial dataset is complete or current. Data provenance and effective dating remain important.
Intercompany flows illustrate the challenge. Transfers within a corporate group can be ordinary liquidity management, but related entities can also be used to layer funds. Monitoring should expose ownership and relationship information to the investigator rather than treating every related-party payment as suspicious or excluding all of them as normal.
Network behaviour across customers
Some risk becomes visible only when accounts are connected. Multiple customers can share a beneficiary, device, address, phone number, funding source or transaction timing. A mule network may appear ordinary at account level but coordinated at network level.
Entity-resolution quality therefore matters. A shared identifier should have a type and confidence. An exact government identifier is different from a common address. A device fingerprint can be strong in one context and weak in another. A common employer can connect hundreds of legitimate people. Investigators need to understand the nature of the link before treating network proximity as evidence.
Graph features can support prioritisation: number of connected accounts, shared counterparties, flow direction, timing and concentration can all be useful. But network outputs should avoid labels such as "criminal cluster" unless the evidence actually supports that conclusion. A safer design describes the observed structure and lets the case process determine significance.
Privacy and access controls deserve special attention because network analysis combines information across customers. The bank should define who can see linked-customer data, how long derived relationships are retained, and how erroneous links are corrected. Local privacy and secrecy laws can affect what may be combined across entities or countries.
Temporal operations and cultural context
Time-based behaviour is useful only when the clock is interpreted correctly. Salary cycles, market hours, public holidays, religious observances, weekends and seasonal trading can all change patterns. Global banks should avoid applying one headquarters calendar to every customer population.
A timestamp also needs technical meaning. Event time, booking time, value date and ingestion time are not interchangeable. If a payment occurred at 23:58 local time but arrived in the analytical platform after midnight UTC, a poorly designed daily-velocity feature can put it in the wrong window. Requirements should identify the authoritative timestamp and time-zone treatment for each feature.
Dormancy is another temporal issue. A quiet account that becomes active may deserve contextual review because the previous baseline is stale, but dormancy alone is not suspicion. Re-verification or enhanced review depends on risk, policy and applicable CDD requirements. The same principle applies to round-the-clock activity: it can indicate automation or account takeover, but it can also be normal for a global digital business.
Reviewer tooling should expose evidence, not decoration
Behavioural analytics are often presented through timelines, distributions and network graphs. Those visualisations should help answer investigative questions. A reviewer needs to see what changed, compared with what, over which period, using which peer group and with what missing data. A graph that looks sophisticated but hides methodology can slow investigation rather than improve it.
A useful case screen can show a customer timeline, baseline range, peer position, material feature changes, related accounts and the transactions contributing most to the signal. It should allow drill-down to source records. It should also show uncertainty: a peer result based on 20 observations should not look as authoritative as one based on thousands of comparable observations.
Usability testing should involve actual reviewers. Measure whether analysts can understand the signal, find the underlying evidence and document a decision accurately. Speed is useful, but faster closure is not an adequate success metric if quality falls.
Outsourcing and third-party analytics
Banks can use vendors for analytics, data or operational review, but responsibility for the control remains with the institution under the applicable outsourcing and AML framework. A vendor may provide an anomaly engine while the bank still owns customer segmentation, thresholds, investigation and reporting decisions.
Due diligence should cover data security, methodology transparency, change notification, service resilience and evidence retention. The bank should understand what happens when the vendor changes a model version or feature definition. Silent vendor changes can make historical outputs impossible to reconstruct.
If review work is outsourced, the institution should define which cases can be handled externally, what evidence reviewers can access, what decisions require bank approval, how quality is sampled and how jurisdictional restrictions on data access are respected. Quality measures should not rely only on throughput. Disagreement analysis, evidence quality and escalation appropriateness are more informative.
Exit planning is part of control design. The bank should be able to retain evidence, model or configuration history required for regulatory and investigative purposes and continue critical monitoring during migration. A provider change should not create a blind period.
Capacity, backlog and alert economics
Behavioural controls compete for finite investigative capacity. Capacity should influence operational design, but queue pressure should not be allowed to redefine risk silently. Lower alert volume can reflect better targeting, or it can reflect reduced coverage.
Capacity planning should examine alert arrival, complexity, ageing, priority and reviewer availability. A surge in a high-risk typology may justify temporary reallocation or triage changes. Persistent backlog may require scenario redesign, additional staffing or process improvement. The response should be documented and risk-based rather than an undocumented threshold increase made solely to reduce volume.
Operational metrics can include ageing by priority, rework, escalation rate, quality-review findings and investigator hours. These are not direct measures of crime detection, but together they reveal whether the control can process its own signals responsibly.
Cross-border behaviour and group-wide views
A multinational customer can look different in each jurisdiction. Local products, currencies and payment customs vary. A global peer model may therefore be misleading if it ignores local context. Equally, purely local monitoring can miss patterns that become visible only when activity is viewed across group entities.
Group-wide analysis should respect legal constraints on data sharing. Where full consolidation is not permitted, institutions may need federated or summary approaches, formal information-sharing channels or escalation between entities. The architecture should reflect what is legally and operationally possible rather than assuming unrestricted global data access.
Correspondent banking creates another limitation because the bank may not know the underlying respondent customer as well as it knows a direct customer. Peer or behaviour comparisons should reflect that data boundary. Confidence should follow visibility.
The central lesson across these advanced cases is that behavioural monitoring is not merely an algorithm. It is an operating system connecting customer understanding, data engineering, analytics, review capacity, governance and legal constraints. Weakness in any of those layers can turn a mathematically sound comparison into a poor banking control.
Practice close: requirements, acceptance criteria and testing
This section converts behavioural monitoring into delivery artefacts. It is deliberately written for BAs, product owners, architects, data engineers, testers, control owners and investigators who need to prove that a comparison control works as designed rather than merely confirm that a model or screen exists.
BA checklist for self-history requirements
A self-history requirement should identify what is being compared and over what period. "Detect unusual behaviour" is too vague to test. A usable requirement might identify transaction count, value, counterparty novelty and corridor change over defined observation windows, specify the minimum history needed, describe treatment of reversals and rejected transactions, and define what happens when the history is incomplete.
The requirement should also define structural-change handling. If a customer profile is materially updated, does the baseline reset immediately, enter a review state or retain the old baseline until an approval? Can the system reconstruct the previous baseline? Are dormancy and newly opened relationships treated differently? These are business decisions, not merely data-science details.
BA checklist for peer comparison
Peer requirements should state which attributes create the segment and why those attributes matter. They should define how small populations are handled, how often assignments can change, whether historical assignments are preserved and what the reviewer sees when a peer result is unreliable or unavailable.
A useful acceptance criterion is not "peer grouping is implemented". It is closer to: given a customer with a specific sector, size band and geography, the system assigns the expected peer segment using the effective-dated profile available at the decision time; when the sector changes, future scoring uses the new segment while historical alerts retain the original assignment.
Statistical checks should match the method. If percentiles are used, test their calculation and boundary behaviour. If a clustering model is used, test versioning, population stability and explainability appropriate to its role. The business requirement should avoid prescribing an algorithm unless the algorithm itself is part of the agreed control design.
Positive, negative and boundary testing
Positive testing should prove that defined risk patterns create the intended signal. Examples include a material break from a stable customer history, rapid movement inconsistent with the prior pattern, a new counterparty network or a customer moving far outside a well-supported peer distribution. The expected result should include the feature values and explanation, not only an alert/no-alert outcome.
Negative testing is equally important. Use legitimate seasonal activity, salary changes, business growth, a new supplier with corroborated commercial context and other benign deviations. A control that detects every deviation but cannot distinguish ordinary change will overload investigators and may harm customers.
Boundary testing should cover just-inside and just-outside values, window cut-offs, time zones, exact duplicate events, reversed payments and peer-group minimum sizes. These cases often reveal implementation defects that ordinary happy-path tests miss.
Cold start, dormancy and structural change
New customers should be tested with insufficient self-history. The expected behaviour may be peer comparison, profile-based monitoring, a lower-confidence output or another defined fallback. The system should not silently manufacture a stable personal baseline from a few events.
Dormant-account tests should verify that stale historical behaviour is not treated as unquestionably current. Structural-change tests should simulate profile updates, acquisitions or employment changes and confirm that re-baselining follows the agreed governance. Historical alerts should remain reproducible after the baseline changes.
Data-quality and failure-mode testing
Behavioural controls depend on joined data. Remove or delay a transaction feed, customer attribute, exchange rate, counterparty mapping or peer benchmark and confirm the system behaves safely. Missing data should not become zero risk unless zero is genuinely the defined value.
Test late-arriving data and replay. If yesterday's payment is loaded today, does it alter the correct historical window? If a duplicate feed is replayed, is activity double-counted? If a customer profile is backdated, does the platform recompute results and, if so, is the recomputation auditable?
Failure-mode testing should also cover unavailable analytical services. The design might queue work, use a controlled fallback or stop scoring until dependencies recover. The correct choice depends on the product, control and risk. What matters is that degraded operation is explicit and monitored.
Testing linked accounts and networks
Create controlled QA data with known relationships: two accounts sharing a household address, multiple accounts sharing a fraud-controlled device, unrelated customers paying the same popular merchant, and a genuine coordinated network. Verify that the system differentiates link types and does not treat every shared attribute as equally strong.
Entity-resolution tests should include false-merge and missed-link scenarios. Historical relationship changes also matter: an address shared last year but not today should retain its effective dates rather than appear permanently current.
Historical back-testing without future-data leakage
A historical test must use the information that would have been available at the historical decision time. Using a customer profile updated months later, a peer group rebuilt from future customers or an investigation outcome that was not yet known contaminates the test.
Create an "as-of" test design: effective-date customer data, peer assignments, feature definitions and model version to the test date. This is particularly important when teams compare a proposed control with a current one. Otherwise, the challenger can appear superior simply because it has information the production control never had.
Reviewer calibration without contaminating production
Reviewer calibration is useful, but synthetic test cases should be handled in a controlled training or QA environment, or through a clearly governed quality-assurance process that cannot be mistaken for genuine customer alerts or regulatory records. Do not secretly inject fabricated customer activity into live production queues merely to test analysts.
Calibration cases can include clear suspicious patterns, clearly legitimate explanations and genuinely ambiguous cases. Scoring should assess the reasoning as well as the final disposition: did the reviewer inspect the baseline, challenge the peer group, test an innocent explanation and document uncertainty appropriately?
Live production outcomes can also support quality review through sampling of real completed alerts and cases, but those outcomes should not be described as perfect ground truth. A closed alert may later prove relevant; a filed SAR or STR is not a criminal conviction.
Customer-impact and fairness testing
Testing should include atypical but legitimate behaviour because peer systems are most vulnerable at the edges of the population. Suitable QA personas can represent irregular income, seasonal work, international family support and unusual business models without assuming that any demographic group is inherently higher or lower risk.
Where a bank wants to test outcomes across protected characteristics, legal and privacy teams should determine what data may lawfully be collected and used in the relevant jurisdiction. There is no universal AML rule requiring one global disparate-impact metric.
Useful customer-impact indicators can include repeated information requests, restriction rates, complaint themes, reopened cases and concentrations of false positives in particular product or customer segments. They should be interpreted with risk context rather than used mechanically.
Sign-off evidence
Before production approval, the delivery team should be able to show the control objective, source-to-feature lineage, peer methodology, data sufficiency rules, versioning, test evidence, known limitations, reviewer explanation, fallback behaviour and governance approvals appropriate to the bank's framework.
The strongest final acceptance test is reconstructability. Choose a test alert and prove that another qualified reviewer can reproduce the customer context, baseline, peer assignment, feature values, logic version and downstream decision from retained evidence. If that cannot be done, the control is not yet operationally transparent even if its aggregate performance looks strong.
Masterclass: when the peer group is wrong
This is a fictional composite training case designed to show how a technically functioning behavioural control can still produce a poor outcome when the customer is compared with the wrong population. The customer, amounts, sequence and decisions are illustrative. They are not taken from an ombudsman decision, enforcement case or regulator finding.
The customer
A self-employed visual artist has banked with the institution for six years. The account was opened when most income came from domestic freelance design work. The original profile therefore records irregular but modest local credits, routine living expenses and occasional international purchases. Over time the artist becomes represented by galleries in several countries, sells work at exhibitions and pays specialist framers, shippers, studio suppliers and collaborators.
The bank's customer profile is not refreshed promptly. Its behavioural platform continues to assign the customer to a broad retail peer group dominated by salaried customers. Several genuine changes now look extreme against that benchmark: large but irregular foreign gallery credits, prolonged low-balance periods between exhibitions, cross-border payments to new professional counterparties, and occasional cash deposits from events where cash sales remain possible.
None of those facts proves legitimacy. Art markets can present financial-crime risk, and unusual cross-border or cash activity deserves proper scrutiny. The defect in this case is different: the comparison layer treats distance from the retail median as if the distance itself were evidence of criminality.
First alert: correct signal, weak context
The first material alert is created after a large gallery credit from another country followed by several payments to a shipper, studio supplier and collaborator. The self-history lens correctly identifies a change. The peer lens also shows strong deviation, but the peer assignment is poor.
A well-designed review would ask whether the customer's occupation, expected income sources and transaction pattern have changed. It would examine available gallery contracts, invoices, tax or business records where proportionate, counterparty information and the broader flow of funds. It would also assess whether the activity contains independent risk indicators beyond being unusual.
Instead, the reviewer sees a high behavioural score, several new foreign counterparties and a peer-percentile chart. The case notes describe the pattern as "inconsistent with peers" without recording why those peers are appropriate. The alert is closed with enhanced monitoring because the reviewer cannot establish suspicion but remains uncomfortable with the score.
Second alert: repetition is mistaken for corroboration
Several months later, another exhibition creates the same pattern. A second reviewer sees the earlier alert and interprets repeated alerts as stronger evidence. This is a common analytical trap. Repeated outputs from the same flawed comparison assumption are not independent corroboration. If the customer remains in the wrong segment, the system can reproduce the same false signal indefinitely.
The second review should therefore distinguish between new evidence and repeated measurement. The prior alert is relevant history, but the reviewer still needs to ask whether the underlying baseline and peer assignment were valid. A strong case-management interface should expose that information instead of displaying only prior dispositions and risk scores.
Customer information changes the picture
The customer is contacted through the bank's normal information-request process. The response includes gallery representation details, exhibition schedules and supporting documentation for several major credits and payments. Some counterparties can be independently corroborated. The activity also aligns with the timing of exhibitions.
That evidence does not mean every future transaction should be automatically trusted. It means the bank now has better information about the relationship. The customer profile should be refreshed, and the monitoring team should assess whether the baseline and peer assignment need controlled change. The old history should remain available for investigation; it should not be overwritten as if the earlier customer state never existed.
The bank also discovers an important data problem: occupation and business-profile changes made in the customer system were not being propagated reliably to the segmentation service. The issue is therefore not only an analyst judgement problem. It is a data-lineage and integration defect.
A proportionate redesign
The remediation does not create an "artist exemption" from monitoring. That would simply replace one weak rule with another. Instead, the bank improves the mechanics that should apply across many atypical customers.
First, peer assignment becomes effective-dated and explainable. Reviewers can see the current segment, the attributes that caused the assignment and previous assignments. Second, material customer-profile changes trigger reassessment of behavioural segmentation rather than waiting for periodic batch refresh. Third, the alert view separates self-history deviation from peer deviation so an analyst can see whether both lenses agree. Fourth, the workflow requires a reviewer to record whether the peer group is considered appropriate when peer deviation materially influenced escalation.
The bank also introduces quality sampling focused on customers with repeated behavioural alerts but weak corroborating risk evidence. The purpose is not to suppress alerts simply because they repeat. It is to identify whether persistent noise points to bad segmentation, stale customer data, poor feature design or a genuinely unresolved risk.
What the case teaches
The case shows why behaviour analytics must preserve the distinction between unusual, unexplained and suspicious. The first is an observation. The second is an investigation state. The third is a judgement that may have legal reporting consequences depending on the jurisdiction and facts.
It also shows why fairness and effectiveness are connected. A bad peer group wastes investigator time, creates customer friction and can still make the bank less safe because analysts become accustomed to clearing noisy alerts. Better segmentation is not merely a customer-experience improvement; it improves the signal that scarce investigative capacity receives.
For BAs and testers, the case translates into concrete acceptance scenarios. Update occupation and business-profile data and verify the change reaches the segmentation service with the correct effective date. Recreate the same transaction pattern before and after peer reassignment and confirm the comparison output changes as designed without rewriting historical alerts. Verify that the case view shows both self-history and peer components. Test that missing peer data produces an explicit degraded state rather than a misleading zero-risk value. Confirm that investigators can reconstruct which profile, peer group and logic version applied when each historical alert was generated.
The final lesson is deliberately modest: peer comparison can be useful, but it is only as good as the customer context, segmentation logic, data lineage and investigation process surrounding it. A bank should never treat mathematical distance from a peer norm as a substitute for understanding the customer and the actual movement of funds.
Knowledge check and glossary
Use these questions to test whether the behavioural-monitoring concepts can be applied without turning statistical deviation into a conclusion of suspicion.
Why use both self-history and peers? Self-history explains whether the customer changed relative to their own past. Peers add context where history is thin or where a change may be common among genuinely similar customers. Neither lens is proof on its own.
Does FATF require peer-group analytics? FATF requires risk-based ongoing monitoring and scrutiny of transactions against the institution's knowledge of the customer. FATF's banking-sector guidance recognises peer-group comparison as one possible technique and expects segmentation criteria and parameters to be documented. It does not prescribe one universal peer algorithm, threshold or statistical method.
What makes a peer group defensible? A clear business rationale, sufficiently comparable members, adequate population size, an appropriate statistical method, known limitations, versioning and a way to identify drift or misclassification.
What should happen when peer data is weak? The result should be suppressed, downgraded in confidence, broadened to an appropriate parent group or routed to contextual review according to the control design. Thin data should not be presented with false precision.
What is the difference between unusual and suspicious? Unusual is an observation relative to an expectation. Suspicious is an investigative or legal judgement under the institution's applicable framework. An alert may remain unexplained for a period without automatically meeting a reporting threshold.
Why can repeated alerts still be misleading? Repeated outputs from the same stale baseline or wrong peer group are not independent corroboration. The bank should test whether repeated signals reflect new risk evidence or a repeated design defect.
What is a structural break? A material change in the customer's behavioural pattern. It may reflect legitimate transformation, criminal use or another cause. Statistical change detection can identify the break; investigation and customer context determine its meaning.
Why is automatic re-baselining risky? If a system automatically learns every new pattern as normal, gradual criminal activity can become embedded in the baseline. Re-baselining should therefore be controlled and historically reconstructable.
How should SAR/STR outcomes be used in model validation? Carefully. A report shows that the institution met a reporting threshold; it is not proof of crime. Likewise, a closed alert is not proof of innocence. Labels should carry their evidential limitations.
What is future-data leakage in back-testing? Using information that was not available at the historical decision time, such as a later KYC update, future peer population or later investigation outcome, to make the historical model look more accurate than it could have been in production.
When is network analysis useful? When relevant behaviour spans accounts through counterparties, devices, addresses, funding sources or timing. Link type and confidence matter because shared attributes can also have innocent explanations.
What should a reviewer see with a behavioural score? Enough explanation to understand what changed, the comparison period, relevant peer group, key contributing features, material missing data and the logic/model version. A number without context is weak investigative evidence.
Does every behavioural analytic need the same model-risk process? No universal AML rule says so. Institutions should apply their own model-risk, technology-risk and financial-crime governance according to materiality, jurisdiction and how the analytic affects decisions.
How should fairness be handled? By ensuring peer groups and features have risk-relevant rationale, examining customer-impact and false-positive patterns, testing atypical but legitimate behaviour and involving legal/privacy teams where protected-characteristic data may be used. Applicable equality and data-protection laws vary by jurisdiction.
Should synthetic calibration cases be hidden in live production queues? No. Use controlled training, QA or clearly governed quality-assurance mechanisms that cannot contaminate genuine customer alerts or regulatory records.
Glossary for delivery teams
Self-history baseline: A representation of the customer's prior behaviour used as a comparison point. It can include value, count, counterparty, geography, channel, timing and sequence features.
Peer group: A population of customers or relationships selected because their relevant characteristics make behavioural comparison meaningful.
Peer assignment: The rule or model that places a customer into a peer segment. It should be versioned and, for historical reconstruction, effective-dated where appropriate.
Deviation: A measured difference between observed activity and a baseline or peer benchmark. It is a signal for analysis, not proof of misconduct.
Cold start: The period in which a new or low-activity customer lacks enough self-history for a stable behavioural baseline.
Structural break: A change indicating that the prior behavioural regime may no longer describe current activity.
Drift: Change over time in the data, population or relationship between features and outcomes that can make an analytical comparison less reliable.
Change point: A statistically or operationally identified point at which behaviour appears to move from one regime to another. The point does not itself explain why the change occurred.
Robust statistic: A measure designed to be less distorted by extreme values, such as a median or percentile-based range in an appropriate use case.
Entity resolution: The process of determining whether records refer to the same person, organisation, device or other entity, with uncertainty and evidence handled explicitly.
Network link: A relationship between entities based on a defined attribute or observed interaction. The meaning depends on link type, strength and context.
Feature lineage: Traceability from the source data through transformations to the derived value used by a rule, score or model.
As-of reconstruction: Recreating a past decision using the data, profile, peer assignment and logic version that were available at that time.
Operational label: A case or alert outcome such as closed, escalated or reported. Operational labels can support analysis but should not automatically be treated as confirmed criminal truth.
Degraded mode: Defined behaviour when required data or analytical services are unavailable, stale or incomplete.
Explainability: Information that enables a reviewer, tester or validator to understand the main factors and data behind an analytical output.
References and further reading
The sources below are public, authoritative references used to frame the chapter. They establish risk-based ongoing monitoring, customer-profile consistency, governance and supervisory expectations. They do not prescribe one universal peer-group algorithm, anomaly model, threshold or investigation workflow; those implementation choices remain institution- and jurisdiction-specific.
Global standards and banking guidance
- Financial Action Task Force (FATF), The FATF Recommendations: https://www.fatf-gafi.org/en/publications/Fatfrecommendations/Fatf-recommendations.html
- Financial Action Task Force (FATF), Guidance for a Risk-Based Approach: Banking Sector: https://www.fatf-gafi.org/content/dam/fatf-gafi/guidance/Risk-Based-Approach-Banking-Sector.pdf.coredownload.pdf
- Basel Committee on Banking Supervision, Anti-money laundering and counter-terrorist financing, Basel Consolidated Guidelines AFS10: https://www.bis.org/committees/bcbs/basel-consolidated-guidelines/module/afs/10
Monitoring effectiveness and innovation
- Wolfsberg Group, Statement on Effective Monitoring for Suspicious Activity, Part I: Moving Beyond Automated Transaction Monitoring: https://wolfsberg-group.org/resources/general/168
- Wolfsberg Group, Statement on Effective Monitoring for Suspicious Activity, Part II: Transitioning to Innovation: https://wolfsberg-group.org/resources/195/202
- Wolfsberg Group, Statement on Demonstrating Effectiveness: https://wolfsberg-group.org/resources/effectiveness/36
Supervisory and jurisdictional material
- European Banking Authority, Guidelines on ML/TF risk factors: https://www.eba.europa.eu/legacy/regulation-and-policy/regulatory-activities/anti-money-laundering-and-countering-financing-1
- US Federal Financial Institutions Examination Council, Customer Due Diligence — BSA/AML Examination Manual: https://bsaaml.ffiec.gov/manual/AssessingComplianceWithBSARegulatoryRequirements/02
- US Federal Financial Institutions Examination Council, Suspicious Activity Reporting — BSA/AML Examination Manual: https://bsaaml.ffiec.gov/manual/AssessingComplianceWithBSARegulatoryRequirements/04
- UK Financial Conduct Authority, Financial Crime Guide (FCG): https://handbook.fca.org.uk/handbook/fcg1/fcg1s1