Alert Triage and Case Prioritisation
A monitoring alert is not a finding of money laundering, terrorist financing or another crime. It is a prompt for review: a rule, model, investigator referral, customer event or external intelligence signal has identified activity that deserves attention. Triage decides how that attention is ordered. When hundreds or thousands of alerts compete for limited reviewer capacity, a bank needs a defensible way to decide what should be looked at first, which alerts can safely wait, which require specialist handling and which need immediate escalation.
That sounds operational, but it is a control-design issue. A bank can have technically sophisticated detection logic and still fail if high-risk alerts sit behind large volumes of routine work, if analysts can select only easy cases, if deadline clocks are not visible, or if the queue cannot distinguish a possible sanctions hold from ordinary post-event AML monitoring. Triage is therefore the bridge between detection and investigation. It translates risk signals into review order without treating the score itself as proof of suspicion.
The central mental model is simple: priority is a decision about review urgency, not a decision about guilt or reportability. A high-priority alert should receive attention sooner because the potential consequence of delay is greater, the evidence is stronger, a time-sensitive obligation may apply, or customer harm may increase while the alert waits. A low-priority alert is not “safe”; it is lower in the current ordering and still needs disposition under the bank’s policy.
Alert, case, suspicion and report are different states
Many control weaknesses start because systems collapse several concepts into one status. An alert is generated because some detection criterion was met. A case is an investigation container that can combine one or more alerts, customers, accounts, transactions and external information. Suspicion is a legal or policy judgement reached when the relevant threshold is met under applicable law. A suspicious transaction or activity report is a regulatory filing made after that threshold is reached where local law requires it.
These states should not be treated as synonyms. A high alert score may justify rapid review but may still end with a documented benign explanation. A low alert score can still become important after enrichment reveals a linked law-enforcement request or a wider network. Likewise, the clock for a regulatory report does not necessarily begin when the monitoring engine generates an alert. For example, US bank SAR rules distinguish a system-generated flag from the later point at which review establishes facts that may constitute a basis for filing. Other jurisdictions use different legal tests and reporting timeframes. Global banks therefore need jurisdiction-aware clock logic rather than one universal deadline field.
A good case model preserves these transitions explicitly. It records when the alert was created, when it entered a queue, when priority changed, when a human or approved automated control made each decision, when a case was opened, when suspicion was formed where relevant, and when any reporting or customer action occurred. That history matters because investigators, QA reviewers, auditors and supervisors need to reconstruct not only the final outcome but also whether the bank acted with appropriate urgency at each stage.
Why risk-based prioritisation is legitimate and necessary
The FATF risk-based approach expects resources and controls to be proportionate to identified ML/TF risks. Wolfsberg’s risk-based approach statement also describes prioritisation of higher-risk customers and activities as a core element of effective financial-crime risk management. This supports risk-based review ordering, but it does not prescribe one mandatory scoring formula, a universal number of priority tiers or a particular queue SLA.
The practical implication is that a bank may design its own prioritisation method, provided the method is consistent with its risk assessment, legal duties and operating model and can be tested. A retail bank with millions of instant-payment events may need different queue logic from a private bank handling fewer, more complex relationships. A correspondent bank may prioritise nested-activity and high-risk corridor signals; a card issuer may have a different split between fraud prevention and AML review. The control should reflect the risk it is meant to manage rather than mimic another institution’s scoring model.
Risk-based prioritisation also has a boundary. Once the applicable legal threshold for reporting has been met, the institution cannot use a lower queue score to delay a filing beyond the local legal timeframe. FATF guidance is clear that risk-based allocation can be used to identify suspicious activity, but where law requires reporting after suspicion is formed, the reporting obligation itself is not optional. Triage therefore belongs mainly before and around investigation prioritisation; it cannot be used to dilute a legal duty that has already crystallised.
What should influence priority
A useful priority model normally combines several dimensions rather than relying on a single amount threshold or customer-risk rating. The exact factors and weights should be documented and validated for the institution.
Detection strength considers why the alert fired. A pattern supported by several independent signals may deserve earlier attention than one weak indicator, but apparent corroboration must be genuinely independent. Five alerts created by the same data defect are not five pieces of evidence.
Customer and relationship risk provides context from KYC, expected activity, beneficial ownership, occupation or business model, country exposure and prior investigations. It should inform priority without becoming a shortcut that assumes every high-risk customer is suspicious or every standard-risk customer is harmless.
Transaction and network characteristics can include value, velocity, flow-through behaviour, rapid onward movement, new counterparties, unusual corridors, cash intensity, device or beneficiary changes where those data are relevant, and links to other alerts or cases. Amount should be interpreted in context. Small-value transactions can form part of structuring, mule activity or terrorist-financing patterns, while a large legitimate transaction may be explainable from the customer profile.
Typology and threat relevance allows the bank to respond to current risks identified through internal intelligence, FIU publications, law-enforcement information, national risk assessments or validated internal experience. A new scam-to-mule pattern or sanctions-evasion technique can justify temporary uplift in priority while evidence is assessed.
Time sensitivity covers legal reporting clocks that have actually started, time-bound law-enforcement or production requests, operational cutoffs for controls that genuinely depend on pre-execution action, and internal service commitments. It is important not to confuse ordinary post-event AML transaction monitoring with sanctions interdiction or fraud controls. Many AML alerts are reviewed after settlement; a payment-scheme cutoff is not automatically an AML deadline.
Potential customer or third-party harm can matter where delay may prolong exploitation, scam victimisation, vulnerable-customer harm or an unnecessary restriction. This factor should be used carefully. A monitoring alert does not itself justify freezing funds or exiting a customer, and customer-impact weighting must remain consistent with applicable law and the bank’s policies.
The model should produce a priority reason that a reviewer can understand. A score such as 87 is not useful by itself. The work item should explain, for example, that priority was raised because the alert involves rapid movement through newly opened accounts, a linked mule-network case and a time-sensitive law-enforcement request. Explainability helps reviewers challenge bad ordering and helps QA distinguish a model error from a reviewer error.
Priority is dynamic, not fixed at alert creation
Alert priority can change as new facts arrive. A customer may be linked to another investigation. A new sanctions designation may affect a connected party. An analyst may identify a victim at risk. A case may reveal that several low-value alerts are one network. Conversely, enrichment may show that an alert is a repeat of a known legitimate pattern and reduce its urgency.
The system therefore needs controlled reprioritisation. Every change should record the old priority, new priority, reason, actor or rule, timestamp and the model or rule version used. The original priority should not be overwritten. Historical reconstruction is essential when reviewing whether the institution acted reasonably with the information available at the time.
Automated reprioritisation can be useful when new data arrive, but governance should define which changes may occur automatically and which require human approval. A model can reorder review urgency; it should not silently convert an alert into a legal conclusion. Material priority changes should be visible to the reviewer and subject to QA sampling.
Deadline control must be jurisdiction-aware
Deadline logic is one of the most error-prone parts of triage because different clocks measure different things. A global bank may have regulatory reporting deadlines, law-enforcement response dates, internal investigation targets, sanctions decision timeframes and customer-service commitments operating in the same case platform. They should not be stored as one generic SLA.
For US banks, the FFIEC BSA/AML manual explains the SAR timing rule and also makes clear that the regulatory period does not necessarily begin when a transaction is first highlighted by an automated system. The bank needs a reasonable review process and must complete it in a reasonable period. In Australia, AUSTRAC’s September 2026 guidance expects reporting entities to prioritise higher-risk suspicious matters more urgently and, once reasonable grounds for suspicion are established, to file within the applicable statutory timeframe. These are jurisdiction-specific examples, not global deadlines.
The architecture should therefore distinguish at least four dates: alert creation, investigation start, the event that starts any statutory reporting clock, and the statutory due date. It should also support multiple concurrent clocks where needed. A countdown should never be calculated from an event simply because it is convenient technically. Business rules need a legal or policy definition of the trigger event.
Pause and resume logic deserves particular scrutiny. Internal service clocks may legitimately pause while approved information is awaited. A statutory reporting clock may not. The case platform should prevent users from pausing a legal clock unless the law or policy explicitly supports it and the reason is recorded. If the bank misses a deadline, the item should remain actionable; lateness does not erase the underlying obligation.
Queue architecture: more than “high, medium, low”
A priority score becomes useful only when the queue can act on it. Queue design should combine urgency, skill, jurisdiction, product and ownership. A high-priority trade-finance alert may need a different reviewer from a high-priority retail mule alert. A multilingual cross-border case may require regional expertise. A correspondent-banking alert may need access to respondent due-diligence information that a retail analyst cannot see.
Skill-based routing is often better than one giant queue. Work can be partitioned into controlled subqueues while still maintaining enterprise visibility. The design should stop local queues from hiding enterprise backlog. Governance needs a consolidated view of all unreviewed alerts, their age, priority and risk context.
Assignment rules also matter. Pooled queues give teams flexibility but can invite cherry-picking if reviewers choose only simple alerts. Named assignment improves accountability but can create bottlenecks when a reviewer is absent. Many institutions therefore use a hybrid: the engine proposes or assigns work based on priority and skill, supervisors can rebalance it, and reviewers cannot silently bypass higher-priority assigned work without a reason code.
The queue should support linking and consolidation. Repeated alerts on the same customer may be better investigated as one case rather than reviewed independently. Network signals may link several customers. But deduplication should not destroy evidence. The case should retain all underlying alert identifiers and explain why they were merged.
Data and system touchpoints
Triage relies on data from multiple systems. Typical inputs include customer and account masters, KYC risk ratings, beneficial ownership, product data, transaction history, channel and device information where relevant, sanctions or PEP screening outcomes, fraud intelligence, prior cases and SAR/STR history where permitted, law-enforcement requests, country risk, adverse information and external threat intelligence.
Data quality problems can distort priority in both directions. A missing customer-risk rating may default to low priority when it should trigger an exception. Duplicated alerts can inflate apparent corroboration. Late transaction feeds can make a quiet account appear suddenly active. Stale beneficial-owner information can hide a higher-risk relationship. The triage layer therefore needs data-quality controls, lineage and exception handling rather than assuming upstream data are complete.
Architecturally, the priority decision should be reproducible. Store input feature values, the policy or model version, reference-data version, calculation timestamp and reason codes. If a model is used, retain enough information to explain the ordering decision under the institution’s model-governance standard. If a rules engine is used, version the rules. A reviewer examining an alert six months later should be able to reconstruct why it was ranked as it was then, not only what today’s model would do.
Automation and machine learning
Automation can enrich alerts, suppress exact duplicates, route cases, calculate deadlines, identify linked activity and rank work. These uses can materially reduce manual effort. However, automated closure or deprioritisation carries higher control risk because errors can remove items from human review altogether.
Any automated disposition should have a defined population, documented rationale, validation, change control and QA sampling. False-negative testing is particularly important. A rule that closes 40 percent of alerts is not successful merely because it reduces workload; the bank needs evidence that the removed population does not contain unacceptable risk.
Machine-learning prioritisation can be useful where there is enough stable data and a clear use case, but historical labels are imperfect. A past SAR/STR filing is not proof that crime occurred, and a closed alert is not proof that activity was benign. Models trained only on previous filing outcomes can reproduce earlier investigative bias and miss new typologies. Validation should therefore look beyond simple “SAR conversion” and include coverage, stability, segment performance, missed-risk sampling and expert review.
Human override remains important. Reviewers should be able to raise or challenge priority where new evidence justifies it, and the system should capture why. Repeated overrides in one direction can reveal a model or rule weakness and should feed tuning governance.
Customer impact without confusing control types
AML transaction monitoring is often post-event, so many alerts do not delay a customer payment at all. Other controls, such as sanctions screening, fraud prevention or enhanced due-diligence restrictions, may affect transactions or account access immediately. A case-management platform may bring these control types together, but the bank should preserve their different legal bases and decision rights.
Where triage does affect a customer outcome, urgency should consider foreseeable harm. A scam victim who is still sending funds, an elderly customer under coercion, or a small business whose account is subject to a time-sensitive review may need earlier human attention. Earlier attention does not mean automatic release or restriction. It means the relevant team should review the case promptly using the correct control and legal framework.
This is also where financial-inclusion principles matter. Risk-based AML controls should be proportionate. Queue design should avoid patterns where whole customer groups systematically wait longer without a defensible risk reason. Monitoring customer impact alongside control effectiveness can identify such unintended outcomes.
From triage to investigation
Triage should end in a clear disposition. Typical outcomes include closure with evidence, request for additional information, escalation to a case, linkage to an existing case, referral to another control team, or immediate escalation where policy requires senior or specialist attention. The exact taxonomy should match the institution.
Escalation criteria should be understandable and testable. An analyst should know what evidence is sufficient to open a deeper investigation and what information must travel with the handoff. The receiving investigator should see the alerts, priority rationale, customer context, transactions, linked parties, prior reviews and any deadline state rather than re-creating the triage work from scratch.
A case may later result in no suspicion, enhanced monitoring, a SAR/STR decision, customer-risk reassessment, relationship restriction or exit, fraud recovery activity, sanctions action, or another outcome. Those downstream results are useful feedback, but they should not be used mechanically as “ground truth” for future priority models. The bank should distinguish regulatory reporting decisions from confirmed criminal outcomes.
Queue health and backlog risk
Backlog is not simply the number of open alerts. Ten thousand low-risk alerts with fresh review dates may present a different problem from two hundred high-priority aged alerts. Queue health should therefore show the distribution of risk and age.
Useful measures include open alerts by priority and age band, oldest alert in each tier, alerts approaching legal or policy deadlines, unassigned items, reassignment rates, case-conversion rates by scenario, QA error rates, manual override rates, reopened cases and reviewer capacity. No single metric proves effectiveness. High closure volume can coexist with poor risk coverage, while a high escalation rate can indicate either strong detection or excessive noise.
Backlog thresholds should lead to action. A bank can define triggers for supervisor review, temporary reallocation, overtime, cross-trained support, model tuning or escalation to governance. Temporary surge actions should not become a substitute for permanent capacity where high volumes are structural.
Capacity planning
Capacity should be planned from expected workload, complexity and quality requirements rather than forcing detection output to match available headcount. The FFIEC examination procedures for US banks explicitly tell examiners to consider staffing adequacy and state that alert and investigation volume should not be tailored solely to existing staffing levels. The broader principle is useful globally even though the exact supervisory rule is US-specific: capacity is part of control effectiveness.
Forecasting should consider planned scenario changes, onboarding growth, product launches, data repairs that may release held alerts, sanctions events, seasonal patterns and known regulatory changes. Productivity estimates should reflect complexity and quality. A reviewer who closes simple alerts quickly is not directly comparable with an investigator handling complex networks.
Cross-training can provide flexible capacity, but specialist cases still need specialist skills. Outsourcing may support defined review populations, subject to due diligence, access controls, training, jurisdictional constraints, quality assurance and clear accountability. The regulated institution retains responsibility for the control even when operational tasks are performed by a service provider.
Roles and governance
The first line or financial-crime operations team normally owns day-to-day queues, staffing, assignment and execution. Compliance or the second line sets or oversees policy, challenges the risk methodology and monitors adherence according to the institution’s governance model. Model-risk or analytics teams may validate machine-learning or statistical prioritisation. Technology teams operate workflow and data services. Internal audit provides independent assurance.
A named queue owner should have authority to respond when priority logic or capacity is not working. Governance forums should review material backlog, breaches, high-risk ageing, model or rule changes, QA findings, customer-impact indicators and remediation. Changes to weightings or auto-disposition rules should follow controlled approval and testing rather than being made ad hoc to reduce volumes.
Documentation should explain the rationale for the prioritisation framework, the meaning of each tier, applicable clock rules, override permissions, deferral rules, escalation routes and contingency arrangements. Reviewers need concise operating guidance; governance and audit need the underlying methodology and evidence.
Failure modes to look for
A useful audit of triage looks for failure patterns rather than only checking whether a priority field exists.
One failure is amount dominance: high-value payments always rank above low-value activity, causing mule networks or structuring to wait. Another is risk-rating dominance: high-risk customers always rank first even when stronger evidence exists elsewhere. Duplicate corroboration occurs when several alerts generated by the same logic are counted as independent evidence. Queue hiding occurs when work is moved into local or deferred queues and disappears from enterprise backlog reporting.
Cherry-picking occurs when analysts select easy items while complex work ages. Clock confusion occurs when internal SLAs are mistaken for statutory deadlines, or when a legal clock starts from the wrong event. Label leakage occurs when model training uses later investigation outcomes that would not have been available at the time of triage. Silent reprioritisation occurs when the system overwrites priority without preserving the earlier decision. Capacity masking occurs when scenarios are tuned primarily to reduce workload rather than to improve risk coverage.
None of these failures is solved by adding another score. They require governance, data quality, workflow design, incentives, testing and management information that show whether the queue is behaving as intended.
BA, architecture and testing considerations
For a business analyst, the most important requirement is to define decision semantics before fields. “Priority”, “severity”, “risk”, “SLA” and “deadline” are often used interchangeably until implementation, when ambiguity becomes defects. Requirements should define who sets each value, what inputs are used, when it may change, what it controls and how it is audited.
The data model should include immutable alert identity, customer and account links, scenario or source, creation time, priority, priority reason codes, priority version, jurisdiction, current owner, queue, deadlines by type, case links, disposition, and audit history. Multi-entity banks also need legal-entity context and access controls for cross-border information sharing.
Architecture should separate the detection engine from the workflow engine where practical. Detection decides that something is unusual enough to generate an alert; workflow decides where it goes and when it should be reviewed. This separation makes queue logic easier to change without altering the detection methodology and makes responsibilities clearer.
Testing should include positive and negative routing tests, boundary tests around priority thresholds, missing-data tests, duplicate-event tests, daylight-saving and timezone tests for clocks, reprioritisation history, permission tests, failover behaviour, and high-volume performance. Model changes should be replayed on historical or synthetic populations before activation. Where shadow testing is used in production, it should not alter live customer or regulatory outcomes until governance has approved the change.
User-acceptance testing should verify not only that an alert appears in a queue but that the reviewer can understand why it is there, see the correct context, identify deadlines, link related alerts, escalate it and reconstruct every material change. QA teams should also test management information against underlying queue populations so dashboards cannot exclude deferred, unassigned or failed-routing items unintentionally.
Mini case: the low-value mule cluster
A retail bank’s priority model heavily weights transaction amount. One week, a new mule typology produces a cluster of low-value inbound credits followed by rapid transfers to newly added beneficiaries. Each payment is below the bank’s high-value threshold, so the alerts land in a standard queue even though network analytics link twelve recently opened accounts using overlapping devices and beneficiaries.
A separate fraud referral identifies one of the beneficiaries as connected to an active scam investigation. The case platform receives the referral, raises the priority of all linked alerts and routes them to a network-investigation team. The investigator confirms that the alerts share genuine independent signals: account age, rapid pass-through activity, beneficiary overlap and fraud intelligence. The bank opens a consolidated case and reviews whether suspicious-activity reporting, customer protection, account controls and law-enforcement liaison are required under the relevant jurisdiction and policy.
The important lesson is not that every linked cluster should become high priority. It is that the original model treated transaction value as a stronger signal than network context without evidence that this ordering was effective. The remediation changes the priority methodology, adds network linkage as a controlled input, introduces QA sampling for aged low-value clusters and documents why each priority factor is used.
Key takeaways
Alert triage is the discipline of ordering review effort, not declaring suspicion. The strongest designs combine risk context, evidence strength, time sensitivity, customer impact, reviewer skill and capacity while preserving the legal distinction between an alert, a case, suspicion and a report.
A defensible triage framework is explainable, versioned and testable. It preserves history, uses jurisdiction-aware deadlines, shows all backlog honestly, prevents silent cherry-picking, supports human challenge and treats automated prioritisation as a governed control rather than an unquestioned answer.
Most importantly, the bank should be able to demonstrate that its queue design helps the right people review the right work at the right time without inventing universal thresholds or using operational convenience to override legal duties. The next chapters build on that foundation by examining false positives, practical red flags and the investigation decisions that follow.
Operational deep dive: priority logic, clocks and capacity
The base chapter explains what triage is for. This deep dive focuses on the parts that most often fail in implementation: combining unlike risk signals into one queue order, handling deadlines without creating false legal assumptions, and planning reviewer capacity without tuning detection merely to fit available headcount.
Designing a priority score without pretending it is truth
A composite priority score is an operational ranking device. It can combine scenario severity, customer risk, transaction behaviour, typology relevance, network context, time sensitivity and customer impact, but the score should never be presented as a probability that a customer is laundering money. The inputs are imperfect proxies and the labels available for calibration are usually imperfect as well.
A practical design starts by defining each factor in business language. For example, a scenario-strength factor might distinguish a single threshold breach from a pattern supported by velocity and network links. A customer-context factor may consider KYC risk, but should not duplicate the same geographic or product risk already embedded in the scenario. A time-sensitivity factor should identify the type of clock involved rather than simply add points because an alert is old.
Normalisation can be useful when factors have very different scales, but the maths should remain explainable. Amounts are often heavily skewed, so a logarithmic or banded treatment may be more sensible than a linear rule that lets one large transaction dominate every other signal. Binary factors can represent genuinely discrete events such as a validated law-enforcement request. Interaction effects should be introduced only where there is evidence that the combination matters and reviewers can understand the rationale.
Calibration should use multiple outcome measures. Case escalation, SAR or STR filing, confirmed fraud intelligence, law-enforcement feedback, QA findings and expert re-review can all provide evidence, but none is perfect ground truth. A filed SAR or STR records a suspicion decision, not a criminal conviction. A closed alert records an operational decision, not proof of innocence. Models trained on these outcomes need explicit acknowledgement of label limitations.
A strong validation asks whether the ranking separates urgency usefully, whether rare but important typologies remain visible, whether the model behaves consistently across customer and product segments, and whether false-negative sampling identifies unacceptable blind spots. A flat escalation rate across all priority tiers can be a warning sign, but it is not automatically proof that the model is wrong because some tiers may be designed around time sensitivity or customer protection rather than report conversion.
Dynamic reprioritisation
Priority should be recomputed when new evidence materially changes urgency. Common triggers include a new linked case, updated KYC information, a newly identified victim, a law-enforcement request, a sanctions event, a fraud referral or new network intelligence. Reprioritisation rules need clear ownership and version control.
The system should preserve the original score and every subsequent score. An audit trail should show which inputs changed and whether the change was made by a rule, model, reviewer or supervisor. Silent overwriting creates a false history and makes later review of timeliness impossible.
Manual overrides should require a reason code and, for material changes, free-text rationale. Override analysis is useful feedback. If experienced reviewers repeatedly raise one alert type, the model may underweight that risk. If they repeatedly reduce another, the design may be generating operational noise. The response should be model or rule review, not automatic assumption that either the model or the reviewer is correct.
Deadline architecture
One of the most important design choices is to keep different clocks separate. A bank may have internal triage targets, regulatory reporting deadlines, law-enforcement response dates, sanctions decision deadlines, fraud-recovery windows and customer-service commitments. The workflow engine should know which clock it is displaying and what event started it.
In the United States, FFIEC guidance for banks explains that the SAR filing period is linked to the initial detection of facts that may constitute a basis for filing, not simply to the moment an automated system first flags a transaction. In Australia, AUSTRAC guidance current in September 2026 expects higher-risk suspicious matters to be prioritised more urgently and requires filing within the applicable statutory period once the relevant suspicion threshold is met. These examples illustrate why a global platform should never hard-code one universal reporting clock.
The data model should therefore store a deadline type, jurisdiction, trigger event, trigger timestamp, due timestamp, calculation rule version and any permitted pause reason separately. Internal SLAs can often be paused for approved reasons; statutory clocks may not be pausable. That distinction belongs in requirements and access control, not reviewer memory.
Deadline escalation should surface work before breach. A typical pattern is warning, supervisor escalation, senior escalation and incident handling, but the exact stages are institution-specific. If a deadline is missed, the case remains active and the system should record the breach, late action and root cause. A missed deadline is not a reason to suppress the filing or investigation.
Capacity modelling
Capacity is part of control effectiveness because even a good queue fails if work cannot be reviewed with appropriate quality. Forecasting should estimate volume by priority and complexity, not only total alert count. A scenario producing 10,000 simple alerts may require less effort than 1,000 complex network cases.
The demand model should include planned scenario changes, new products, customer growth, data repairs, seasonal effects, external events and known regulatory changes. Reviewer supply should consider productive hours, training, leave, QA rework, specialist skills and supervision. Raw closures per analyst are a poor productivity measure if the work mix differs materially.
Quality-adjusted productivity can use complexity bands and QA results to avoid rewarding fast but shallow review. It should still be treated as an operational planning measure rather than an exact scientific quantity. The aim is to make staffing assumptions transparent enough for challenge.
Leading indicators can trigger action before a backlog becomes critical: rising high-priority age, increasing unassigned work, sustained overtime, deteriorating QA results, missed internal targets or growing dependence on temporary staff. Temporary surge capacity may be appropriate for an event-driven spike. Repeated use of surge arrangements during normal volumes points to a structural capacity problem.
Queue simulation and shadow testing
Before changing priority logic, teams can replay historical or synthetic alerts through the candidate model. The replay should use data available at the time of each alert to avoid look-ahead bias. Comparing old and new ranking can reveal which populations move and why.
Simulation is useful but cannot reproduce every behaviour of live operations. Reviewer selection, shift patterns, unexpected data gaps and new typologies may change outcomes. For material changes, shadow mode can run the new prioritisation alongside production without affecting live ordering. Differences can then be sampled and reviewed before cutover.
Any production experiment needs stronger governance than an ordinary product A/B test because queue ordering can affect legal timeliness and customer outcomes. Institutions should not deliberately expose one group to materially weaker financial-crime controls for experimentation. Where controlled comparison is used, safeguards, approval and rollback criteria should be explicit.
Fairness and proportionality
Priority models can create unintended disparities if they rely on proxies correlated with protected characteristics or if historical outcomes reflect earlier investigative bias. Fairness review should therefore examine whether materially different waiting times or escalation rates arise across customer segments and whether those differences have a defensible risk basis.
This does not mean every segment must have identical outcomes. Different products and risk exposures can justify different treatment. The control question is whether the difference is proportionate, evidence-based and consistent with applicable law and policy. Features with weak financial-crime relevance should not be retained merely because they improve historical prediction.
FATF’s 2025 changes to Recommendation 1 strengthened the emphasis on proportionality and financial inclusion. That makes it especially important not to equate risk-based prioritisation with blanket exclusion or indefinite delay for whole customer groups.
Practical review questions
When reviewing a triage model, ask: Can a reviewer explain why the alert is in this tier? Are the factors independent enough to avoid double-counting? Does the system preserve the version of the logic used? Are statutory and internal clocks separated? Can new evidence raise priority quickly? Are low-value network patterns visible? Are deferred alerts still visible to governance? Does capacity planning reflect complexity? Can the bank show false-negative testing and QA results, not only closure volume?
If those questions cannot be answered from system data and governance evidence, the control is not yet mature even if the queue looks orderly.
Advanced practice: incentives, specialist routing and surge governance
Triage can be technically well designed and still fail because people respond to the way work is assigned and measured. Queue governance therefore has to consider reviewer incentives, expertise, temporary volume shocks, third-party support and the difference between an internal service target and a legal deadline.
Preventing cherry-picking without blaming reviewers
If analysts are rewarded mainly for the number of alerts they close, they have a natural incentive to select simple work. The resulting behaviour can look like strong productivity while complex alerts age in the background. A sound control measures whether assigned higher-priority work is actually being reviewed, whether difficult cases are concentrated on a small number of people, and whether analysts repeatedly bypass work without a documented reason.
Selection-pattern analysis can compare the mix of alerts a reviewer receives with the mix they complete. That analysis should be interpreted carefully. A reviewer may legitimately handle fewer alerts because the assigned work is complex, because they are supporting an investigation, or because they have a specialist role. The objective is not to create another simplistic performance score but to identify persistent patterns that warrant supervisory review.
Where cherry-picking is driven by the operating model, the first response should be to correct the operating model. Named assignment, supervisor-controlled rebalancing, complexity-adjusted workload measures and quality-weighted performance objectives can reduce the incentive to chase easy closures. Persistent deliberate avoidance after those controls are in place can then be treated as an individual performance or conduct issue under the bank’s policies.
Performance measures that support control effectiveness
No single metric should determine analyst performance. Useful measures can include completed work, timeliness, QA accuracy, quality of written reasoning, appropriate escalation, adherence to procedures and contribution to complex cases. Downstream case or SAR/STR outcomes may provide useful feedback but should not become a crude target. A reviewer should not be incentivised to escalate weak cases simply to improve a conversion measure.
Team-level metrics can reduce individual gaming, but they also need safeguards. A high-performing team can still have one queue or customer segment that receives poor attention. Management information should therefore allow drill-down by priority tier, scenario, product, legal entity, geography and age.
Governance should review incentives whenever a material process or technology change alters the nature of the work. Introducing automated enrichment, for example, may change the time needed per alert. Holding analysts to an old closure target after the work has become more complex can recreate the same pressure in a different form.
Skill-based routing
Priority answers when work should be reviewed; routing answers who should review it. These are related but separate decisions. A high-priority alert can still be mishandled if it reaches a reviewer without the right product, language, jurisdiction or typology knowledge.
Skill profiles can identify capabilities such as correspondent banking, trade finance, retail mule investigations, virtual assets, sanctions, complex ownership, particular languages or regional law. The workflow engine can use those profiles together with priority and workload to propose assignments. Supervisors should retain authority to reassign when real-world context requires it.
Skill matrices should not become permanent silos. Cross-training and supervised rotation can build resilience and reduce key-person dependency. More junior analysts can participate in complex reviews under appropriate supervision while senior specialists remain accountable for the decision. The bank should monitor whether certain skills are consistently overloaded because that can be an early capacity warning.
Outsourced or third-party review
Banks may use service providers for parts of alert review, especially where volume is high. Outsourcing does not transfer the regulated institution’s accountability for the control. The bank needs due diligence, contractual standards, access controls, training, quality assurance, incident management and an exit plan consistent with applicable outsourcing and data-protection requirements.
Scope should be based on capability and risk rather than a blanket rule that all low-priority work can be outsourced. A provider may be well equipped for some specialised work and poorly equipped for some routine-looking alerts that depend on internal customer context. The bank should decide which populations are suitable based on evidence, not merely labour cost.
Quality sampling should compare provider decisions with internal standards and track the direction of disagreements. Recurring under-escalation, weak narratives, missing evidence or misunderstanding of local requirements should trigger remediation. Changes in provider staff, location or subcontracting can affect risk and should be governed under the bank’s third-party framework.
Surge governance
Alert volumes can rise suddenly after a scenario change, a data repair, a major fraud event, a new sanctions designation, a system catch-up or an external intelligence release. A surge plan should define how the bank will protect higher-risk work while restoring normal service.
Triggers can include sustained volume above forecast, high-priority ageing, unassigned work or a material external event. Response options may include temporary reallocation of trained staff, extended coverage, supervisor-led prioritisation, pausing non-critical internal activities or using pre-approved third-party capacity. Any simplified review process should have clear boundaries and should not weaken legal or policy requirements.
A surge plan should also define an exit condition. If the same emergency measures remain active for months, the problem is no longer a surge; it is a structural capacity gap. Governance should then decide whether permanent staffing, automation, scenario redesign or another long-term change is required.
Post-surge review should sample decisions made during the period, assess whether high-priority or deadline-sensitive work was protected, identify customer impact and record lessons for the next event. The purpose is not to prove the surge plan worked but to discover where it did not.
Internal service levels versus legal deadlines
Priority tiers often have internal targets such as review within a certain number of hours or days. Those targets help manage workload but should not be described as regulatory deadlines unless a specific law, rule or official requirement creates the timeframe.
A bank can set tighter internal targets than the law requires. It may also set different targets by alert type, risk and product. Changes to internal SLAs should follow governance and be tested against capacity. Repeated SLA breach can indicate inadequate resources, poor routing or unrealistic design, but it does not automatically mean a regulatory breach.
Legal clocks should be separately identified and protected. Where a reporting obligation has arisen, applicable local law determines the due date. The workflow should display both the internal target and the legal deadline where both exist so users do not confuse them.
Management information that reveals rather than reassures
A useful governance pack shows more than total open alerts and average age. It should allow leaders to see the distribution of age by priority, approaching deadlines, unassigned work, deferred items, complex-case load, QA errors, manual overrides, system failures and material customer impact.
Averages can hide long tails. An average queue age of two days can coexist with a small number of critical items waiting several weeks. Percentiles, age bands and explicit high-priority oldest-item measures are more informative.
Metrics should also disclose scope. If one business line, country or deferred queue is excluded, the dashboard should say so. A control cannot be governed honestly if material work is invisible because it sits outside the metric definition.
The strongest triage operating model is therefore not the one with the most sophisticated score. It is the one where ordering, people, incentives, skills, capacity and governance reinforce each other and where management information makes failure visible early enough to act.
Practice close: requirements, controls and testing
This section turns the chapter into delivery artefacts that a business analyst, product owner, architect, tester or control owner can use when designing or reviewing a triage capability.
Requirements checklist
The business requirement should define what “priority” means before specifying a field or score. It should identify the factors that may influence priority, the source system for each factor, who owns the methodology, when priority can change, whether a reviewer can override it and what evidence must be retained.
Deadline requirements should list each deadline type separately. For every clock, document the applicable jurisdiction or policy, the event that starts the clock, the calculation rule, the due-date field, whether pause is permitted, who can change the date and what escalation occurs before or after breach. Avoid one generic slaDueDate if the process contains both internal targets and legal reporting deadlines.
Queue requirements should cover routing by priority, jurisdiction, product and skill; assignment and reassignment; unassigned work; deferral; linkages between alerts and cases; supervisor intervention; business-continuity arrangements; and the management-information view across all queues. Deferred work should remain visible and have a review-by date.
Audit requirements should preserve the inputs and version of the prioritisation logic, every priority change, assignment history, deadline changes, case links, disposition and user or automated actor. Historical reconstruction should be possible without recalculating the case using today’s rules.
Acceptance criteria examples
Good acceptance criteria are observable. Examples include:
- When a validated linked-case signal arrives, the alert is reprioritised according to the approved rule and the previous priority remains in the audit history.
- When required source data are missing, the alert follows the defined exception path rather than silently defaulting to a lower priority.
- A reviewer cannot pause a statutory deadline unless the configured jurisdictional rule permits it and the permitted reason is recorded.
- Deferred alerts remain in enterprise backlog reporting with their original creation date, current review-by date and deferral reason.
- A supervisor can identify the oldest unreviewed item in every priority tier and trace it to the underlying alert.
- The user interface explains the main reason for priority sufficiently for a trained reviewer to challenge the ordering decision.
- Related alerts consolidated into a case retain their original identifiers and disposition history.
Criteria such as “queues are monitored” or “alerts are prioritised correctly” are too vague for objective sign-off.
Test strategy
Functional testing should verify score and rule calculations, routing, assignment, reprioritisation, manual override, case linkage, deferral and closure. Boundary tests should cover thresholds and tier transitions. Negative tests should confirm that incomplete or malformed inputs follow controlled exception handling.
Clock testing should cover time zones, daylight-saving changes where relevant, weekends and holidays where the rule uses them, missing trigger events, duplicate trigger events and late-arriving data. Testers should verify that internal SLA clocks and statutory clocks remain distinct in both the user interface and reporting layer.
Integration tests should trace a transaction or customer event from source through the detection engine, enrichment, priority decision, workflow and case system. Data-lineage checks should confirm that the feature value seen by the priority engine matches the authoritative source at the relevant time.
Volume and resilience testing should simulate expected and stressed alert loads. The objective is not only response time. Test whether high-priority items continue to route correctly, whether deadline escalation still runs, whether retry logic creates duplicates, and whether failed enrichment produces a visible exception instead of an apparently complete low-priority alert.
Model or rule changes should be evaluated first using historical or synthetic populations in a controlled environment. Shadow testing in production can compare a candidate priority against the live priority without changing review order. Any test that could alter a live customer outcome, legal deadline or regulatory decision requires explicit governance, segregation and rollback controls; unguarded seeded cases should not simply be mixed into live queues.
QA and ongoing assurance
QA sampling should include more than completed alerts. A risk-based sample can include aged alerts, deferred items, manual overrides, automatic deprioritisations, high-priority closures and alerts affected by data-quality exceptions. This helps detect blind spots that completed-case sampling alone cannot see.
QA should distinguish the error source. A reviewer may make a poor decision even when the priority was correct. The priority model may order work poorly even when the reviewer handles it correctly. Upstream data may be wrong. Routing may send work to an unsuitable team. Remediation depends on identifying the right layer.
Independent validation should periodically challenge the methodology, especially after material changes in typology, product mix, customer base or data. Internal audit may then assess whether governance, change control, evidence and execution operate as designed.
Management information for sign-off
Before go-live, confirm that governance can see at least: alert volumes by priority, age bands, oldest items, approaching deadlines, unassigned and deferred work, manual overrides, routing failures, QA results and capacity indicators. The exact dashboard can vary, but material populations should not disappear because they sit in a different queue or operational status.
Management information should be reconciled to the workflow population. If the operational dashboard says 10,000 alerts are open, teams should be able to explain how that number relates to source alerts, cases, deferred work and exclusions. Reconciliation is a control, not merely a reporting exercise.
BA questions for design workshops
A useful workshop should be able to answer these questions in plain language: What makes one alert more urgent than another? Which inputs are risk signals and which are legal clocks? What happens if a required input is missing? Who can change priority? Can reviewers bypass assigned work? How are linked alerts consolidated? What prevents low-value networks from remaining low priority forever? Which team owns the queue when systems fail? How is customer harm considered? What evidence will prove the design still works six months after go-live?
If the answers depend on individual memory or undocumented judgement, the requirement is not complete. The goal of the design is not to eliminate human judgement but to make the conditions, evidence and accountability around that judgement visible and testable.
Masterclass: when closure targets distort the queue
This is a fictional composite case built for training. It combines common control weaknesses seen across financial-crime operations: closure-focused productivity measures, pooled queues, weak network context and management information that shows throughput more clearly than risk. The organisation, events and figures are illustrative and do not describe a particular enforcement action.
The operating model
A retail bank runs transaction monitoring through several pooled queues. Alerts receive a priority score, but analysts can choose which item to open next. Team performance is discussed mainly in terms of alerts closed per day and the percentage of work completed inside an internal SLA. Quality assurance reviews a sample of completed alerts for procedural accuracy.
The model appears efficient. Closure volumes increase, the average age of open alerts remains stable and most completed items meet the internal target. But the metrics do not show the complexity of the work selected by each analyst, the age distribution within each priority tier or the alerts that remain repeatedly unselected.
At the same time, the bank begins receiving alerts on newly opened accounts showing rapid inbound and outbound movement. Individual transactions are relatively small. Several alerts require manual linkage across customers and beneficiaries, so they take longer to investigate than simple threshold alerts. Reviewers under closure pressure naturally favour work that can be resolved quickly.
What the queue hides
Over several review cycles, the complex alerts age. Nothing in the process explicitly instructs reviewers to ignore them. The weakness comes from the combination of pooled selection and performance measurement. The system ranks the alerts, but it does not enforce assignment or require a reason when higher-priority work is bypassed.
A fraud intelligence referral later identifies one beneficiary used by several of the accounts. Network enrichment then links additional customers through shared beneficiaries and device information. The bank realises that alerts previously treated as separate low-value events form a wider mule pattern.
The control failure is not that every earlier alert should automatically have been escalated. The failure is that the operating model allowed more complex work to wait without visible challenge, while management information interpreted high closure volume as healthy control performance.
The investigation
The bank opens a consolidated investigation. Reviewers reconstruct the alert timeline and compare the priority assigned to each item with the order in which work was actually completed. They find repeated examples where easier, lower-priority items were selected before older, more complex alerts.
QA then re-reviews a risk-based sample of aged work. Some alerts close correctly after deeper analysis; others require escalation. The exercise demonstrates why backlog sampling is important: an unreviewed alert has no final outcome, so completed-case QA alone cannot measure what the queue is missing.
The investigation also examines the priority methodology. Transaction amount carried more weight than network context. That design was not necessarily unreasonable when introduced, but the bank had not recalibrated it after mule activity changed. New fraud intelligence and linkage data were available but were not integrated into the triage decision promptly.
Remediation
The bank changes several parts of the operating model rather than simply increasing the priority score.
First, high-priority work moves to controlled assignment. Analysts can request reassignment, but bypassing an assigned item requires a reason that supervisors can review. Second, performance measures add complexity and quality so reviewers are not rewarded for choosing the easiest work. Third, management information shows age by priority tier, unassigned and deferred work, and the oldest high-priority items instead of relying on one average age.
The bank also introduces network-linkage information into triage, subject to validation and version control. The new signal does not automatically create suspicion. It raises review urgency when multiple independent indicators support a connected pattern.
Finally, QA expands beyond completed alerts. It samples aged, deferred and automatically deprioritised items to look for missed risk. The results feed scenario and priority governance.
Lessons for delivery teams
For operations, the case shows why productivity measures can alter reviewer behaviour even when procedures are correct on paper. For compliance, it shows why oversight should examine who is waiting in the queue, not only who was reviewed. For architects, it shows the need to preserve assignment, priority and override history. For data teams, it shows that linkage features require lineage and validation. For testers, it shows that queue-order tests should include selection behaviour, ageing and reassignment, not just whether the calculated score is numerically correct.
The larger lesson is that triage is a socio-technical control. Scoring logic, workflow, incentives, data and human judgement operate together. Improving one component while leaving the others unchanged can simply move the weakness somewhere else.
Knowledge check and glossary
Use these questions to test whether the triage concepts are understood as control principles rather than memorised terminology.
Is a high-priority alert the same as a suspicious transaction? No. Priority determines review urgency. Suspicion is a legal or policy judgement made after relevant facts are assessed. A high-priority alert can close with a benign explanation, and a lower-priority alert can become important after new evidence is linked.
Why should an alert and a case be separate concepts? An alert records a detection event. A case is an investigation container that may combine several alerts, customers, accounts and external information. Keeping them separate preserves traceability and avoids treating each system signal as an independent investigation conclusion.
Does FATF prescribe one global alert-priority formula? No. FATF requires a risk-based approach and proportionate controls, but institutions design their own monitoring and prioritisation methods within applicable law, risk assessment and supervisory expectations.
Should transaction value always be a major priority factor? Not necessarily. Value can be relevant, but amount-only ranking can miss structuring, mule activity, terrorist-financing patterns and network behaviour. The weight given to value should be justified by the institution’s risk and outcome evidence.
When does a regulatory reporting clock start? It depends on the applicable jurisdiction and obligation. A system alert is not universally the trigger. Requirements must define the specific legal event and rule for each jurisdiction rather than using one global assumption.
Can an internal SLA be treated as a statutory deadline? No. An internal target can be stricter than the law and is useful for operations, but the system and management reporting should clearly distinguish internal targets from legal deadlines.
Why is historic SAR or STR filing an imperfect machine-learning label? Filing records a suspicion and reporting decision, not a court finding that crime occurred. Similarly, closure of an alert does not prove the activity was benign. Models need broader validation and missed-risk testing.
What is cherry-picking? It is the selection of easier work while more complex or urgent assigned work waits. It can be encouraged unintentionally by pooled queues and performance measures focused on closure volume.
How should cherry-picking be addressed? First examine the operating model: assignment rules, incentives, workload measures and supervisor visibility. Persistent individual avoidance can then be handled under performance or conduct processes, but flawed incentives should not be ignored.
Why should deferred alerts remain visible? Deferral is a decision to wait, not removal of risk. Governance needs to see the population, age, priority, review-by date and reason so postponement does not become an invisible backlog.
What makes reprioritisation auditable? The system retains the old and new priority, reason, relevant input changes, rule or model version, actor and timestamp. Overwriting the original score destroys the decision history.
Why is capacity part of control effectiveness? Correctly ordered work still fails if reviewers cannot examine it with appropriate skill and timeliness. Capacity planning should reflect volume, complexity, quality and specialist needs rather than forcing detection output to fit available headcount.
Can outsourcing remove the bank’s accountability? No. A service provider may perform review tasks, but the regulated institution remains responsible for its AML/CFT control obligations and should govern the provider accordingly.
What should QA sample? Completed alerts are important, but aged, deferred, automatically deprioritised and overridden alerts can reveal risks that completed-work sampling misses.
Does a higher case-conversion rate always prove a better priority tier? No. Conversion can be a useful indicator, but some tiers may exist because of time sensitivity, customer protection or other risk considerations. Effectiveness should use several measures and expert review.
Why separate detection and workflow logic? Detection decides that activity is unusual enough to generate an alert. Workflow decides who reviews it and when. Separating the concerns makes change ownership clearer and reduces the risk that workload pressure alters detection logic without proper governance.
Glossary
Alert: A detection event or referral that requires review; it is not itself a finding of suspicion.
Case: A structured investigation record combining relevant alerts, parties, transactions, evidence, decisions and outcomes.
Priority: The operational ordering of work according to approved urgency and risk factors.
Priority reason: Human-readable explanation of the factors that caused the assigned priority.
Reprioritisation: Controlled change to review urgency after new information or circumstances arise.
Internal SLA: An institution-defined target for completing a process step. It should not be confused with a statutory deadline.
Statutory reporting clock: A legally defined timeframe that starts from the trigger specified by the applicable law or regulation.
Deferral: Approved postponement of review with a documented rationale, review-by date and continuing visibility.
Queue health: The condition of review inventory measured through factors such as age, priority, deadlines, assignment, quality and capacity rather than raw volume alone.
Skill-based routing: Assignment of work according to required expertise as well as urgency and workload.
Manual override: A controlled human change to a model or rule-generated priority with documented rationale.
False-negative testing: Review designed to assess whether alerts or activity that were not escalated may contain material risk the control missed.
Shadow testing: Running a candidate model or rule against live data without allowing it to change live ordering or customer/regulatory outcomes until approved.
Label limitation: Recognition that operational outcomes such as SAR/STR filing or alert closure are imperfect proxies for actual criminal or benign activity.
References and further reading
The sources below are public, authoritative references used for the chapter. Jurisdiction-specific material is labelled accordingly; it should not be read as a universal rule for every bank.
Global standards and supervisory principles
- Financial Action Task Force (FATF), The FATF Recommendations: https://www.fatf-gafi.org/en/publications/Fatfrecommendations/Fatf-recommendations.html
- Financial Action Task Force (FATF), 2025 update to the Standards on proportionality and financial inclusion: https://www.fatf-gafi.org/en/publications/Fatfrecommendations/update-standards-promote-financial-conclusion-feb-2025.html
- Basel Committee on Banking Supervision, Basel Consolidated Guidelines, AFS10: Anti-money laundering and counter-terrorist financing, published 1 January 2026: https://www.bis.org/committees/bcbs/basel-consolidated-guidelines/module/afs/10
- Wolfsberg Group, Statement on the Risk-Based Approach: https://wolfsberg-group.org/resources/rba/203
- Wolfsberg Group, Statement on Effective Monitoring for Suspicious Activity, Part I: https://wolfsberg-group.org/resources/general/168
- Wolfsberg Group, Statement on Effective Monitoring for Suspicious Activity, Part II: https://wolfsberg-group.org/resources/195/202
United States: bank SAR review and examination guidance
- Federal Financial Institutions Examination Council (FFIEC), BSA/AML Manual: Suspicious Activity Reporting: https://bsaaml.ffiec.gov/manual/AssessingComplianceWithBSARegulatoryRequirements/04
- Federal Financial Institutions Examination Council (FFIEC), BSA/AML Manual: Suspicious Activity Reporting Examination Procedures: https://bsaaml.ffiec.gov/manual/AssessingComplianceWithBSARegulatoryRequirements/04_ep
These US sources are used for the chapter’s examples on alert review, staffing and SAR timing. The US filing rules should not be applied as global deadlines.
Australia: current suspicious-matter and ongoing-monitoring guidance
- AUSTRAC, Suspicious matter reports, current guidance updated 9 September 2026: https://www.austrac.gov.au/industry-and-business/obligations-and-guidance/your-amlctf-program/reporting-us/suspicious-matter-reports
- AUSTRAC, How to monitor your customers: https://www.austrac.gov.au/industry-and-business/obligations-and-guidance/your-amlctf-program/customer-due-diligence/ongoing-customer-due-diligence/how-monitor-your-customers
Practical supervisory example on monitoring data completeness
- UK Financial Conduct Authority, FCA fines Metro Bank £16m for financial crime failings, published 12 November 2024 and updated 5 December 2025: https://www.fca.org.uk/news/press-releases/fca-fines-metro-bank-16m-financial-crime-failings
This FCA enforcement release is included as a practical reminder that monitoring effectiveness depends on complete and reliable data. It is not used as evidence for a universal triage methodology.