Where traditional rules failed in real banking. A practical lesson in why banks turned to ai for banking and payments practitioners.
Study purpose
This chapter is written as a serious banking study guide, not a technology brochure. The aim is to explain where traditional rules failed in real banking in a way that a business analyst, architect, developer, tester, risk specialist, operations lead, compliance reviewer or student can use inside a real bank. The discussion stays close to banking decisions, data, controls, customer impact, model governance and audit evidence.
The chapter also explains how this topic connects to AI and machine learning without pretending that every banking problem should be solved by AI. Some questions need deterministic rules. Some need scorecards. Some need statistical models. Some need human judgement. The practical skill is knowing which method belongs where, and how the evidence travels from source data to final bank action.
How to study this chapter
This chapter explains why deterministic rules remain essential in banking but became insufficient for many modern risk, payments, fraud, compliance and operations decisions. A rule is powerful when the bank knows the exact condition and the exact action. It is weak when behaviour is changing, signals interact, thresholds create cliffs, and operational teams keep adding exceptions to repair yesterday's miss.
In practical banking terms, this means the bank must identify the decision point, the data available at that moment, the control owner, the allowed action, the expected customer or regulatory impact and the evidence that will be stored afterwards. A model output becomes useful only when it changes a real workflow in a controlled way. If the output cannot be tied to an action, a user, a policy rule, a monitoring metric and an audit trail, it is not yet production-grade banking AI.
For delivery teams, the safest approach is to describe the use case as a chain: source event, data validation, feature or rule calculation, model or score output, policy orchestration, human review where needed, customer or operational action, reporting and feedback. This keeps the project grounded. It also prevents the common mistake of discussing AI as if it floats above the bank instead of sitting inside payment hubs, credit platforms, risk engines, case tools, ledgers, data warehouses and monitoring dashboards.
What traditional rules did well
Traditional rules gave banks control. They could reject a payment without mandatory fields, stop a product outside eligibility, escalate an overdue complaint, block a transaction above an agreed threshold, or require enhanced due diligence for a high-risk profile. These rules were understandable, testable and auditable.
Where they started failing
The failure appeared in the middle ground: not clearly safe, not clearly bad, but behaviourally unusual. A transaction can be below a threshold but suspicious in context. A customer can pass one affordability rule but still show stress across account behaviour. An alert can match a rule but have very low investigation value.
Threshold cliff effects
Rules create sharp boundaries. A transaction of 9,999 may pass while 10,001 is reviewed. A score just above a cut-off may approve while a nearly identical score is referred. Thresholds are sometimes necessary, but when used alone they create unnatural jumps in customer treatment and operational workload.
Rule explosion
When a bank adds a new rule for every new case, the rule estate grows quickly. One rule catches a pattern, another suppresses noise, another applies to a product, another applies to a country, another is temporary but never removed. Over time, nobody can confidently explain the whole decision logic.
Payments example
In payments, rules can check mandatory data, currency, scheme reachability, cut-off, sanctions screening status, duplicate reference and account validity. But a rule alone may not understand the combination of customer behaviour, beneficiary novelty, device change, transaction velocity, scam indicators and previous failed attempts.
Fraud example
Fraud is a classic failure area for static rules. Criminals adapt. If the bank blocks only high-value transactions to new beneficiaries, fraud may move below the limit, split across transactions or use social engineering. A model can rank combined risk, but policy rules still decide which actions are allowed.
AML example
Transaction monitoring rules can create high alert volumes. A rule may be required for a typology, but too many broad thresholds produce low-quality queues. Learning systems can help prioritise alerts, but the bank must not suppress legal obligations without defensible evidence and compliance ownership.
Credit example
Credit rules can enforce eligibility and affordability floors. But customer risk is rarely one variable. Repayment behaviour, income volatility, bureau history, account conduct, product usage and macro sensitivity interact. Scorecards and ML can combine those signals better than one hard threshold.
Operations example
Operations teams often rely on rules for service level, value, age and customer type. But real work queues also need pattern recognition: duplicate breaks, recurring root causes, high-risk customer impact, likely repair path and downstream settlement or reporting risk.
Why AI was not a simple replacement
The answer is not to remove rules. Mandatory policy, legal restrictions, field validation and scheme rules must remain deterministic. AI helps where ranking, pattern recognition and prediction are needed. Banks need orchestration between rules, models and human review.
What business analysts must capture
A BA should document each rule's purpose, owner, trigger, input data, exceptions, override path, evidence, downstream impact and retirement condition. Rules fail quietly when ownership disappears. They fail loudly when customer impact, fraud loss or regulatory exposure appears.
The main lesson
Traditional rules failed when banks expected them to behave like learning systems. Rules are excellent for known obligations. They are weak for changing behaviour, interacting signals and prioritisation. The future is not rules versus AI; it is governed decision orchestration.
Chapter-level control checklist
| Control question | Why it matters | What good looks like |
|---|
| What decision is being supported? | Prevents vague analytics from entering production | One named decision point, one accountable owner and one defined action |
| What data is known at that moment? | Prevents look-ahead bias and weak evidence | Point-in-time source data with lineage and quality checks |
| What must remain deterministic? | Protects legal, policy and scheme obligations | Mandatory rules remain rules and are not silently overruled by a model |
| What does the model output mean? | Avoids blind trust in a number | Output type, reason, limitation and confidence are clear to users |
| Who can override? | Keeps human judgement accountable | Override reason, authority, evidence and outcome are captured |
| How is performance monitored? | Detects drift, bias and operational harm | Dashboards track outcomes, exceptions, false positives, false negatives and incidents |
| What evidence is retained? | Supports audit, validation and regulatory review | Input, score, version, rule hits, decision, user action and final outcome are stored |
Business analyst study prompts
- Draw the process before and after the model is introduced. Mark exactly where the bank decision changes.
- List the deterministic controls that must remain outside model discretion.
- Identify which source systems provide the data and whether the data is available before the decision.
- Write the reason a user would trust, challenge or override the output.
- Define the customer impact if the model is wrong in both directions.
- Describe what operations, risk, compliance, finance and technology each need to see.
- Explain what would happen if the model service is unavailable during a peak period.
- Define the monitoring report that proves the use case is still safe after go-live.
Main lesson
Rules still matter, but they fail when behaviour changes faster than policy thresholds and when many weak signals matter together. The best banks will not adopt AI by replacing every existing rule, scorecard or control. They will adopt AI by understanding where learning systems improve judgement, where deterministic rules remain stronger, where human review protects customers, and where evidence must be retained for audit and regulatory challenge.
For Malla Banking Academy, the takeaway is simple: this topic belongs to banking first and technology second. AI and ML become valuable only when they are connected to capital, provisions, payments, fraud, sanctions, liquidity, reconciliation, reporting, customer treatment and operational resilience. That is the difference between a generic AI explanation and a bank-grade learning chapter.
A rule can be correct and still miss the pattern
A deterministic rule has a precise job: given a known condition, it returns a specified action. A bank should keep that mechanism for eligibility, required fields, legal restrictions and hard policy boundaries. The weakness appears when the bank asks a fixed condition to identify a changing pattern. A threshold may work for a known type of misuse and then become less useful when behaviour moves below it, spreads across accounts or changes channel. The failure is not proof that machine learning is always preferable. It is evidence that the bank must separate a hard requirement from a prediction, and evaluate each against the outcome it is meant to control.
Consider a fictional bank that blocks a card after five failed purchase attempts in ten minutes. The rule is reproducible, and it may stop a crude attack. It misses three attempts on each of several cards from the same device, or a low-volume sequence that changes merchant and time of day. Lowering the threshold to three may create more customer challenges without detecting the cross-card pattern. A model could combine device, account, merchant, location and timing signals, but it also needs labelled outcomes, a decision deadline and a fallback when a feature source fails. The threshold can remain a backstop while the model ranks cases for an approved policy action.
A different rule rejects an application when a required field is absent. That rule should not be weakened because a model predicts a good repayment outcome. It protects a known process condition. The useful design question is which outcome needs a fixed rule, which needs a statistical estimate and which requires a person to review conflicting or incomplete evidence.
Card fraud: a velocity rule meets a coordinated pattern
Assume a fictional issuer sees an authorization at 10:15. Its old rule counts purchases above a specified amount on the same card in the preceding hour. The rule catches a sudden burst on one card, but a criminal uses several compromised cards and keeps each amount below the threshold. A simple change to the threshold may increase false positives for ordinary customers making several small purchases. The suspicious relationship lies across device, beneficiary or merchant context and time, rather than one amount on one card.
The analyst should map the exact event that the bank can act on. For card authorization, the response window may be short and the final action is not the model's score. The score is an input to the issuer's decision policy. The bank may approve, challenge, refer or decline under approved rules. A challenge has a customer cost; a referral has operational cost and a time limit. The model cannot turn an unavailable device feed into a valid low-risk signal. It must return a failure or missing-feature status, and the policy needs a documented response.
Build a dated feature contract. An attempt count must specify the card or customer identifier, eligible authorization states, window boundary, time zone and deduplication key. A reversal or retry should not be counted as a new purchase without a reason. A merchant category may change in reference data; a historic decision needs the version available then. Training must use what the service could have observed at the authorization time. If it uses a later chargeback or investigation note as an input, the back-test will appear accurate for a reason production can never reproduce.
The pilot needs more than an alert hit rate. Compare eligible authorizations, confirmed fraud losses after labels mature, legitimate declines, step-up completion, investigator load and customer complaints. A large reduction in fraud loss may reflect a separate rule change or a different transaction mix. Test by channel, product and relevant customer segment. An analyst should be able to reconstruct one declined transaction from event ID through feature values, model version, policy version and final action. A shadow model may generate scores for comparison, but it must be clearly marked as non-decisioning.
A rule remains useful for prohibited transactions, mandatory authentication conditions or a defined contingency path. A model may recognize combinations that a single threshold misses, yet it also introduces performance drift and validation work. The design succeeds when the bank knows what each component can decide, how the outcome is measured and who can reverse an incorrect customer action.
AML monitoring: a threshold is not an investigation
A bank may generate an alert when transaction volume exceeds a defined scenario threshold. This gives investigators a consistent starting point. It can also flood a queue with activity that has an ordinary explanation, while missing a customer who spreads transactions across products or counterparties. Increasing the threshold reduces volume mechanically. It does not demonstrate that the monitoring system has retained coverage of relevant risk. A model that ranks alerts has the same limitation if the bank judges it solely by the number it suppresses.
The FFIEC suspicious activity reporting overview describes alert management as the evaluation of identified unusual activity. It recognizes multiple identification methods and expects processes to assess them. That is a U.S. examination reference, not a global rule for every institution. The bank must apply its own jurisdiction's obligations and risk assessment. The model should support the investigator's work rather than assert that a low score means a transaction is lawful.
Suppose a fictional corporate customer regularly pays one overseas supplier. A new pattern of several smaller transfers to related counterparties may not cross the old single-payment amount rule. The analyst needs to define the relationship evidence: customer ownership, counterparty links, geography, frequency and baseline period. A relationship graph could help group alerts, but a broken customer identifier might join unrelated companies. The investigation must preserve source records and uncertainty. A model's ranking is not a suspicious activity report decision.
The review workflow distinguishes initial trigger, model priority, investigator notes, escalation, closure and later correction. A case closed without escalation is not proof that the behaviour was innocent. An unreviewed low-priority case has even less reliable outcome evidence. A validation sample should include high- and low-ranked cases, different scenarios and time periods, with documented selection. The team checks whether a changed risk profile or new typology has reduced coverage. Investigator capacity matters, but it should not be treated as a license to silently discard alerts.
When an AI ranking layer is unavailable, the existing scenario monitoring may continue under a contingency plan. The case tool should retain the original trigger and queue status. After recovery, it must reconcile alerts created during the outage and prevent duplicate closures. This is a practical division of labour: deterministic scenarios surface defined conditions, analytical methods can add context, and authorised staff decide the case disposition under applicable policy.
Lending: the limit of a one-line approval rule
A lender might once have referred every applicant with income below a fixed level. Such a rule is easy to explain but ignores obligations, variability of income, existing debt and product terms. Replacing it with a score alone is not a cure. A customer can have a low predicted default risk in a historical model and still fail a current affordability assessment. The bank must keep eligibility, identity, fraud, affordability and credit-risk questions distinct.
Take two fictional applicants with the same verified monthly income. One has a modest fixed repayment and stable cash flow; the other has substantial existing commitments and variable income. A single income cut-off treats them alike. A more detailed rule set can account for obligations and evidence, but a collection of thresholds may create cliffs at each band boundary. A credit model can estimate risk using multiple dated factors, yet its score depends on the population and outcome definition on which it was built. It cannot substitute for a policy rule or customer-specific evidence that the bank must consider.
The business analyst defines the decision date, product, target outcome, observation horizon and population. A default label observed months later can assess the model, but cannot be an input to the original decision. An account balance updated after the application is not the balance that was available at the time. If historic approved borrowers are the only people with observed repayment outcomes, validation must acknowledge the missing outcomes for rejected applicants. A model that looks strong on that selected sample may behave differently when the bank expands approvals.
For a covered U.S. adverse credit action, Regulation B section 1002.9 requires specific principal reasons. The bank's explanation must reflect factors actually used in the decision, including policy and human actions. Other jurisdictions have different rules; this is a U.S. example. A model explanation that highlights income stability would be misleading if the final refusal came from an affordability rule about existing commitments. The decision log needs to separate score, eligibility result, affordability result, human override and final notice reasons.
Test an applicant at a threshold boundary, one with stale bureau data, one with a missing income source and one referred for manual review. An authorised underwriter may correct a verified data error within approved authority. That correction should not erase the original score and policy result. Track approval, arrears and complaints by dated cohort, product and channel after outcomes mature. A model may add useful discrimination, but the bank's decision remains a controlled combination of evidence, policy and accountability.
Payment processing: the difference between validity and prediction
A payment rule can reject an instruction with a missing mandatory field or an invalid account format. Those are deterministic checks with a defined error route. A predictive model cannot make an invalid instruction valid. It may help identify a likely repair, predict an operational delay or prioritize investigation. Those uses require different authority from message validation, and the bank should not merge them into a vague "AI approved" status.
Imagine a fictional corporate payment with a beneficiary name that differs from the bank's reference record. A model suggests a likely intended value from historical instructions. The suggestion is not evidence that the new value is authorized. The bank must check current mandates, source quality, privacy and product rules before an operator or approved rule can change an instruction. Where repair is not permitted, the item is returned or referred under the relevant workflow. Changing beneficiary account details without a controlled evidential basis can misdirect funds.
A legacy routing rule may send payments down the cheapest available path. It could fail when it ignores cut-off, liquidity, scheme eligibility, counterparty reachability or settlement risk. A model may forecast a delay or likely rejection, but the routing engine must still enforce hard constraints. The decision record should show the original instruction, validation result, eligible routes, predictive inputs, policy version, selected route and eventual status. Clearing acceptance, settlement, booking and customer status are different events. A predicted success is not settled funds.
Test a payment below and above a cut-off, one with a stale reference value, one whose route becomes unavailable, and a duplicate retry. The fallback path should preserve idempotency and avoid sending two instructions. Reconcile accepted, rejected, returned and outstanding items against the payment hub and ledger, not only the model dashboard. If a model outage changes routing, the bank should know which payments used fallback and whether later settlement messages altered their final state.
This example is implementation-neutral. Different payment systems and schemes impose different message and timing rules. The teaching point is the boundary: deterministic checks define whether an instruction may proceed, predictive tools may improve prioritization or route choice within that boundary, and operations must be able to prove the eventual outcome.
Operations: rules can conceal a broken process
An operations team may use a rule to close a reconciliation break whenever the amount difference is below a small tolerance. The rule may be appropriate for a known rounding issue. It becomes dangerous when the same tolerance is used for an unmatched payment, an incorrect beneficiary or a settlement timing difference. Similar amounts do not establish the same cause. A model that groups breaks can help an investigator see patterns, but it should not convert every low-value item into an automatically resolved case.
A fictional bank receives a settlement confirmation without a matching internal payment record. The first questions are identifiers, value date, currency, amount, account and message status. A rule can match exact identifiers and amounts under a defined policy. If several records are plausible matches, a model might rank them for review, with its supporting features and uncertainty. The original break must remain open until an approved action resolves it. An operator may request source correction, wait for a late posting or raise an investigation. Each path has different accounting and customer implications.
A daily dashboard that reports fewer breaks after automation may be misleading if the process simply changes the closure code. Measure incoming items, auto-matched items, manual cases, reversals, ageing, reopened cases and unresolved value. Reconcile the counts with source systems, and sample automated closures for correctness. A rule that improves queue time but increases later reopens has shifted work rather than removed it.
The same caution applies to complaint triage. A keyword rule may direct "fraud" to a fraud queue, while a customer's message describes a card transaction and a separate account-access problem. A classifier can propose categories, but the case system needs ownership, service deadlines, escalation and an appeal or correction route. A low-confidence classification should not silently lose the complaint. A human may still need to decide a customer remedy.
Operational learning should begin with the specific failed item and its evidence. The owner asks whether the defect came from bad source data, an ambiguous rule, a stale model, a changed policy or a hand-off between teams. That diagnosis determines the remedy. Replacing one threshold with a prediction does not fix a broken ledger, missing event or undefined case owner.
Choosing a hybrid control and proving its value
A disciplined design separates four layers. Data validation checks whether inputs are present, current and allowed. Deterministic policy enforces known requirements. A model estimates an uncertain outcome or ranks work. A person reviews cases where judgement, legal authority or contested evidence requires it. Some journeys omit the model, and some can apply a model without a human reviewing every case. The allocation depends on risk, product, law and operational capacity.
Write an action contract before training. For each model output, name the caller, permitted population, score meaning, threshold owner, final action, fallback and retained evidence. A number without a defined action is analytics, not a bank decision. A confidence or explanation field does not automatically make the decision fair or accurate. The owner needs validation on the intended population, performance monitoring after deployment and a way to restrict the model when its assumptions fail.
The revised U.S. interagency Federal Reserve SR 26-2 guidance describes risk-based model governance for covered banking organisations. Its scope and applicability must be checked for the institution and model use. The NIST AI Risk Management Framework is voluntary and offers a broader structure for mapping, measuring and managing AI risk. Neither source says that every fixed rule should be replaced with machine learning.
A comparative pilot should use the same dated population for the old rule and proposed design. Record how many cases each flags, how many true outcomes are observed after a suitable period, what manual workload follows, which customers are affected and what data was unavailable. A challenger run in shadow mode can be compared without changing customer actions. Where outcome labels are delayed or selectively observed, state that limitation instead of presenting a single "accuracy" percentage. A pilot should include failure cases and a rollback plan, not only a favourable back-test.
For acceptance, test normal flow, edge thresholds, stale data, feature outage, conflicting model and policy results, duplicate requests, manual override and appeal or correction. The log should preserve source timestamps, model and rule versions, action authority and final disposition. A reviewer should be able to reconstruct a single customer case. Measure the actual workflow after launch, not just the model score distribution. If a new method reduces alert volume but raises missed-risk indicators or complaints, the owner must investigate and may restrict use.
A worked decision comparison
Use an illustrative card authorization to make the distinctions concrete. The payment is 42 currency units at a new online merchant. The cardholder has made two ordinary purchases that morning. A fixed amount rule does not trigger. A model sees the new merchant, an unusual device and a recent failed login. The bank's policy may issue a challenge. The customer completes it and the transaction proceeds. The model did not "approve" the payment; the issuer's control chain did.
Now change one fact at a time. If the device signal is unavailable, the service returns a missing-feature status and the approved fallback applies. If the challenge fails, the bank records the failure and final payment status. If a later fraud claim arrives, it is an outcome for monitoring, not a feature that existed at the original authorization. If the customer disputes the challenge, operations can retrieve the dated evidence and correct a process defect. None of these states is captured by a single hit-or-miss rule metric.
An analyst can document this as a small decision table with columns for observed data, hard control, model result, policy action, customer response and final status. Populate it for a normal case, an outage, a stale feature and a duplicate retry. Then ask which cases the old rule missed, which legitimate customers the new design inconvenienced and whether the net effect warrants a change. The answer is a measured business decision under bank ownership, not a claim that AI is intrinsically better.
Source notes for further study
For U.S. organisations within scope, see the Federal Reserve's SR 26-2 and the OCC's 2026 revised model-risk guidance.
Additional banking practice note 1
A real bank should never treat this chapter as only a model-building exercise. The model sits inside policy, architecture, workflow, risk appetite, customer communication, operational support and evidence retention. That full chain is what makes the solution bank-grade rather than experimental. In the context of where traditional rules failed in real banking, this means the team should document the exact portfolio, channel, product, process and control boundary before making design decisions. A retail credit example, a corporate treasury example, a sanctions alert example and an instant payment example may all use data-driven scoring, but the risk owner, evidence requirement and customer impact are different.
A useful working question is: if this output is challenged later, who can explain why the bank trusted it? The answer should include the business owner, data owner, model owner, validation evidence, monitoring result, user action and stored audit trail. If that answer is weak, the bank may have analytics, but it does not yet have a controlled banking capability.
Additional banking practice note 2
The practical difficulty is usually not the algorithm. It is agreeing the definition, finding the trusted source, proving the data timing, making the output usable for staff, preventing misuse, monitoring outcomes and explaining the result months later to someone who was not part of the delivery team. In the context of where traditional rules failed in real banking, this means the team should document the exact portfolio, channel, product, process and control boundary before making design decisions. A retail credit example, a corporate treasury example, a sanctions alert example and an instant payment example may all use data-driven scoring, but the risk owner, evidence requirement and customer impact are different.
Additional banking practice note 3
This is also why business analysts matter so much in banking AI work. They translate between risk language, product language, operations language, data language and technology language. Without that translation, a technically good model can still fail because the bank cannot use it safely. In the context of where traditional rules failed in real banking, this means the team should document the exact portfolio, channel, product, process and control boundary before making design decisions. A retail credit example, a corporate treasury example, a sanctions alert example and an instant payment example may all use data-driven scoring, but the risk owner, evidence requirement and customer impact are different.
Additional banking practice note 4
The strongest implementation pattern is staged adoption. First understand the current process. Then run the model silently. Then compare with existing decisions. Then expose it as decision support. Then automate only low-risk paths when monitoring evidence proves that the use case is controlled. In the context of where traditional rules failed in real banking, this means the team should document the exact portfolio, channel, product, process and control boundary before making design decisions. A retail credit example, a corporate treasury example, a sanctions alert example and an instant payment example may all use data-driven scoring, but the risk owner, evidence requirement and customer impact are different.
Additional banking practice note 5
Customer impact must remain visible. A false positive can delay a genuine payment, decline a good customer, create a complaint or overload operations. A false negative can allow fraud, credit loss, financial crime exposure or regulatory breach. Both sides of error need business cost and control ownership. In the context of where traditional rules failed in real banking, this means the team should document the exact portfolio, channel, product, process and control boundary before making design decisions. A retail credit example, a corporate treasury example, a sanctions alert example and an instant payment example may all use data-driven scoring, but the risk owner, evidence requirement and customer impact are different.
Additional banking practice note 6
The chapter should therefore be read as part of the larger AI and ML banking journey. Data foundation, feature design, model governance, production monitoring, fallback, human oversight and audit evidence are not separate topics. They are the operating model that allows AI to be adopted safely. In the context of where traditional rules failed in real banking, this means the team should document the exact portfolio, channel, product, process and control boundary before making design decisions. A retail credit example, a corporate treasury example, a sanctions alert example and an instant payment example may all use data-driven scoring, but the risk owner, evidence requirement and customer impact are different.
This application uses JavaScript for the full interactive experience. This text summary is served for accessibility and search indexing.