How customer behaviour changed the risk problem

How customer behaviour changed the risk problem. A practical lesson in why banks turned to ai for banking and payments practitioners.

Risk moved from a profile to a sequence of events

A customer profile remains useful. It records identity, products, stated occupation or business activity, known counterparties, credit obligations and information needed for customer due diligence. A profile is a dated description, however. It cannot by itself explain a new device, an unusual transfer, a missed repayment, or a sudden change in account use. A behavioural view follows events over time and asks whether the change is relevant to a particular decision. It may improve fraud detection, credit monitoring or customer support when the bank knows what the signal means and how recently it was observed.

The change is not a claim that every digital event is risk. A new phone can be a normal upgrade. A large payment may be a property purchase. A delayed salary credit may reflect a public holiday or an employer's payroll cycle. A model can rank such events for attention, but the bank must check context before interpreting a deviation as fraud or distress. The customer's ability to explain or correct information is part of a responsible workflow. An unusual pattern can be an investigation lead, not a finding about character or intent.

For this lesson, a behavioural signal is a defined observation, such as a beneficiary added, a credit limit used, or an account accessed from a newly registered device. A risk decision is a controlled action, such as requesting additional authentication, referring a transaction, reviewing an account or changing an offer under approved policy. A label is a later outcome used to evaluate the signal, such as confirmed fraud or a defined credit delinquency event. Those three objects must stay separate. A bank should not relabel every referred payment as fraud, or every declined applicant as likely to default.

A learner can assess a use case by naming the population, observation window, target event, decision deadline, permitted action and person accountable for the outcome. "Detect changing behaviour" is too broad to test. "At payment initiation, use prior verified customer activity to decide whether a new-payee transfer needs step-up authentication before release" is testable. The bank can then identify required data, mandatory checks, fallback, error costs and evidence to retain. This is a descriptive training example, not a rule for any particular payment scheme.

Different signals answer different questions

Digital channels expose account access, session and device signals. Payment systems expose instruction amount, beneficiary, channel, timing and outcome. Credit systems expose drawings, utilisation, repayment status, modifications and collections contact. Service systems expose disputes, complaint themes and customer-reported problems. Customer due diligence records establish a separate view of expected activity and changing customer facts. A single event can appear in several systems, with different timestamps and definitions. The bank must choose signals for the decision at hand rather than pour every field into one score.

Consider an account holder who logs in from a new device. The access team may ask whether the device was enrolled through a verified process. The fraud team may combine that event with a new payee and an attempted transfer. A credit team has little reason to interpret a device change as repayment capacity. A compliance team may consider whether actual account activity differs from the customer's risk profile under its risk-based programme, but a device change alone does not establish suspicious activity. A shared event feed can serve several teams while each uses a different purpose, policy and retention rule.

Timing gives a signal meaning. "Five transfers today" depends on which time zone defines today, which attempts were accepted, whether reversed items count and whether retries share an instruction identifier. "New beneficiary" depends on account ownership, approval time and whether a matching payee existed in another channel. "Salary stopped" depends on the agreed salary-classification logic and normal pay schedule. A dashboard that treats missing events as zero activity may miss an ingestion failure. Every useful feature needs source, calculation, window, freshness, null treatment and a point-in-time replay method.

The bank also has to respect purpose boundaries. A device identifier collected to secure login should not automatically become an unrestricted credit feature. Applicable privacy, consumer and data-protection obligations vary by jurisdiction and product. The owner should record why a source is appropriate, who can access it, how long it is retained and whether an alternative could achieve the decision with less data. More data can worsen reliability if it is stale, hard to explain or unevenly available across customers.

A complete customer journey

Imagine a fictional customer, Leela, who usually receives salary near month-end, pays household bills and occasionally transfers money to relatives. On a Sunday she changes her phone, adds a new payee and requests a payment for 18,000 illustrative units. She also contacted support after receiving a message that claimed to come from the bank. The access, support and payment events are individually ambiguous. Together they justify a proportionate review of whether she is being impersonated or manipulated. They do not justify a conclusion that she is a fraudster.

The channel authenticates the session under its approved controls and records the device-enrolment evidence. The payee service records the payee creation time and any verification outcome. At initiation, the payment service applies eligibility and mandatory controls. A fraud service may request features that were available before that decision and return a versioned score or risk band. A separate policy chooses the allowed action, for example step-up confirmation, a short referral where supported, or continuation. The bank's rail, law and policy determine what can actually be done before the payment is executed. If the risk service is unavailable, the bank follows a preapproved fallback rather than treating an absent score as low risk.

Suppose Leela completes a secure callback and confirms she intended the transfer. The bank records the contact method, agent, time, explanation and final action without storing unnecessary sensitive conversation text in a model feature. If she reports coercion, staff follow the bank's scam-response process and communicate the payment state accurately. The later investigation may find an authorised payment, a scam or an account-takeover attempt. Until then, the label is unresolved. Training data should not convert the initial referral into confirmed fraud merely because a case was opened.

At each hand-off the bank keeps a stable event or decision identifier. The payment outcome, fraud disposition and customer communication must reconcile. A customer may see a pending screen while the payment network returns a final status; operations needs to resolve any mismatch. An approved review action also needs a deadline. A model result that arrives after a final payment decision can support investigation, but it cannot retroactively prevent the payment. The preceding lesson on real-time banking explores that deadline in more detail; this lesson concentrates on interpreting behaviour without making unsupported conclusions about the customer.

The same journey could create a false alarm. Leela might have verified her new device, added a contractor as payee and called support about an unrelated phishing email. A model should be tested against such combinations. The evaluation records the number of genuine customers referred, delay duration, customer complaints and confirmed harmful outcomes. A bank that only reports "alerts generated" cannot tell whether the control improved protection or shifted cost to customers and staff.

Change from static forms to ongoing observation

At onboarding, the bank gathers identity and product information, evaluates eligibility and establishes expected account use under applicable policy. After onboarding, the account can change. A small business may expand into another market. A student may start employment. A customer may experience financial difficulty. A static application form should remain in the historical record; it cannot be treated as a live description of every later event. Ongoing observation can prompt a refresh, verification or support, but it must not silently rewrite the original facts.

A behavioural credit signal has a different horizon from a payment-fraud signal. A missed instalment is a contractual fact, subject to posting corrections and grace or due-date definitions. A decline in salary credits is a possible early-warning signal, but it may reflect a new employer, a different receiving bank or irregular work. An increase in utilisation may indicate stress or normal seasonal activity. A credit team can use a defined cohort and observation period to decide whether a model meaningfully predicts a later outcome. It should not use a fraud model's alert as a direct substitute for an affordability assessment.

For U.S. covered credit decisions, Regulation B's adverse-action provisions require specific principal reasons when adverse action is taken. A creditor cannot use an opaque behavioural score as an excuse for an uninformative reason. This is a U.S. example; other jurisdictions have their own consumer and privacy rules. The decision record must distinguish model evidence from eligibility, affordability and human judgment so the stated reasons match what actually determined the outcome. A contact-centre note, for instance, may alert staff to a customer needing help while being unsuitable as a credit-decline reason.

The FFIEC BSA/AML examination manual describes U.S. customer due diligence and risk-based ongoing monitoring in a separate regulatory context. A pattern that differs from an expected profile may require investigation under the bank's programme. It is not automatically a reportable event or a fraud conviction. Compliance analysts need transaction details, customer context, prior activity and an escalation record. The bank should not merge that assessment with a lending score merely because both use transaction history.

Constructing a point-in-time behavioural feature

An analyst begins with a business question, then defines the feature in terms of records available before the decision. For an illustrative payment control, "number of distinct new payees with an accepted transfer in the previous 30 days" needs an account scope, payee identity rule, accepted-status definition, start and end boundaries, event-time source and treatment of returns. The feature service should return a value, the window's end time, a freshness status and a version. Historical replay must apply the same rule to dated events. Otherwise an excellent back-test may simply have read future outcomes.

The source sequence is often untidy. One channel accepts an instruction, another service rejects it, and a reporting feed publishes a correction later. The analyst should distinguish customer attempts, accepted bank instructions, network submissions and settled outcomes. A retry may have a new technical message ID but the same business payment. Deduplication needs an explicit key and reconciliation to the bank's authoritative outcome. A source-record correction should create a traceable restatement or new processing run, not silently alter the evidence behind a prior decision.

A joined customer view can fail when identifiers differ. A person may have multiple accounts, a joint account, a card token and a business role. The bank needs a governed identity crosswalk with effective dates and confidence or exception handling. A fuzzy match that combines two people's transactions can create a convincing but false behavioural change. It can also leak another customer's activity into a decision explanation. Access controls and test data should prevent that. When the crosswalk is unresolved, the model should receive a missing or limited-view flag, not invent complete history.

Missingness has multiple causes. A customer may have no previous transactions, a feed may be delayed, consent or access may be absent, or a product may not expose the field. Those cases cannot all be encoded as zero. A thin-history customer may need an alternative verification path. A late feed should raise a data-quality incident. The bank should measure feature availability by channel, product and relevant customer segment. If a new mobile release changes event capture, a sudden shift in scores might reflect instrumentation rather than new risk.

An acceptance set should include a transaction exactly at a time-window boundary, two identical retries, a reversed transfer, a payee created in another channel, a daylight-saving transition and a correction posted after the decision. The expected outputs should be calculated independently from a small dated event table. The test should also verify that a replay retains the original inputs and identifies later corrections separately. A model can be mathematically correct while the feature contract is wrong.

Outcome labels are delayed and selected

Fraud can be reported by a customer, suspected by an analyst, confirmed after investigation or reversed after a dispute. Those are different statuses. A credit outcome may require several months to reach a defined delinquency or default threshold. A compliance case may remain open without a definitive "innocent" label. The team should choose a target that corresponds to the intended risk decision, specify a maturation period and retain unresolved outcomes rather than forcing them into a binary field. The outcome's source and as-of date belong in the training record.

Selection changes what the bank observes. Suppose investigators review only high-score payments. Their cases may acquire detailed labels, while most low-score transactions remain "not reported." Training on investigated cases alone can make the model reproduce the previous referral policy. A later low rate of complaints is not proof that every unreviewed payment was safe. The analyst can compare multiple outcome sources, record label coverage by score band and channel, and state the uncertainty in a back-test. A model's apparent precision can be driven by who was investigated.

A fictional validation cohort has 50,000 eligible payment attempts. After a sufficient observation window, 80 have confirmed scam or fraud outcomes under a documented definition. An incumbent policy refers 500 attempts and captures 32 of the 80 confirmed cases. Its observed referral precision is 32/500, or 6.4%, and its observed capture is 32/80, or 40%, if the labels are complete. A challenger refers 700 and captures 42 on the same cohort. It adds 200 referrals for ten additional observed captures, a 5% incremental yield. These figures do not prove the challenger should be deployed. Loss values, customer disruption, investigation capacity and incomplete labels may reverse the business conclusion.

The denominator must be explicit. If 5,000 attempts could not be scored because device data was missing, reporting precision only among the 45,000 scored attempts hides an availability problem. A fair comparison should state received, eligible, scored, referred, acted upon and outcome-matured populations for both policies. It should report case counts and value at risk, including the cost of falsely delayed genuine payments. The bank also needs a prospective or shadow period to see whether the historical relationship persists. A threshold that looks good on a prior month may fail during a new scam campaign.

Context, proxies and fair treatment

Behavioural signals can proxy for circumstance rather than risk. An account that receives irregular income may belong to a freelancer, seasonal worker or small business. A customer who changes device often may have a shared or repaired phone. Accessibility needs can change interaction timing and authentication behaviour. A model that treats those patterns as suspicious without context may generate more friction for some groups. The bank should examine outcome and error rates across relevant segments where collection and use of such information is permitted, then investigate why differences occur.

Fairness analysis is not a single parity number. Different groups may have different feature availability, intervention types, response times or ability to appeal. A decline in aggregate false positives can coexist with a severe increase in one thin-history segment. The owner should compare the incumbent and challenger on the same population and report the trade-offs by product and channel. A reviewer should also look at the customer journey after a referral: whether a secure contact method is available, whether an alternative proof path exists and whether an unresolved case sits indefinitely.

A relationship between a feature and an outcome does not establish a permissible use. Teams must evaluate legal restrictions, purpose, relevance, data quality and explanation. The same event may be useful to stop a suspected account-takeover payment but inappropriate as a reason to decline credit. A customer-service note may contain unstructured sensitive information; summarising it with a model can introduce errors and expose unnecessary detail. A human can use the note to route support without making it an unrestricted training field.

Interventions themselves change the data. If the bank blocks a payment, it cannot observe whether it would have been fraudulent after execution. A step-up challenge may deter a scam or simply inconvenience a legitimate customer. A retrospective "loss avoided" metric is therefore an estimate requiring assumptions, not a measured fact for every blocked item. Monitoring can track confirmed cases, customer confirmations, complaints, manual overrides and subsequent outcomes, while clearly separating observed from estimated effects.

The NIST AI Risk Management Framework is a voluntary framework for mapping, measuring, managing and governing AI risks. It supports attention to intended use, affected people, performance and monitoring. It is not a banking regulation or an approval to use a particular customer feature. For U.S. banking organisations within its scope, Federal Reserve SR 26-2 provides revised interagency model-risk guidance. Its applicability and proportionality must be checked for the institution and use. The bank's legal and risk owners decide the required controls for its jurisdiction.

One customer, several decision owners

Return to Leela's case. The access team verifies the new device and owns authentication evidence. Payment operations owns the instruction state and scheme response. The fraud team owns the risk assessment and any permitted intervention. Customer support owns the secure contact and complaint record. Compliance owns a separate assessment if transaction activity raises an issue under its programme. These teams can share event identifiers, but they should not share a single undifferentiated "risky customer" status. Such a status obscures what was observed, which purpose justified review and who can clear it.

Imagine a second event a week later: Leela's salary credit is late, and an instalment falls due. A credit team may check account conduct and contact history under its policy. It should not treat the prior scam alert as a fact of inability to pay. The payment team should not infer fraud from the late salary alone. A support agent may see that Leela asked for help and route her to a suitable assistance process. The bank's record should distinguish the new credit evidence from the earlier fraud investigation, with access limited to what each role needs. The customer could have a normal payroll delay, a changed employer, or actual difficulty; the correct next step depends on verification.

Now suppose Leela disputes the Sunday payment after it was executed. Operations needs the network and ledger outcome, not only the fraud score. The fraud team may revisit the case with new customer evidence. A model owner may later use the confirmed label in validation and training. Customer support must tell Leela what the bank knows and what remains unresolved, subject to its dispute process and local rules. The system should link these events without changing the historical pre-payment score or suggesting the bank possessed the later evidence at the earlier decision time.

This separation prevents a dangerous feedback loop. If a caseworker marks Leela's account "high risk" because the model referred a payment, and the model later trains on that caseworker label, apparent accuracy may come from the system repeating its own suspicion. A trustworthy outcome record shows the independent evidence behind confirmation. The bank can then evaluate whether the original signal helped detect harm, caused unnecessary friction or revealed a data defect.

Decision precedence and fallback

An event-driven control needs a precedence order. Identity, eligibility, sanctions or other mandatory checks remain governed by applicable law and bank policy. A behavioural score may route an otherwise eligible transaction for additional review or authentication. A human may resolve a case within defined authority. A missing score is neither a pass nor a failure unless the approved fallback says so. The bank should document what happens when multiple controls produce conflicting results and what the customer sees during each state.

For a payment, a model response can arrive before or after the last reversible step. The policy must specify a response deadline, the acceptable feature age, and an action if either is breached. An approved fallback might request step-up verification, apply a simpler rule or defer a channel action where permissible. It should not promise a manual review before instant execution if staff and scheme timing cannot support it. A retry must not create two executions from one customer instruction. The bank can test the decision service and the payment system's idempotency separately.

For credit, the clock may allow a different fallback. A missing bureau response or income record may lead to a pending application and request for evidence. The model should not silently replace missing verification with a low-risk guess. If a customer is declined, the final reason must reflect the actual decision rules and evidence, not an unrelated behavioural flag. In account monitoring, the bank may review an early-warning signal before customer contact; it need not turn every unusual data point into an automatic limit change.

The hand-off record should retain source event time, ingestion time, scoring time, decision time, model and policy versions, rule hits, feature freshness, action, override, customer message and final outcome. Access to raw device or service data should be restricted. A later correction should be attached with its own timestamp and reason, leaving the original decision replayable. The audit question is precise: "What information was available when the bank acted, and who had authority for that action?"

Monitoring the decision and its consequences

A model dashboard must report the full funnel. Count eligible events, events with usable data, scored events, high-risk results, interventions, completed human reviews, final bank actions and matured outcomes. If an upstream feed fails, the proportion scored can fall while measured fraud precision appears to improve. A dashboard that only displays precision would miss the service failure. Segment the funnel by product, channel, customer tenure and relevant cohorts. Label the denominator and observation cutoff beside each rate.

Operational indicators arrive quickly: source-file delay, stale features, scoring timeout, fallback rate, queue length and unresolved status. Harm outcomes arrive later: confirmed fraud, recovery, charge-off, dispute resolution, complaint or customer attrition. A bank should not merge the clocks into a single green status. The owner sets alert thresholds based on the approved use and investigates change before retuning the model. A new fraud pattern, product launch, source mapping defect and different customer mix can all shift scores for different reasons.

Suppose a mobile app release changes device identifiers. The share of "new device" signals doubles overnight. The team checks release timing, enrolment records and other independent indicators. It can isolate affected decisions, apply an approved fallback or roll back a feature version. Retraining immediately on the changed data would risk normalising the defect. The incident record should capture detection time, affected population, customer actions, containment, correction and follow-up validation. A reviewer can then separate model drift from broken instrumentation.

A threshold change should be tested against queue capacity. If referrals rise from 500 to 700 per month, the extra 200 are not evenly distributed across hours. A burst on a weekend may overwhelm a team even when the monthly total looks manageable. The bank should forecast peak cases, staffing, action deadlines and the cost of abandoned or late reviews. It can choose a different intervention for cases that staff cannot examine in time, but that choice needs approval and customer-impact assessment.

Model performance monitoring also needs an incumbent comparison. Apply both versions to a dated shadow cohort with identical eligibility and feature availability. Compare confirmed harmful outcomes, false referrals, loss value, customer complaints and segment effects. Record known label gaps and the outcome maturation window. A higher area-under-curve statistic does not alone show that a threshold or policy improves the bank's decision. Release authority should review the business effect as well as statistical performance.

A business analyst's requirements pack

The BA can translate a behavioural use case into six linked artifacts. First, a decision definition names the event, timing, product, eligible population and owner. Second, a data contract lists fields, meanings, effective timestamps, freshness limits and rejected-record handling. Third, a policy matrix states mandatory controls, model result bands, permitted actions, precedence and fallback. Fourth, an evidence record describes identifiers, versions and retention. Fifth, acceptance cases exercise ordinary, boundary and failure scenarios. Sixth, a monitoring plan names denominators, outcomes, segments, owners and response actions. These artifacts may live in different tools, but their terms must agree.

A requirements review should challenge one concrete decision. If a new-payee signal is missing, does the channel stop, continue with a fallback, or ask for verification? If a customer is already under an authorised protective hold, can a score release the payment? If a caseworker overrides an alert, where is the reason stored and when does it expire? If a payment status remains unknown, which team reconciles it and what is communicated? Each answer should be traceable to an approved policy and a test. The BA does not choose a regulatory rule from an AI suggestion.

A test set can use Leela's journey with variations: verified device versus unverified device; new payee created in another channel; duplicate request; feature feed delayed past the deadline; caseworker contact completed before and after execution; customer later reports coercion; investigation reverses an early label. The expected result includes both bank action and customer message. It also checks that a correction changes future learning without altering the earlier decision record. That is stronger than asserting only that a scoring API returned 200.

A complaint that tests the whole chain

An illustrative complaint arrives from Leela: "Your app said my payment failed, so I sent it again, and both transfers left my account." The bank first determines the status of each instruction, not whether the fraud model was accurate. The payment team correlates the customer's requests, bank instruction IDs, network acknowledgements, ledger postings and beneficiary confirmations. It asks whether the first status was unresolved rather than failed, whether the second request reused the idempotency key and whether a duplicate control should have intervened. The customer-facing correction and any recovery or dispute route depend on the verified outcomes and applicable rules.

Behavioural data still matters. The second transfer may have been flagged as unusual because it repeated an amount within minutes. If the model referred the second request but the UI labelled the first as failed, the bank cannot celebrate a fraud alert as a successful control. The customer's repeated action was induced by an inaccurate status. The root cause may lie in messaging, state management or reconciliation. The incident review should separate the model output, policy action, payment execution and UI statement. A shared identifier and ordered event timeline make that possible.

The bank then examines whether similar customers experienced the same sequence. Search for unknown statuses followed by duplicate amount and payee combinations within a defined period, but avoid assuming every such pair was an accidental duplicate. Some customers intentionally make two equal transfers. Investigators sample cases and check network outcomes. The bank can quantify confirmed impact, potential impact and uncertainty separately. If a design defect is found, product and operations owners decide remediation and communication under their policy. The model team may improve a feature, but it cannot repair an incorrect payment-state promise alone.

This complaint is a useful acceptance test for the behavioural system. The original risk decision must remain replayable. The later complaint and correction should update case and learning records without altering what the bank knew at the time. Staff should be able to explain to Leela whether funds moved, what the bank is doing and what remains uncertain. A model score is one piece of evidence, not the answer to a customer's double-payment problem.

Contrasting three forms of behavioural change

A sudden change can reflect harm, opportunity or a data error. In a scam scenario, a customer adds a new payee after a suspicious message and attempts an atypical transfer. A timely, proportionate contact may prevent loss. In a life-event scenario, the customer starts a business and receives many new payments; the bank may need to refresh its understanding of expected activity and offer suitable services. In a data-error scenario, a channel changes transaction codes and makes ordinary transfers appear as cash withdrawals. The same aggregate spike requires different actions in each case.

The distinction begins with independent evidence. A verified customer report can support a scam investigation. A documented business change can explain volume. A source-system release and record reconciliation can prove a mapping defect. A statistical anomaly only tells the bank that observed data differ from a comparison. It cannot identify which explanation is true. The design should route the event to a person or process capable of obtaining the missing evidence and correcting the record. An automatic adverse decision based solely on the anomaly can turn a data defect into customer harm.

Even stable behaviour can hide risk. A criminal may imitate a customer's usual amount and channel, or a genuine customer may be persuaded to authorise a transfer using familiar devices. A model that relies on "different from personal history" will miss some harms. The bank needs controls based on transaction, payee, network, customer contact and other appropriate evidence. Conversely, a new customer has little history; the system should not label the absence of data as risky behaviour without an alternative path. Test both failure modes.

The bank should calibrate intervention to uncertainty and reversibility. A warning that asks a customer to review payee details has different consequences from freezing an account. A human referral during a reversible step differs from a post-execution alert. The action policy should state what evidence is enough for each step and how the customer can challenge it. This is a policy decision under local law and product rules, not a number that a generic model can set for every bank.

Validation and governance questions

An independent reviewer should know the development population, time periods, outcome definition and excluded records. They should reproduce feature values on dated cases and compare performance on an independent period. The review examines calibration, discrimination, stability, missingness and error by relevant segment. It tests sensitivity to a changed payment mix, a new channel and delayed labels. It also checks the decision policy around the model: thresholds, human capacity, fallback and complaints. An accurate model can still be unsuitable for a workflow with no usable intervention.

The bank's model inventory should state what the output is for, where it is used, who can change it and what would trigger suspension. A feature or policy change deserves versioning even when model weights stay fixed. A release pack should include approved data contracts, validation findings, test evidence, deployment version, rollback owner and monitoring plan. If a critical feed fails, operations must know whether to disable scoring, use a limited-data model, invoke a rule or hold a journey where permissible. No fallback should be invented during an outage by the engineer on call.

Customer-impact review is not an optional sentiment exercise. Count delays, abandoned journeys, appeal outcomes, repeated authentication challenges and complaints attributable to the control. Compare them with confirmed prevented or detected harms, stating where benefit is estimated rather than observed. Include frontline feedback about confusing reason messages and cases that cannot be resolved before the decision deadline. A threshold that reduces loss while creating an unmanageable queue or disproportionate friction needs redesign.

The final review question is whether the bank can explain one affected customer's journey without exposing other customers' data. Select a decision ID and reconstruct the event order, features available, required rule results, model version, policy action, human review, customer message, payment or credit outcome and later corrections. If that chain cannot be reconstructed, the bank cannot reliably measure whether behavioural analytics helped. Repair the data and workflow contract before broadening automation.

Further reading

Related learning paths

This application uses JavaScript for the full interactive experience. This text summary is served for accessibility and search indexing.

How customer behaviour changed the risk problem · Malla Banking Academy