When AI supports decisions and when it must not decide. A practical lesson in why banks turned to ai for banking and payments practitioners.
A recommendation becomes consequential at the action boundary
A bank can use artificial intelligence (AI) to find a policy passage, summarize a case, rank an alert, estimate a risk or propose a next action. Those outputs do not have the same authority. The critical boundary is the moment an output changes a customer's position, a payment's status, an investigator's conclusion or a regulatory record. A model can be excellent at ranking cases and still be unsuitable for making the final decision. A bank has to define the action, its owner and the evidence required before choosing how much autonomy to give the system.
Imagine a fictional bank with three workflows. An underwriter receives a modelled credit-risk estimate on a loan application. A payments service receives a fraud score before an instant transfer. A financial-crime analyst receives a generated summary of unusual account activity. Each output may assist, but a final credit decline, payment intervention or suspicious-activity reporting decision has its own policy and legal context. The same label, "AI decision," conceals those differences. This lesson follows the bank action rather than the model architecture.
"AI must not decide" is not a universal statement that all bank decisions legally require a person. Some defined actions can be automated under applicable law and approved policy. Conversely, placing a person on an approval screen does not make a poorly understood process safe. The boundary depends on jurisdiction, use case, data, consequence, reversibility, required checks, explanation and the reviewer's real ability to challenge the output. A bank should not present an internal risk preference as a legal prohibition, or treat the absence of a specific AI prohibition as permission to skip existing obligations.
The practical test begins with one sentence: "At this point, the bank will [action] for [population] based on [evidence] under [policy]." If the team cannot fill those brackets, it has not specified a decision. If it can, it can then ask which part is deterministic, which part can be model-assisted, which part needs human judgment and which facts must be communicated or retained. The answer may differ for the same model used in a different product.
Five levels of model involvement
A useful operating taxonomy separates five levels. Information support retrieves or organizes material without changing a customer record. Prioritization orders a queue, but staff still investigate and decide. Recommendation proposes a reasoned action for a named reviewer. Bounded automation performs a specified action within approved limits and exception rules. Autonomous execution lets the system initiate and complete an action with limited contemporaneous human control. These levels describe workflow authority, not model sophistication. A simple rule can execute autonomously; a complex model can remain a research aid.
Information support is not risk-free. A policy assistant may cite an outdated rule, omit an exception or reveal confidential data. Its interface should show source identity, version and the exact passage. Staff need a way to inspect the source and record when it was wrong. A retrieval result that is later pasted into a decision memo has crossed from convenience into evidence. The bank should define whether the source is authoritative and whether the generated summary can be retained in a case file.
Prioritization changes who receives attention first. A fraud queue model that ranks suspected scams can reduce delay for some cases, but an overloaded queue may never reach lower-ranked customers. The bank should measure coverage, waiting time, stale cases and harm among cases it did not review. An algorithm may not formally "decide" whether to investigate, yet its ranking can create a de facto decision through capacity. The owner must set a minimum review policy for mandatory categories and a way to detect systematically neglected groups.
Recommendation adds a proposed action. A credit reviewer may see "refer for income verification," with the dated evidence and reason. The reviewer can inspect the underlying record, request more information, disagree and document why. A recommendation that appears as a large green "approve" button while the contrary evidence is hidden is not neutral assistance. Test the interface for automation bias: how often do reviewers accept a model output despite an injected source error or a conflicting rule?
Bounded automation can be appropriate for a reversible, well-specified action with measured performance and clear exception handling. For example, a bank might automatically route a routine reconciliation exception to the correct work queue using a model, while a person decides whether to post an accounting adjustment. The routing has its own error cost, including delay. If the model's confidence is insufficient or a source is missing, it should use an approved fallback. The bank still owns the outcome and must monitor the automated path.
Autonomous execution demands closer scrutiny as consequence and irreversibility increase. A model that directly releases an instant payment, declines credit, freezes an account or files a regulatory report would cross different control boundaries. No single rule says those four actions are equivalent. A bank may automate parts of them under applicable frameworks, but it cannot infer authority from the model's accuracy alone. The decision inventory should record each action's legal basis, mandatory rules, delegate, customer remedy, audit trail and ability to pause the automation.
A loan application: score, policy and reasons
Consider a fictional applicant, Nila, seeking a small unsecured loan. She supplies identity and income information. A credit model estimates a probability of a defined default outcome over a stated horizon. The estimate is a risk signal, not a complete approval. The bank separately checks product eligibility, affordability, identity, any applicable restrictions and the approved lending policy. If the income feed is unavailable, the model may still return a number using other variables. That number must not be presented as verified income.
A useful decision record separates four layers. The source snapshot records what was available at the application time. The model layer records feature definitions, version, score and its interpretation. The policy layer records eligibility, affordability and threshold results. The action layer records approval, request for evidence, referral or decline, including a human override where authorized. The final reason must reflect the layer that actually drove the outcome. If affordability failed while the score was acceptable, the bank should not tell the customer that the score was too low.
For covered U.S. credit decisions, Regulation B, 12 CFR 1002.9 requires specific principal reasons for adverse action. A complex algorithm does not remove that obligation. The requirement's applicability and notice form depend on the transaction and creditor; this example is not global legal advice. The bank should test its reason-code mapping on combinations of model score, rule failure, missing evidence and manual review. A plausible generated explanation is inadequate if it describes a factor that was not a principal reason for the actual decision.
A human underwriter can add contextual evidence, but the role must be real. The person needs access to the relevant dated source records, the policy, the model's known limitations and the authority to request evidence or change the action. An override needs reason, evidence, approver and later outcome analysis. If the reviewer sees only a score and a timer, signing the case does not cure opaque decisioning. Conversely, mandatory referral of every application can slow service without adding judgment if staff simply repeat the score.
Credit performance is observed later. A loan that performs for six months does not prove the original decision was correct for all applicants; a declined applicant has no repayment label in the bank's book. Validation must account for the selected funded population and defined outcome window. The risk model should be compared with an incumbent policy on comparable dated cohorts. Monitor both credit losses and customer impacts, including referrals, missing-data outcomes, complaints and reason-code accuracy. Do not optimize approvals by hiding a rise in harmful or unexplained declines.
A payment: the deadline changes what support can do
A customer initiates a fictional instant transfer to a new payee at 02:10. The bank has a brief interval to apply authentication, account, scheme and other mandatory controls before execution. A fraud model may combine point-in-time device, beneficiary and transaction signals. It can recommend a step-up, hold or review where the actual rail and policy permit. A score returned after the last preventable point can support investigation, but cannot retroactively prevent the original payment. The BIS CPMI fast-payments report explains the speed and availability characteristics; it does not prescribe a universal intervention deadline.
The policy engine must define precedence. A required screening or eligibility rule cannot be overridden by a reassuring model score. A high fraud score should not be treated as proof that the customer authorized a scam. If step-up verification is allowed, the channel must show the real payment state and a way to complete or cancel the instruction safely. If a model dependency times out, the bank follows an approved fallback, which may differ by amount, customer or rail. "Fail open" and "fail closed" both carry customer and risk costs; neither is a universal answer.
Payment idempotency is separate from scoring. If the customer retries a spinning screen, the first instruction may have been accepted. The bank must not create a second transfer because an inference call was late. A repeat scoring request can return a consistent decision, while the payment platform controls duplicate execution and enquiries about final status. An "unknown" rail outcome must not be rendered as a certain failure. The decision log needs the instruction ID, timestamps, rule hits, score version, final action, customer message and eventual rail outcome.
Review the errors in both directions. A false block can strand a customer's urgent payment and create a complaint. A missed scam can cause loss even when the model's overall accuracy looks high. Compare interventions by amount, new-payee status, channel and vulnerable customer journeys where data and law permit. The staff queue needs capacity at the hours when referrals arrive. A model that sends a case for review after a non-deferrable payment deadline has not provided an executable control.
Financial-crime alerts require investigation
A transaction-monitoring model can rank unusual activity or identify a pattern a fixed rule misses. The alert is a starting point for a defined review process. It does not establish that an account holder committed a crime, that a report must be filed or that a relationship must be terminated. A U.S. bank can use FFIEC BSA/AML Examination Manual guidance on suspicious-activity reporting to understand the monitoring, investigation and reporting control expectations in that jurisdiction. The exact legal and supervisory duties elsewhere must be checked locally.
Consider a fictional alert linking repeated small transfers to a group of recipients. The model's graph may be useful, but an analyst must check identity resolution, time windows, legitimate customer context, duplicates, reversals and earlier cases. A shared address or device does not by itself prove a criminal relationship. The case system should distinguish observed transactions, model inference, analyst hypothesis and verified external information. It should also allow an inconclusive disposition. Forcing every case into "suspicious" or "clean" turns uncertainty into a misleading training label.
A generated case summary can save reading time but may omit a counterexample or invent a transaction. The analyst should see source-linked facts and be able to inspect the underlying records. If a summary says "three payments" when the fourth was reversed, the system should distinguish executed, returned and attempted events. Its citation needs an event ID and timestamp, not merely a confident paragraph. The case should retain the version of the summary, any corrections and the human rationale. Sensitive reporting information must remain within the authorized workflow and access boundary.
The reporting decision belongs to the authorized function under the applicable process. A model may prioritize which cases receive attention and identify additional context. It should not silently file, suppress or withdraw a report merely because a probability crossed a threshold. If the bank chooses any automation around case closure, it needs a specifically approved scope, controls for mandatory review, audit evidence and periodic sampling of closed cases. A fall in alert volume can reflect better prioritization, a broken feed or missed risk; investigate the reason before calling it an improvement.
Generative AI is an assistant with an evidence problem
A generative model can draft an internal memo, retrieve relevant policy, compare records or propose test cases. It may produce fluent text that is wrong or unsupported. The NIST Generative AI Profile discusses confabulation and other risks in such systems. NIST's framework is voluntary and cross-sectoral. It informs risk management, but does not replace bank policy, law or model governance for a specific use.
For a policy assistant, retrieval should point to a controlled corpus with document owner, version, effective date and access rule. The answer should link the exact passage used. If a policy was superseded, the assistant should not blend the old and new instructions. If the relevant source is absent or contradictory, it should say so and route the question to an owner. A model that cites a real document but invents the proposition attributed to it is still unreliable. Validate both retrieval and answer fidelity on representative questions and known exceptions.
A bank should decide which information may enter an external or internal model service. Customer records, investigative notes and confidential policies require approved access, retention, logging and vendor controls. Redaction may reduce exposure, but a redacted narrative can still identify someone through context. A prompt-injection string inside a retrieved document is untrusted content, not an instruction from the bank. The system should isolate retrieved material from operating instructions and prevent a generated answer from invoking an action outside the user's authority. Test the boundary with a document that instructs the assistant to send data elsewhere or change a case status.
The output interface should show that a draft is a draft. A human reviewer needs source links, uncertainty, unresolved conflicts and a way to edit or reject it. The reviewer should not be forced to sign a polished memo after a perfunctory glance. Monitor unsupported citations, omissions, wrong policy versions, unauthorized data exposure and reviewer correction rates. A model upgrade may change wording and error patterns without changing its API; retest the relevant use cases and preserve a rollback path.
The action boundary can move without a new model. A summary copied into a customer letter has become a communication. A recommendation displayed beside a one-click approval button can influence outcomes. A generated test case accepted as the only coverage can let a missing exception reach production. Governance therefore tracks the whole workflow, including UI affordances and downstream reuse, not merely whether the model API technically returns "text" or "score."
EU creditworthiness classification and the changed timeline
The EU AI Act lists AI systems intended to evaluate the creditworthiness of natural persons or establish their credit score in Annex III, with an exception for systems used to detect financial fraud. Classification depends on intended purpose and the detailed legal rules, not on a bank calling a system "decision support." The official Annex III text gives the scope. A bank should map its specific model, deployment and role before claiming an exemption or applying a high-risk obligation.
As checked on 29 September 2026, the European Commission says an AI Omnibus entered into force on 27 July 2026 and moved the Annex III high-risk AI rules' application date to 2 December 2027. The Commission's implementation update is the source for that date. The general AI Act and other obligations have different start dates. A banking team must not copy an earlier 2 August 2026 date into an Annex III credit model checklist without checking the amended framework and any transitional provisions.
The high-risk classification is not a statement that a person must manually approve every credit application. The framework specifies risk management, data governance, documentation, logging, transparency, human oversight, accuracy, robustness and other obligations for systems in scope, according to the applicable dates and roles. A deployer and a provider may have different duties. Existing credit, data-protection and consumer rules may already constrain a decision before those particular high-risk provisions apply. Legal counsel should determine the exact requirements for the institution and use.
The fraud-detection exception in Annex III 5(b) also has a boundary. A model genuinely used to detect financial fraud is different from a creditworthiness model relabeled as "fraud." If its output drives eligibility or pricing, document the actual intended purposes and effects. A shared platform can host both functions, but its documentation, access, validation and decision records should not conflate them. The classification exercise is part of product design, not a footer added shortly before release.
A U.S. supervisory boundary
The U.S. model-risk framework also changed in 2026. Federal Reserve SR 26-2, issued with the OCC and FDIC, supersedes SR 11-7 and SR 21-8. The letter says it is expected to be most relevant to Federal Reserve-regulated banking organizations with more than $30 billion in total assets. Its revised guidance emphasizes risk-based practices tailored to a bank's model-risk profile and complexity. This is a scoped supervisory source, not a universal AI law or a statement that all small banks must use the same validation program.
The bank should document model purpose, assumptions, limitations, data, testing, use and governance in proportion to risk. A fraud score supporting a reversible referral and a credit model driving an adverse decision need different controls even if both are statistical models. A generative assistant may create additional use and operational risks that require assessment. A vendor's claim that a model is "compliant" does not establish that its particular deployment, threshold, training data and workflow meet the bank's obligations.
Regulation B's adverse-action requirement is separate from model-risk supervision. A bank cannot replace a principal reason with "the AI decided." It also should not claim that a model explanation technique automatically produces legally sufficient reasons. Validate the actual reason path, including policy overrides and missing evidence. The audit trail needs the decision-time inputs, model version, policy version, reviewer action and communication. A later model update must not rewrite what the bank knew when it acted.
These EU and U.S. examples teach scope discipline. A globally accessible course cannot prescribe one human-approval rule for every credit, payment or monitoring action. It can show how to identify intended use, jurisdiction, controlling rules, consequence and evidence. The owner should obtain a current legal determination for the deployment. Meanwhile, the delivery team can build an architecture that separates recommendation from action and makes the boundary testable.
Human oversight has to change the outcome when warranted
A reviewer is effective only if they can understand the question, inspect decisive evidence, notice limitations, disagree and act before the deadline. The bank should define the role's authority, training, workload, escalation route and expected evidence. A human who sees only "approved by AI" cannot exercise independent judgment. A reviewer who has ten seconds to read a long case before an irreversible action may be present in the workflow but absent from its control.
Use a counterfactual review test. Give staff a case with a deliberately wrong model recommendation and a visible contradictory source record. Measure whether they find the conflict, correct the proposed action and record a reason. Repeat with a missing source, a stale document, a rare customer circumstance and an ambiguous label. Do not evaluate only how quickly reviewers click through ordinary cases. Review quality includes appropriate disagreement and evidence requests.
The bank can reserve human effort for cases where it adds value. A low-consequence internal routing action might be automated, with sampling and correction. A credit application with conflicting income evidence might require an underwriter. A payment approaching an execution deadline may need a customer step-up or an approved deterministic fallback rather than an unavailable analyst. A suspicious-activity report needs its authorized review process. The route depends on actionability and applicable rules, not on a blanket preference for or against humans.
An override should preserve the model's original output and the reviewer's decision as separate events. Record who had authority, what evidence changed the judgment, when the action occurred and the eventual outcome. If overrides are concentrated in one customer segment, investigate data quality, model performance and policy. If reviewers almost never disagree, that may reflect an excellent model, an interface that discourages challenge or an ineffective review step. Sample cases to distinguish those explanations.
A customer challenge path is a further control. A false payment block may need immediate status correction; a credit decision may need a notice and reconsideration procedure; a case-management error may need confidential correction. The bank should not promise that every outcome is reversible. It can preserve the original record, assess harm and take remediation within its authority. Support staff need accurate action status without access to restricted case details.
A worked capacity and harm exercise
Suppose a fictional fraud queue receives 10,000 eligible payment attempts during a day. A new model refers 200 before the execution deadline, compared with 120 under the current policy. Analysts can timely review only 150. Historical matured investigation data suggests 20 confirmed fraud cases among the 200 new referrals, versus 14 among the 120 current referrals, on the same eligible cohort. These classroom numbers are not expected industry rates and do not prove future performance.
The proposed referral rate is 200/10,000, or 2%. The current rate is 1.2%. The challenger adds 80 referrals and finds six more confirmed cases in historical replay, an incremental 7.5% observed yield among the added 80, if the six are indeed in that added set. But capacity leaves 50 new referrals unreviewed by the deadline. If those 50 are simply released, the designed "human in the loop" path does not operate for them. If all 50 are held, genuine customers may experience delays and complaints. The policy must specify which cases get immediate attention, what happens to the overflow and what the customer sees.
Do not divide 20 confirmed cases by 200 and call it the real-world precision without label maturity and investigation coverage. Some referred cases may not be resolved. The model may have influenced which cases investigators examined, so observed labels are selected. Compare losses, prevented harm, friction, queue age and complaints by dated cohort. A prospective shadow or limited release can gather evidence under controlled conditions, but even a good replay does not make an unstaffed queue safe.
Now change one assumption: the same model ranks cases for post-event investigation rather than pre-payment intervention. A next-day queue may have more review capacity and a different action. The accuracy numbers can be identical while the operational value changes. The model contract must state whether a score arrives before or after the preventable point and whether the proposed action is possible. A dashboard should not aggregate pre-execution prevention and later detection into one "AI stopped fraud" total.
A credit model has another denominator. If 1,000 people apply, 600 are funded and 18 later default under a defined window, the observed default rate is 18/600, or 3%, among funded loans. It says nothing directly about outcomes for the 400 who were not funded. A policy that rejects more applicants may lower the observed funded default rate while making unfair or unexplained decisions. The review must examine the applicant flow, the source of reasons, customer impact and the selection limit alongside portfolio performance.
The decision contract an analyst can write
A business analyst can express each use as a decision contract with a trigger, eligible population, data cutoff, available signals, deterministic rules, model output, allowed actions, authority, deadline, fallback and evidence. Add the customer or regulatory communication, challenge path, monitoring measures and stop condition. This forces an abstract "use AI" initiative into an implementable banking workflow. A contract should be short enough that product, risk, operations and technology can agree on each field.
For the fictional loan, the trigger is a completed application; the cutoff is the application decision time. The model returns a calibrated estimate within its validated population. Eligibility and affordability are separate policy checks. The action can approve, request evidence, refer or decline. The creditor retains the actual reasons and sends any required notice. If the income feed fails, the system does not equate missing income with zero income or invent a reason. The approved fallback requests evidence or routes to review.
For the payment, the trigger is an authenticated instruction and the deadline is the last preventable point on the specified rail. Mandatory controls have precedence. The fraud model returns a score with timestamp and feature availability. The policy can allow, step up, refer or reject only where those actions are permitted and executable. A timeout follows a versioned fallback. The log connects instruction ID, score, rules, final action and rail result. A retry does not create a second payment.
For the monitoring case, the trigger is an alert or analyst observation. The model may rank and summarize; the authorized analyst investigates. The case disposition, escalation and any reporting decision follow applicable policy. A failed model service does not suppress mandatory monitoring coverage. The case record distinguishes generated text from verified evidence. Access controls prevent a customer-facing agent from seeing restricted details.
A decision contract also specifies who may change it. A data owner approves feature meaning, a model owner manages development, a validator challenges performance, a policy owner sets actions and operations owns the queue. Segregation depends on the bank's operating model. A threshold change can have the same customer impact as a new model. The bank should compare expected volume and harm before approval, deploy a versioned change and be able to roll it back without altering historical decisions.
Test the boundaries before release
Test cases should target transitions that can turn advice into action. Start with a routine case where all sources are available, rules pass and the model responds in time. Assert the output type and the final bank action separately. Then introduce a required rule failure while the model returns a favorable score. The rule must retain its precedence. Introduce a high score with missing or stale features. The policy should use its declared missing-data path, not an accidental default value.
Test time. Deliver a fraud score just before and just after the payment decision deadline, then confirm the latter cannot be counted as prevention. Deliver a credit data correction after an application was declined. The bank should preserve the original decision snapshot and assess whether reconsideration or customer remediation is needed. Change a policy document's effective date between retrieval and reviewer action. The assistant should show the version it relied on and prevent an obsolete passage from masquerading as current guidance.
Test conflict and contest. Put a contradictory document into a generated case summary. Verify that the reviewer can inspect both passages, mark the summary wrong and record the corrected rationale. Place prompt-injection text in a retrieved document and check that it cannot send data or change an account. Present an appeal from a customer whose transfer was delayed. Support should see the actual payment state and permitted next step, not a model accusation or an inaccurate success message.
Test failure. Disable the feature service, model endpoint and case system separately. Confirm that the policy uses the approved fallback, the queue retains mandatory work and the audit record identifies the failure. Exercise peak-hour volume so the expected human-review path is actually staffed. Test a rollback after a threshold change and reconstruct decisions made under both versions. Run accessibility and mobile checks on the review screen: a hidden reason or unreachable override is an ineffective control.
A release pack should contain the inventory of decisions, current source and legal scope, intended population, data lineage, validation evidence, approved policy, action precedence, reviewer training, fallback test results and monitored outcomes. The bank should name an owner for each open limitation. If an essential source cannot be reproduced, a required rule can be bypassed, or a consequential action lacks an accountable owner, the use should not move into that action level.
Monitoring changes the autonomy decision
After deployment, monitor the entire chain rather than a single model metric. Track input completeness, feature age, score distribution, rule exceptions, action rates, human override reasons, queue backlog, customer complaints and matured outcomes. Segment by product, channel and relevant population. A shift may come from customer behavior, a policy change, a broken feed or a new attack. Investigate the cause before retraining. An automatic retrain on corrupted labels could make the failure harder to detect.
Set stop conditions before launch. A bank may pause an automated path when feature freshness falls below an approved limit, a required control is unavailable, queue capacity is exceeded or a material error rate threshold is breached. The thresholds themselves belong to the specific bank's risk appetite and service commitments; this chapter does not invent universal numbers. The incident owner needs a way to locate affected decision IDs, apply a safe fallback, preserve evidence and assess remediation.
Suppose a feature mapping error starts treating "income not received" as "income zero." The credit model now refers many applicants, and staff accept its recommendation without checking the source. Monitoring should detect the missing-data spike and changed referral mix. Incident response identifies the deployment time, affected applications and communications; stops the bad path; restores the correct mapping; and reviews customers harmed. Simply training a new model would not fix the source semantics or the earlier adverse actions.
A model may also degrade gradually while technically available. A fraud ranking may begin to disadvantage customers with a new device type because its development population was older. Compare score and intervention distributions with confirmed outcomes and complaints once labels mature. If the bank cannot yet observe outcomes, use leading operational measures and state that outcome quality remains unknown. Change autonomy only when evidence supports it. Moving from advice to automatic action is a new policy decision, even if the model bytes stay the same.
An analyst's final review
For any proposed banking AI use, ask what the system actually changes. Identify the specific customer, transaction, case or regulatory record. Name the decision owner and the model's role. Draw a line between source fact, inference, recommendation, policy action and final outcome. Record when each was known. Check whether required rules and customer rights apply in the relevant jurisdiction. Establish the permitted fallback and the path for challenge, correction and incident response.
A good recommendation can be accepted, challenged or ignored for a reason. A safe automated action has a narrow scope, verified inputs, policy precedence, monitoring and a way to stop. A prohibited or unsuitable action does not become acceptable because an output is fast, fluent or statistically impressive. The bank remains accountable for what it does with the output. Its evidence must let a reviewer reconstruct the action without pretending that a later outcome was known at decision time.
The practical result is a graded use of AI. Retrieve and summarize where source fidelity can be checked. Rank cases while measuring who is left behind. Recommend actions when a reviewer can exercise real authority. Automate bounded actions when the consequence, rules and fallback are understood. Reassess the boundary whenever the product, law, model, data or customer impact changes. This is more precise than promising that AI will always decide or never decide.
This application uses JavaScript for the full interactive experience. This text summary is served for accessibility and search indexing.