Improving STP rates using AI

Improving STP rates using AI. A practical lesson in business impact and controls for banking and payments practitioners.

Plain language meaning

Improving STP rates using AI explains how banks can reduce manual touchpoints by predicting missing data, identifying repairable defects, prioritising exceptions, suggesting enrichment and improving routing quality without bypassing mandatory controls.

This topic is about straight-through processing across banking operations. It includes payments, lending, onboarding and servicing where relevant, but it must not become only a payments-routing chapter.

In a real bank, this topic cannot be handled as a loose data-science or technology idea. It affects customer outcomes, fraud and AML control, operational queues, service continuity, privacy, security, model governance, audit replay, management reporting and regulatory confidence. AI should improve speed and quality, but the bank must still prove source data, permitted use, approved logic, human accountability, fallback handling and retained evidence.

Where it sits in the banking AI journey

This card belongs to Business Impact and Controls. The working flow is Operational break, AI defect classification, Repair or enrichment suggestion, Control check, and STP improvement.

Read the flow as a bank operating model. Each stage needs a source system, a data owner, a timing rule, a quality gate, a model or rule boundary, an exception path, a customer-impact view, a fallback option, a monitoring requirement and a retained record. That is what separates useful AI adoption from uncontrolled automation.

Banking data and evidence

The important data points are break reason, missing field, customer record, reference data, routing result, manual touch, repair code, and STP metric. These items matter because they can influence risk scoring, operational repair, fraud action, AML triage, customer treatment, reporting, model monitoring and management decisions.

The evidence pack should include break report, AI suggestion log, repair note, control result, manual override log, STP trend, and root-cause analysis. A strong bank can replay the journey from source data to transformed input, AI output, rule result, human action, system outcome and monitoring result. A weak bank only knows that a process ran and hopes the process was right.

Controls that make AI adoption safe

The core controls are mandatory-field rule, approved enrichment source, human review for material change, repair audit log, exception sampling, customer-impact check, and STP dashboard governance. These controls keep the topic anchored to banking purpose, approved policy, data governance, model-risk expectations, operational resilience, customer fairness, privacy, security and auditability.

The practical design should define what AI may recommend, what it must never decide alone, which deterministic rule remains authoritative, who owns thresholds and overrides, how degraded service is handled, how customer harm is detected and what evidence is retained. Without that control design, faster AI can simply make weak processes fail faster.

Architecture and data-operation lens

Banking AI depends on the architecture around it. Storage, streams, feature definitions, training sets, model versions, thresholds, feedback labels and rollback paths must be governed before the bank relies on AI output. The model is only one part of the control chain.

A bank-grade design connects channels, source systems, core records, payment hubs where relevant, fraud systems, AML platforms, case tools, data platforms, feature stores, model-serving endpoints, policy engines, audit logs and management dashboards. It also records degraded operation, recovery actions and lessons learned.

Regulatory and governance lens

Federal Reserve SR 26-2, dated 17 April 2026, gives revised model-risk guidance for traditional models and non-generative AI models used by banking organisations, including development, validation, monitoring, change control and governance.

The Federal Reserve's 2026 model-risk guidance states that generative and agentic AI are outside that guidance, while broader bank risk-management and governance practices still need to control tools and processes not covered by the guidance.

NIST AI RMF 1.0 uses Govern, Map, Measure and Manage functions for AI risk management, and NIST AI 600-1 adds generative-AI risk actions for source grounding, content provenance, data protection, cybersecurity and human oversight.

FFIEC Architecture, Infrastructure and Operations guidance expects financial-institution technology environments to be governed, resilient, secure, monitored, documented and aligned to business risk, including emerging technologies such as artificial intelligence and machine learning.

BCBS 239 remains current for effective risk data aggregation and risk reporting, and the Basel Committee's January 2026 newsletter reiterates the importance of accurate, comprehensive and timely data capabilities in banks.

FFIEC BSA/AML examination guidance expects suspicious activity monitoring systems and supporting technology to be risk-based, explainable by management, independently tested where appropriate and aligned to the bank's risk profile.

FinCEN's 12 June 2026 Section 314(b) materials clarify information sharing for possible terrorist activity, money laundering and fraud-related specified unlawful activity within the statutory safe-harbor framework for participating financial institutions.

OFAC's Framework for Compliance Commitments describes sanctions compliance programme components including management commitment, risk assessment, internal controls, testing and auditing, and training.

Diagram walkthrough

Read the diagram from left to right as Operational break, AI defect classification, Repair or enrichment suggestion, Control check, and STP improvement. It is a banking control map. The point is to show how data, AI or ML output, rules, human action, operational routing and audit evidence should connect.

Use it as a 30-minute study method. For each box, ask which system creates the data, which definition is used, which model or rule acts, what can go wrong, who can override it, how a fallback works, which customer or regulatory impact exists and what record proves the final state.

Most important mistake to avoid

The common failure is improving STP by hiding exceptions. True STP improvement removes avoidable breaks while keeping sanctions, AML, fraud, credit, accounting and customer-impact controls visible.

The correction is disciplined scope. Keep the chapter anchored to banking purpose, prove the data path, make ownership visible, test failure behaviour, record the evidence and make the final outcome explainable without relying on memory, assumptions or developer-only knowledge.

Measure straight-through processing at the right boundary

A bank receives 10,000 payment instructions in a day. It can claim a high straight-through processing rate only after defining which instructions are eligible, where processing begins, what counts as manual intervention and whether the payment reaches a final state. A message that passed an AI classifier but later entered an exception queue is not necessarily straight-through. If an AI system changes the denominator by filtering difficult cases out, the published improvement may be misleading.

Consider a repair assistant that suggests a missing beneficiary detail from approved data. It should identify the source and confidence of the suggestion; a bank-approved rule or operator decides whether a value can be used. Some fields may be safe to normalise deterministically, while changing a beneficiary account without evidence would be a different and potentially harmful action. The system records the original instruction, proposed change, reviewer or rule, resulting message and eventual settlement status. It must preserve the exception reason if the item cannot be repaired.

An evaluation compares a dated pre-change and post-change cohort, adjusting for channel, payment type, cut-off, seasonal volume and policy changes. Measure eligible items, auto-resolved items, manual reviews, returns, rejects, reversals, customer complaints and downstream reconciliation defects. A faster hand-off to another queue is not end-to-end improvement. Review a sample of automated resolutions for correctness and a sample of unprocessed cases for missed opportunity. Distinguish model errors from missing reference data or ambiguous instructions.

Acceptance tests include a duplicate instruction, expired reference record, conflicting beneficiary data, late settlement response and operator reversal. A fallback must preserve payment status and prevent duplicate execution when the assistant is unavailable. Define a process owner, control threshold and rollback rule before rollout. The outcome is better only if fewer appropriate cases need manual work without raising misdirected payments or obscuring exceptions. ISO 20022's external code-set governance is a useful example of controlled payment reference values, not proof that any local AI repair is authorized.

Define straight-through processing precisely

Straight-through processing, or STP, means a transaction or case completes an agreed workflow without manual repair. The definition must identify the start, end and exclusions. A payment accepted by a channel but held for screening is not necessarily straight through. A payment that settles and then returns may count differently from one that never left the bank. An AI project cannot claim improvement until the bank agrees on the denominator, lifecycle and what manual intervention means.

AI can help classify payment messages, predict repair needs, extract fields from documents, recommend routing or prioritize exceptions. It does not replace mandatory screening, authorization, reconciliation or payment-scheme rules. The model's role should be defined: does it suggest a correction to a trained operator, supply a confidence-scored field for review or automatically select a validated route under policy? The final system action and outcome must be recorded separately from the recommendation.

STP is a business outcome, not a model accuracy metric. A classifier can correctly predict many exceptions while STP stays unchanged because the repair queue is slow. An auto-repair can raise apparent STP while introducing returns or customer harm. Measure end-to-end completion, rework, holds, returns, cycle time and customer impact together.

Map the payment journey

A typical journey includes instruction capture, validation, enrichment, screening, routing, settlement or rejection, reconciliation and reporting. Identify the points where manual intervention occurs. Invalid beneficiary details, missing purpose codes, inconsistent names, format errors, insufficient reference data or ambiguous routing can create repair cases. Some are predictable from structured fields; others require document review or external confirmation.

Start with a process map by payment type, channel and corridor. A domestic transfer, card transaction and cross-border message have different rules and deadlines. Choose one bounded use case for an AI pilot. Count eligible instructions, clean completions, manual repairs, holds, returns and cancellations over a representative period. Separate upstream data defects from policy-required reviews. AI should not be credited with eliminating a review that was legally or operationally necessary.

Capture event and decision timestamps. An instruction may be accepted at 09:00, repaired at 09:30 and settled at 10:00. A channel status of "accepted" is not proof of settlement. Link business instruction IDs to AI output, human edits, payment-hub actions and later returns. Without lifecycle reconciliation, an STP rate can be inflated by counting only records that reached the hub successfully.

Pick the AI intervention

An exception classifier can predict which incoming messages are likely to fail validation or require human repair. It can route them to the right specialist sooner, but prediction alone does not make a payment straight through. Measure repair time and final completion. A document extraction model can read invoice or trade fields, but low-confidence or conflicting values should go to a reviewer. Preserve source spans and corrections.

A recommendation model can suggest a payment route or standardized code based on historical successful cases. It should respect deterministic eligibility, sanctions and scheme constraints. A historical route may no longer be valid after a correspondent or rule change. Version reference data and policy; a high model confidence cannot override an invalid destination.

A generative assistant can draft an explanation for a repair analyst or summarize an exception. It should cite authoritative payment data and policy, avoid inventing beneficiary details and require review before consequential use. A fluent draft can save reading time, but the effect on STP may be indirect. Report analyst productivity separately from transactions completed without intervention.

Data and labels

Training data should include all eligible instructions, not only successfully settled payments. A model trained on completed cases can learn a biased view of errors. Define labels such as validation failure, manual edit, returned payment and eventual clean completion with timestamps. A manual repair code may be inconsistent across teams; sample source cases and standardize taxonomy before modeling.

Features available at intake may include message fields, payment type, channel, amount, currency, beneficiary reference, prior exception history and source-quality flags. Do not use later repair notes or final settlement status in a pre-submission model. Preserve point-in-time reference mappings and effective rules. A model can appear accurate offline if it sees a field populated only after an analyst fixed the message.

Missingness matters. An empty purpose code can be a genuine omission or a field not required for a particular corridor. A failed reference-data lookup is not the same as an unknown beneficiary. Define conditional requirements and validity flags. Keep raw customer narratives protected and use minimum necessary data in model serving and logs.

Baseline and evaluation

Establish a current-process baseline by product and corridor. Report STP numerator and denominator, manual-touch rate, repair time, return rate, settlement time and customer complaints. Distinguish technical processing failures from compliance holds and customer cancellations. A period with fewer complex payments may show higher STP even if the process did not improve. Compare cohorts with similar mix or report adjusted and raw rates.

For the model, measure precision and recall of predicted repair needs, field extraction accuracy or recommendation acceptance. Then measure whether the intervention changes the workflow: fewer manual edits, faster resolution, fewer returns and no increase in wrong payments. A model can have high prediction accuracy but low operational value if it flags cases that would have been fixed automatically anyway.

Review costs and risks. An incorrect auto-filled beneficiary identifier can misdirect funds, whereas a mistaken queue priority may mainly delay processing. The acceptable confidence and human check differ. Evaluate rare but high-impact errors separately from averages. Use thresholds and fallback approved by payment and control owners.

Automated repair boundaries

Some repairs can be deterministic: normalize a known format, validate a check digit or populate a field from an authoritative source. An ML model can propose an ambiguous mapping, but automatic acceptance should require evidence and policy. Record original and proposed values, confidence, source, approval rule and final message. A rule-based fix may be more reliable than AI for a well-specified format.

Do not infer a missing account number from a plausible name. A model may suggest a beneficiary match, but identity uncertainty and payment irreversibility require strong controls. A mandatory screening match must go through its governed disposition. STP is not improved responsibly by suppressing holds or hiding manual work in another system.

Set an abstention path. If source data is missing, model confidence is low or reference version is stale, send the item to review or use the approved conventional process. Monitor abstention rate and queue capacity. A system that reports high STP only by excluding difficult cases from its denominator is not improving the bank's total workflow.

Human review design

Show the analyst the original message, proposed edit, source evidence, model confidence or uncertainty and relevant rule. Capture accept, modify or reject with reason. A reviewer should be able to identify an unsupported model suggestion, not just approve a polished draft. Audit the actual message submitted after review.

Feedback from edits can improve future models, but an analyst's correction is not always a universal label. It may reflect a changed correspondent rule or new customer information. Categorize source defect, policy exception, model error and customer correction. Route each to the appropriate owner before retraining.

Measure analyst time, rework and error. A tool that adds an extra confirmation step may reduce STP even if predictions are good. A tool that speeds simple cases while leaving complex cases to specialists may improve overall service. Segment workload and quality to understand the effect.

Payment example

A cross-border payment message lacks a standardized purpose code and includes a free-text description. A model suggests a code with a source-linked rationale; a rules service checks whether the code is permitted for the corridor. High-confidence suggestions under an approved policy may be accepted for a narrow class of messages, while ambiguous cases go to a specialist. Screening and other mandatory checks remain independent.

Track the instruction from intake to final settlement or return. If the model suggestion is accepted and the payment later returns for wrong purpose, the apparent first-pass STP gain is offset by rework. Measure net clean completion after a defined follow-up window. Record original text, model version, rule version, final code and any human action under access controls.

If the purpose-code reference becomes stale, the model may keep suggesting a formerly valid code. A freshness check should disable the automatic path and route to review. Correcting the reference later should not rewrite earlier payment records. Reverse lineage identifies messages that consumed the obsolete mapping and their actual outcomes.

Document-processing example

A trade-finance workflow receives invoices and supporting documents. OCR and extraction models identify invoice number, amount, dates and counterparties. A deterministic check compares fields with the payment instruction and policy. If the model is uncertain or documents conflict, a trained reviewer resolves the case. The bank retains source spans and the final approved data, not just a generated summary.

Evaluation should include poor scans, multi-page tables, multiple currencies, handwritten marks and unusual templates. A 98 percent field-level accuracy can still produce many documents with at least one critical error. Measure complete-case correctness, manual review, downstream return and customer time. An automatic path should be restricted to cases supported by validation evidence.

Monitoring in production

Track eligible volume, AI coverage, auto-accepted suggestions, human edits, abstentions, manual repairs, returns and final STP by product and corridor. Monitor feature and reference freshness, model confidence distribution and exception queues. A rising auto-acceptance rate may reflect improved model utility or weakened review discipline. Sample accepted cases against source evidence.

Watch for shifts after payment-standard, correspondent or product changes. A code mapping can remain syntactically valid but become semantically obsolete. Update reference data and revalidate the model-policy combination. A generative assistant's citation quality should be monitored separately from payment completion.

Report financial and customer outcomes with caution. Time saved per case can be measured, but a claimed loss reduction requires a credible comparison and mature outcomes. A higher STP rate can coincide with more returns or wrong-party payments. Management should see the full set of guardrails, not one headline rate.

Incident and fallback

Suppose a beneficiary-reference feed fails while the model API continues to respond. The model's suggestion may be based on stale mapping. The system should flag invalid data and use the approved manual or rule-only route. Identify every instruction scored with the stale reference, preserve actual decisions and assess returns or customer impact. Retraining the model is not the immediate remedy.

A model service timeout should not cause the payment hub to skip validation. The existing controlled process continues within its capacity; if queues exceed limits, operations escalates and communicates accurate statuses. When the model returns, held items are reconciled before processing. A late suggestion must not overwrite an analyst's completed repair.

Improvement experiment

Choose a narrow exception type with sufficient volume and consistent labels. Run a shadow model and compare its suggestions to analysts' actions without changing payments. Review disagreements and source evidence. Then pilot a bounded assisted workflow, with predeclared success criteria for clean completion, review time, return rate, adverse incidents and customer treatment. Include a comparable baseline period or group and describe differences in mix.

At the end, compute STP using the original eligible population and agreed lifecycle. Report model coverage and exclusions, downstream returns and delayed corrections. Inspect rare high-impact errors. Expand only when the model, policy and human process together demonstrate a better outcome. The objective is a payment completed correctly with less unnecessary manual work, not a higher metric created by moving exceptions out of sight.

The denominator audit

Suppose 100,000 payment instructions enter a channel. The hub accepts 98,000; 2,000 fail early validation. Of the accepted instructions, 90,000 complete without manual intervention, 5,000 enter repair, 2,000 receive required screening review and 1,000 remain pending at the reporting cutoff. If a team reports 90,000 divided by 98,000, it has excluded early failures. If it reports 90,000 divided by 95,000, it may have excluded screening and pending cases. Neither figure is automatically wrong, but each answers a different question. Define and publish the intended denominator and exclusions.

Now assume an AI classifier resolves 1,000 of the 5,000 repair cases automatically, but 100 of those later return for incorrect details. A first-pass metric might show an improvement of one percentage point, while clean end-to-end completion improves less. Count returns, rework and customer delay within a defined follow-up period. If the model only shifts work from one queue to another, label that operational effect honestly.

Compare like-for-like traffic. A period with more domestic low-complexity payments will have higher STP even if the model does nothing. Report rates by corridor, channel, message type, source and value band, plus total counts. A standardized mix-adjusted estimate can complement the raw result, but disclose weighting assumptions. Never hide a critical low-volume corridor behind a high overall average.

Error taxonomy

Build a taxonomy of why a payment failed STP: missing required field, invalid format, ambiguous beneficiary, reference lookup failure, mandatory screening hold, customer cancellation, liquidity or limit issue, external network rejection and internal technical fault. Separate preventable data defects from legitimate control interventions. The AI use case should target a category it can influence safely.

Within model-handled cases, record extraction error, classification error, stale reference, low confidence, unsupported recommendation, reviewer correction and downstream return. A single "AI failed" bucket is too coarse for improvement. A document extraction error may need training or OCR work; a stale routing reference needs data governance; a wrong auto-acceptance threshold needs policy change.

Some exceptions occur after apparent completion. A message can pass initial validation and later be returned by a correspondent. A screening case can reopen. Define the follow-up horizon and update the outcome report when late events arrive. Retain the initial decision and subsequent lifecycle rather than overwriting one status field.

Model validation for auto-repair

Validate at the field and whole-message levels. If a model extracts five required fields with 98 percent accuracy each, the probability of an entirely correct message may be substantially lower when errors compound. More importantly, errors are not independent and some fields have greater consequence. Report critical-field accuracy, complete-case accuracy and severity-weighted error examples. Inspect rare formats and languages.

For a proposed code or route, test against current authoritative reference, eligibility and policy. A historical successful case is not proof that the same route works now. The model can provide a ranked suggestion, but a deterministic validation layer should reject impossible or prohibited outputs. Test abstention with unseen products and stale reference data. A model that always returns a confident answer can be operationally worse than one that refers uncertain cases.

Evaluate human review. Does the interface show original evidence and changed fields clearly? Can the analyst reject a suggestion without extra work? Are reviewers overly influenced by a fluent explanation? Sample accepted and rejected suggestions and compare to source records. Measure time and errors, not just click-through acceptance.

Feature and label leakage

An exception model trained to predict manual repair should not use a field populated only after repair. A code "repair complete" or final routing result is an outcome, not intake evidence. Freeze an observation cutoff at the moment the model would run. Point-in-time joins to product and correspondent reference data must reflect versions available then. Test historical rows against original incoming messages.

The label "manual touch" can reflect old staffing practices as well as message quality. One team may manually review all payments from a corridor even when data is valid; another may auto-process them. A model trained on that label may reproduce policy history rather than predict correctable defects. Include reason codes and domain review. A future policy change can invalidate the label's meaning.

Returned payments are useful outcomes but have delayed and selective observation. A payment blocked before submission cannot be returned by the network. A model that excludes blocked cases from evaluation can appear safer. Track the full eligible population and distinguish unobservable outcomes. Do not label every nonreturned payment correct if it has not yet matured.

Security and privacy

Payment messages and supporting documents contain account, identity and commercial information. Limit model inputs to the intended purpose, restrict training extracts and protect logs. A generative assistant should not send full payment narratives to an unapproved external service. Retain source references for authorized audit without making every engineer a reader of raw customer data.

A malicious free-text field can include instructions aimed at an LLM summarizer. Treat message content as untrusted data, constrain tools and require review. A model suggestion that changes beneficiary details carries higher risk than a summarization draft. Design permissions and approval around the action, not the marketing name of the model.

Vendor models may change behavior or availability. Version responses and test against a controlled sample after updates. Define what happens when extraction or scoring times out. The payment hub must still enforce validation and mandatory controls. A vendor outage is not a reason to mark all instructions straight through.

Financial evaluation

Estimate saved manual minutes from observed case handling, not a generic average. Multiply by the volume of cases actually avoided, then account for additional review, model operation, monitoring and remediation. A suggestion that analysts still verify may save time without increasing STP; report it as productivity. A model that auto-resolves cases may improve STP but incur returns; include their cost and customer effect.

Compare with simpler fixes. If most exceptions stem from one upstream form missing a required field, improving the form or deterministic validation may deliver a larger, safer gain. AI is justified when ambiguity, variety or prediction adds value that rules alone cannot provide under the same constraints. The baseline should include those feasible process improvements.

Do not claim revenue or loss benefits solely from STP percentage. Payment volume, fees, customer retention and operational costs may move differently. Use measured causal evidence where possible and state assumptions. Management should see both gains and guardrails.

Root-cause loop

When the model repeatedly flags a source channel's missing purpose code, send the pattern to the channel owner. Fixing intake can eliminate exceptions for all payments, including those outside model coverage. Track the reduction in source defects separately from model auto-repair. A good AI system can reveal process problems without becoming a permanent patch for them.

Analyst corrections should be categorized and linked to original messages. A repeated model misclassification may call for better training data; a new correspondent rule calls for reference and policy update; a privacy issue calls for data-flow repair. Do not automatically retrain on every accepted suggestion as if acceptance proved correctness. Later returns and quality sampling matter.

Operations drill

Simulate a peak hour with twice the usual exception volume, a stale reference table and a model timeout. Count instructions by original eligible population, AI-assisted cases, manual queue, mandatory screening, final settlement and returns. Verify that the queue can handle abstentions and that customer statuses match actual payment state. After recovery, reconcile each instruction and identify decisions made with stale inputs.

Repeat after a schema or payment-standard change. An old model can keep producing plausible codes that are no longer valid. The release test should include current reference data, new message examples and a rollback to the conventional controlled process. If the bank cannot operate safely without the model for a defined period, the STP improvement has created a fragile dependency rather than a reliable capability.

A worked release scorecard

For a four-week pilot, report eligible payment instructions by corridor and channel, then count those the model evaluated, abstained on, suggested for repair, automatically repaired and sent to analysts. Reconcile these categories to final payment states. Calculate first-pass STP, clean completion after a follow-up window, median and tail processing time, return rate, wrong-party or wrong-field incidents, analyst minutes and customer contacts. Keep the pre-pilot baseline under the same definitions.

For each automatic repair, retain original field, proposed field, confidence, source evidence, reference version, deterministic validation result and final message. Sample routine cases and all severe errors. An aggregate 99 percent acceptance rate can hide a small but unacceptable number of critical beneficiary mistakes. Report error severity and reversibility. An incorrect internal code caught before submission differs from a payment sent to the wrong destination.

Segment the scorecard. A model may work well on common domestic messages but poorly on a low-volume cross-border corridor. If the pilot excludes that corridor, say so. If analysts continue reviewing all suggestions, the result may be time saved but no increase in strict STP. If the model automatically repairs a narrow class, report that class's coverage and the total-bank effect separately. This prevents a local metric from being presented as a universal improvement.

Review changes in traffic mix and policy. A new channel form could reduce missing fields during the pilot independently of the model. A correspondent update could increase returns. Use a comparable period or controlled group where feasible and disclose remaining differences. The business decision should be based on actual net outcomes and known uncertainty.

Approval boundaries

The payment owner approves which exception types can use AI suggestions and which can be automatic. Model validation challenges training labels, field accuracy, edge cases and drift. Compliance owns mandatory screening boundaries. Data owners govern source and reference versions. Operations owns queue capacity and fallback. An audit function should be able to follow a sampled payment from original message to final status without relying on the model team.

If the system expands from one code to another, repeat the risk and accuracy assessment. A suggestion for a descriptive purpose category is not equivalent to changing a beneficiary account. A new corridor may have different standards and return behavior. Scope restrictions should be enforced in policy and monitored, not merely written in a launch document.

Set a review date and rollback trigger. A rise in returns, critical field errors, stale reference use or customer complaints may require disabling the automatic path. The conventional validation and repair process should remain workable at expected incident volume. An AI-enabled STP program succeeds when it improves correct completion and can prove that improvement under ordinary and degraded conditions. Preserve the pilot's source and model versions so later reviewers can reproduce the claimed result. Reconcile pending payments at the follow-up cutoff before reporting final clean completion. A payment in a manual queue is neither a clean completion nor a confirmed return. Report its age, owner and accurate customer-facing status.

Clean completion after automated repair

An AI extractor proposes a purpose code for 500 cross-border payment messages. The bank's deterministic checks accept 400 suggestions; 100 go to analysts. Of the 400, 20 later return because a correspondent rejects the code. A first-pass automation metric of 400 is not the same as 380 clean completions. Count eligible instructions, manual touches, returns and pending cases over a declared follow-up window.

Sample the 20 returns for source evidence, model confidence, reference version and approval rule. If the code list was stale, correct reference governance and review affected messages; retraining the extractor alone will not fix it. Compare the net outcome with a simpler channel validation improvement. The value of AI is correct end-to-end completion with less unnecessary work, not a high auto-acceptance rate. Measure the 100 analyst cases' handling time and error rate, then compare the pilot with a simple form validation fix that prevents missing purpose codes upstream. Keep required screening and eligibility checks in both paths. A returned payment counts as rework even when the model's suggestion was initially accepted.

Related learning paths

This application uses JavaScript for the full interactive experience. This text summary is served for accessibility and search indexing.

Improving STP rates using AI · Malla Banking Academy