Validation, reconciliation, and exception handling

Validation, reconciliation, and exception handling. A practical lesson in the data foundation for banking and payments practitioners.

How to study this topic

Validation, reconciliation, and exception handling are the banking controls that decide whether data is fit to move forward into reporting, analytics, AI, and operational decisions. Study this chapter as a banking control topic first and a data science topic second. The model is not the beginning of the story. The beginning is the bank's duty to know what happened, where the data came from, whether it is complete, whether it is reconciled, whether it is allowed to be used, and whether the result can be explained later.

This topic belongs to Data Foundation, so the focus is broad banking. It covers lending, deposits, cards, treasury, finance, regulatory reporting, risk, compliance, customer service, branch operations, digital channels and internal operations. Payments are included only where the title naturally needs transaction examples. The point is not to turn the lesson into payment processing. The point is to explain how AI adoption works across a bank.

The banking meaning

validation, reconciliation, and exception handling is important because banking data carries legal, financial, customer and regulatory meaning. A balance is not just a number. A balance may be available, booked, ledger, cleared, unsettled, blocked, earmarked, overdrawn or adjusted. A customer status is not just a label. It may affect onboarding, credit treatment, servicing, conduct obligations, complaints, vulnerable customer handling and regulatory reporting.

Sources and boundaries

Typical sources include customer master records, account records, loan servicing feeds, deposit balances, card transactions, treasury positions, finance ledger extracts, risk exposure files, fraud case outcomes, compliance alerts, branch operations, and digital channel events. These sources do not have equal authority. A digital channel may show customer intent, but a book-of-record platform may show final outcome. A CRM note may show relationship context, but it may not be structured enough for automated model use. A risk system may hold a rating, but the rating may have a valid-from date, expiry date, override reason or review cycle.

Why AI depends on this discipline

The practical AI use cases include risk reporting, credit model training, fraud monitoring, customer segmentation, finance analytics, operational dashboards, regulatory evidence, and model monitoring. These can help a bank act faster and with more consistency, but only if the model input has banking quality. A model trained on weak data can still produce a neat score, rank, category or recommendation. The output may look professional while the underlying evidence is damaged.

AI does not remove the need for banking controls. It makes those controls more important because one bad input can influence many downstream decisions at speed. The safest banks treat model input preparation as part of governance, not as background plumbing.

Validation controls

Validation asks whether the data is acceptable for its intended use. For this topic, useful controls include schema validation, mandatory field checks, referential integrity, business-date alignment, and control totals. These controls should run before data becomes a model feature, dashboard metric, risk indicator or operational priority. The controls must check both technical shape and banking meaning.

Validation should produce actionable exceptions. It should say what failed, why it matters, which downstream consumers are affected, whether the model must stop, whether degraded use is allowed, and who owns correction. A generic red status is not enough for production banking.

Operational design

Operational resilience also means fallback. If the feature feed is unavailable, should the bank stop the model, use the last good value, route to manual review, switch to a rule-based control, or degrade the service? The answer must be designed before failure.

Governance and audit

BCBS 239, model-risk guidance and AI-risk frameworks all point in the same practical direction: banks need controlled data, clear limitations, documented governance and monitoring. The details differ by jurisdiction and model type, but the core discipline is stable.

Common mistakes

The final mistake is rushing feature reuse. Centralised features are powerful, but reuse without purpose control can create hidden risk. The same feature may be safe for portfolio monitoring and unsafe for direct customer decisioning. A bank must know the difference.

Bank-ready checklist

If these questions are answered well, AI adoption becomes much safer. The bank is not merely feeding data into a model. It is turning banking evidence into controlled decision support.

Source anchors for further study

BCBS 239 supports the banking discipline of accurate, complete, timely and adaptable risk data aggregation and reporting.

Model-risk guidance such as the Federal Reserve supervisory guidance emphasises input quality, data constraints, limitations, validation, monitoring, documentation and governance.

NIST AI RMF is useful as a general AI risk framing for governance, mapping, measuring and managing AI risk, but banking still needs bank-specific controls around customers, products, books of record, risk and regulatory evidence.

Practical banking note on business meaning

For validation, reconciliation, and exception handling, the business meaning point matters because banking AI does not live in isolation. It sits beside customer journeys, product ledgers, risk controls, finance numbers, operations queues, compliance obligations, audit evidence and management decisions. A weak data point can move quietly through all of these areas unless the bank designs controls that stop, flag or limit it.

The practical discipline is to keep four layers separate: observed banking fact, derived data signal, model interpretation and business action. Observed facts come from systems such as customer master records, account records, loan servicing feeds, deposit balances, and card transactions. Derived signals reshape those facts into features or indicators. Model interpretation produces a score, classification, ranking or recommendation. Business action decides what the bank will actually do. When these layers are blurred, nobody can explain the result properly.

Practical banking note on lineage

For validation, reconciliation, and exception handling, the lineage point matters because banking AI does not live in isolation. It sits beside customer journeys, product ledgers, risk controls, finance numbers, operations queues, compliance obligations, audit evidence and management decisions. A weak data point can move quietly through all of these areas unless the bank designs controls that stop, flag or limit it.

Practical banking note on customer impact

For validation, reconciliation, and exception handling, the customer impact point matters because banking AI does not live in isolation. It sits beside customer journeys, product ledgers, risk controls, finance numbers, operations queues, compliance obligations, audit evidence and management decisions. A weak data point can move quietly through all of these areas unless the bank designs controls that stop, flag or limit it.

Practical banking note on risk control

For validation, reconciliation, and exception handling, the risk control point matters because banking AI does not live in isolation. It sits beside customer journeys, product ledgers, risk controls, finance numbers, operations queues, compliance obligations, audit evidence and management decisions. A weak data point can move quietly through all of these areas unless the bank designs controls that stop, flag or limit it.

Practical banking note on model limitation

For validation, reconciliation, and exception handling, the model limitation point matters because banking AI does not live in isolation. It sits beside customer journeys, product ledgers, risk controls, finance numbers, operations queues, compliance obligations, audit evidence and management decisions. A weak data point can move quietly through all of these areas unless the bank designs controls that stop, flag or limit it.

Practical banking note on operational fallback

For validation, reconciliation, and exception handling, the operational fallback point matters because banking AI does not live in isolation. It sits beside customer journeys, product ledgers, risk controls, finance numbers, operations queues, compliance obligations, audit evidence and management decisions. A weak data point can move quietly through all of these areas unless the bank designs controls that stop, flag or limit it.

Practical banking note on audit evidence

For validation, reconciliation, and exception handling, the audit evidence point matters because banking AI does not live in isolation. It sits beside customer journeys, product ledgers, risk controls, finance numbers, operations queues, compliance obligations, audit evidence and management decisions. A weak data point can move quietly through all of these areas unless the bank designs controls that stop, flag or limit it.

Practical banking note on privacy and permitted use

For validation, reconciliation, and exception handling, the privacy and permitted use point matters because banking AI does not live in isolation. It sits beside customer journeys, product ledgers, risk controls, finance numbers, operations queues, compliance obligations, audit evidence and management decisions. A weak data point can move quietly through all of these areas unless the bank designs controls that stop, flag or limit it.

Practical banking note on feature reuse

For validation, reconciliation, and exception handling, the feature reuse point matters because banking AI does not live in isolation. It sits beside customer journeys, product ledgers, risk controls, finance numbers, operations queues, compliance obligations, audit evidence and management decisions. A weak data point can move quietly through all of these areas unless the bank designs controls that stop, flag or limit it.

Practical banking note on monitoring

For validation, reconciliation, and exception handling, the monitoring point matters because banking AI does not live in isolation. It sits beside customer journeys, product ledgers, risk controls, finance numbers, operations queues, compliance obligations, audit evidence and management decisions. A weak data point can move quietly through all of these areas unless the bank designs controls that stop, flag or limit it.

Three controls answer different questions

Validation asks whether a record meets a defined rule: required fields, code sets, format, relationship and business constraints. Reconciliation asks whether records and amounts agree across an expected boundary: source to landing zone, channel to hub, hub to external message, posting to ledger, or offline to online feature. Exception handling decides who investigates a mismatch, what action is permitted, and when a model or payment may proceed. A green schema check cannot prove that an entire product partition arrived. A balanced ledger cannot prove that a customer feature was available at the earlier model decision. Keep the control question and owner distinct.

In an AI pipeline, the relevant boundaries include source extraction, identity joins, feature calculation, model input, policy action and later outcome labels. For each, state expected counts, amounts where meaningful, accepted and rejected rows, reconciliation keys, tolerance, timing and escalation. A high-volume feed may use automated checks and sampled evidence; material customer decisions still need a path to reconstruct an individual case. Reconciliation should not hide a missing record because a duplicate elsewhere offsets the total. Compare by product, date, currency, source and business key as appropriate.

Payment example: one instruction, several records

A customer instructs a EUR 12,000 cross-border payment. The channel records the order, the hub records validation and screening outcomes, a pacs.008 is released with its identifiers, a correspondent status is received, and the bank posts a customer debit and later receives account reporting. The data product used for repair prediction must link these facts without treating a network acknowledgement as a customer credit. A return, if one occurs, is a separate financial movement linked to the original payment. Reconcile the number of channel orders accepted to hub orders, outbound messages and expected accounting entries, with legitimate holds, rejects and cancellations explained.

Suppose the hub shows two release events but the gateway shows one accepted message because the hub republished an event after a restart. An event consumer should identify the duplicate publication and avoid a second feature count or customer notification. If two genuinely separate messages were sent, operations needs an urgent duplicate-payment investigation. The reconciliation cannot infer which from amount alone; it needs internal payment key, message ID, UETR, payload version, send time and gateway response. The model training label duplicate must distinguish these cases. An exception owner records the determination and financial outcome.

Credit example: source totals and feature values

A nightly account extract says 50,000 active loans, while the warehouse receives 49,800. The difference may be a source filter, failed partition, migration or account closure. Before publishing a repayment feature, compare source control totals and sample the missing population. A bank should not assume the 200 missing accounts are random; they may belong to a new product or high-risk segment. If the affected feature is unavailable, an online credit model uses its approved missing-data or referral path, not a fabricated zero arrears value. Record the exact batch, accounts and decisions affected.

At the individual level, a repayment-history feature might count a reversal incorrectly as a missed instalment. The account ledger, schedule, payment posting and reversal need a business rule for that product. Reconcile the computed feature for sampled accounts to source evidence and an independent calculation. If the feature store and training warehouse use different rules, the model may pass offline validation while making wrong live decisions. A correction after an adverse decision should trigger a defined review and possible customer remedy under the bank's process; it must not overwrite the first decision snapshot.

Online and offline reconciliation

The model owner should choose a fixed set of decision timestamps and compare each online feature value with a reconstructed historical value under the same definition. Include normal cases, late events, changed customer identities, missing feeds and version transitions. A difference can arise because the online service used information not stored for replay, or because the offline job accidentally included future data. Investigate which path reflects the approved contract before changing either. A numeric tolerance may be appropriate for rounding or currency conversion, but it must be documented and tested so it cannot hide a materially different borrower or payment.

Record feature version, source event IDs, calculation time, availability cutoff, score and policy action for the online decision. An offline job should report which inputs it reconstructed and any missing evidence. If the bank cannot reproduce a material score, that is an exception with an owner and a release consequence. A model metric computed from a final cleaned warehouse does not settle a dispute about what the live service knew at decision time. The audit trail should allow a validator to move from model input back to each authoritative source.

Exception states and decision rights

An exception queue should distinguish detected, assigned, under investigation, awaiting source, resolved, waived under authority and closed with evidence. A waiver is not the same as a fix. Record who approved it, why the affected model use can continue, scope, expiry and monitoring. A temporary waiver for a monthly portfolio forecast may be unacceptable for an automated loan decline. Set materiality by customer or financial consequence as well as record count. A single wrong identity join can be more serious than thousands of low-impact formatting rejects.

When several controls fail, avoid closing a case just because one issue is resolved. A payment with a corrected account number may still be under sanctions review. A loan with a repaired income feed may still lack current obligations. The system should display separate control states and block an action until the applicable policy gates pass. If an exception is re-opened after a later correction, keep the first disposition and the reason for reopening. This supports fair customer communication and incident analysis.

Timeliness and operational service levels

The deadline for reconciliation depends on the decision. A pre-release fraud score may require a feature within seconds; an end-of-day financial close may allow a later statement; a monthly model-monitoring report can wait for matured outcomes. Define when the clock starts, the expected evidence, the escalation time and the fallback. A feed arriving late with an earlier business date was still unavailable at the initial decision. Operations should see the last approved state and its time, not a dashboard that silently changes to the late value.

If a reconciliation job fails at a cut-off, the system should identify affected payments or applications and route them under a documented policy. Retrying the job must be idempotent. A late correction may require re-evaluation, customer notification, accounting adjustment or no action after materiality review; the appropriate owner decides. The data team should not infer the business remedy from a successful rerun. Keep source and publication timestamps so the bank can explain both the first failure and later resolution.

Tests that prove the controls work

Build a matrix with a valid case, missing partition, duplicate event, wrong customer join, amount mismatch, late status, reversal, changed feature definition and model-service timeout. For each, state the validation result, reconciliation result, exception owner, permitted action and evidence. Test a pair of equal-amount payments to the same beneficiary so the matching process cannot rely on amount alone. Test a file whose total count is correct because one partition is duplicated and another missing. Test a corrected source that changes a feature after a decision. The expected behavior should be specific enough for a developer and operations analyst to agree.

Review a sample end to end. Recover raw input, transformation, quality check, exception case, feature snapshot, model score, human or policy action, customer status and later outcome. If any link is missing, the bank should document the limitation and decide whether the model use can proceed. The control report should show not only the number of breaks but aged cases, waivers, repeated defects, customer effects and whether a previous incident recurred. A falling exception count can mean a better feed or a disabled rule; independent sampling helps tell the difference.

Release evidence and accountability

The source owner certifies what was extracted or published. The data platform owns transport and transformation. The feature owner defines the model-ready value and replay. Finance reconciles postings and balances. Operations resolves payment or customer cases. Model risk challenges the model's use of the data. The product owner decides the customer action and fallback. An audit sample should show all these responsibilities without inventing a single universal data owner. The release pack contains control definitions, source totals, exception logs, feature comparison results and sign-offs for material waivers.

The Basel Committee's BCBS 239 principles are a primary reference for risk data aggregation and reporting at banks within scope. They do not define every payment or feature reconciliation rule. The bank applies relevant principles to its architecture, model use and local obligations. A controlled AI decision is one whose input quality, cross-system consistency and exception outcome can be explained at the time it was made, including what remained unknown.

A worked exception from detection to closure

At 08:00 a batch of transaction records arrives for a behavioral credit model. The source control report expects 120,000 records and a defined sum of posted amounts by currency. The landing zone has 120,000 records, but one product has 500 fewer and another has 500 duplicates. A total-count check alone passes. Product-level counts, unique source IDs and currency totals fail. The platform quarantines the batch before feature publication and opens an exception with the affected products, candidate accounts, source file and failed controls. The model service continues only under a previously approved fallback for unaffected populations; affected lending decisions are referred rather than scored from partial history.

The source owner discovers that a restart resent one partition while omitting another. A corrected extract arrives at 09:30. The platform validates its sequence and idempotency, reconciles counts and amounts by product, and publishes a new approved version. The feature owner compares calculated repayment histories for affected accounts with the last approved snapshot. The model owner identifies any decisions attempted between 08:00 and 09:30 and verifies they followed the referral policy. Operations contacts customers only if a delayed or wrong action warrants it under the product process. The first failed file and exception remain in the evidence trail even after the corrected file passes.

An independent reviewer samples an affected account, an unaffected account and a duplicated record. They can reconstruct the raw source lines, parsing result, quarantine reason, corrected batch, feature values and final customer action. The review asks why the product-level reconciliation caught the issue and whether a source restart should have triggered a sequence alarm sooner. The team adds a regression case for that failure mode. This is a complete control story; a spreadsheet cell marked resolved without the affected decision trail would not be.

When a variance is legitimate

Not every difference is a break. The customer instruction amount can differ from the booked amount because of FX, while a statement can aggregate entries differently from the hub. A source can report a correction after the bank's first snapshot. Define approved matching rules and tolerances with the underlying business reason, effective date and evidence. A permitted variance should be visible as such, not hidden by widening a general tolerance until every case matches. Test the tolerance around its boundary and include a case where two unrelated entries happen to net to zero. A model trained on reconciled data needs to know whether a match was exact, allowed with variance, or resolved manually.

The last check is whether exception labels themselves are reliable. If one team closes a case as source corrected and another uses no issue found for the same condition, a model learning repair causes receives inconsistent targets. Maintain a controlled taxonomy and train reviewers. Sample dispositions against source evidence and customer outcome. Changes to the taxonomy need versioning so historical labels can be interpreted. The bank should not use a high model accuracy score to mask weak operational truth at the end of the data chain.

Three distinct controls

Validation asks whether an input obeys its field and relationship contracts. Reconciliation asks whether the complete expected population survived across systems. Exception handling determines who owns a failed item and how its banking action proceeds. An AI pipeline can pass schema validation while losing a file partition, and it can reconcile counts while giving every transaction the wrong currency scale. Operate these controls together, but keep their evidence distinct.

Consider a payment hub publishing 10,000 accepted instructions. The intake service receives 10,020 messages because 20 are retries. Deduplicate by stable instruction ID and event version, then reconcile 10,000 eligible instructions with the hub. A raw message count of 10,020 should not be forced to match 10,000 by dropping arbitrary rows. Reconcile by channel, processing window, status and amount as well as total count. Later returned instructions remain linked to originals and should be explained in a status bridge rather than disappearing from the original population.

Field and relationship validation

Validate required IDs, timestamp order, amounts, currency, account state and source code lists before feature construction. A parseable date in the wrong timezone can move a payment into a different velocity window. A positive amount with a negative refund status can be legitimate, but needs an explicit event-type rule. Check one-to-one and many-to-one joins against the business relationship at the cutoff. A customer may hold multiple accounts; an account should not unexpectedly resolve to two active primary customers in a model designed for one.

For training data, validate the sequence of observation and outcome. Features must be available before the application or payment decision, while labels mature after it. A case status updated after the decision belongs in the outcome lineage, never in the historic feature vector. Check that splits keep connected accounts or customers from leaking between training and evaluation where the model's proposed deployment requires entity independence. Record population exclusions with reason codes and balances so a clean dataset does not conceal who was omitted.

Reconciliation at each boundary

Set control points at source publication, landing, normalization, feature computation, score request, model response and final business action. Not every count will match exactly: one instruction can emit multiple lifecycle events, and one account can generate several feature rows over time. Document expected transformations and their keys. Reconcile source business IDs to served decision IDs, then to holds, releases and timeouts. A model call returning successfully is not proof that the hub acted on its response.

For a nightly credit score, freeze 50,000 eligible facilities and 2 billion units of exposure. If the feature join produces 49,950 rows, list the 50 missing IDs and their exposure; do not accept a rounded 99.9% success metric. If the score service returns 49,945 rows, distinguish five technical failures from the 50 join failures. The business owner determines whether to rerun, use an approved previous score or suspend downstream use. Preserve the original run ID and a separate corrected run ID.

Balance checks can find an error count checks miss. A truncated extract may retain most accounts while dropping a few large facilities. Conversely, a duplicate join can preserve distinct facility count while doubling the exposure summed across rows. Compare both unique IDs and financial totals to source controls; sample the largest discrepancies and one ordinary record. A reconciliation should have a named cutoff and source revision so two valid reports taken at different moments are not mistaken for a defect.

Exception states and ownership

Give each exception a state such as detected, assigned, contained, corrected, replayed, reviewed and closed. The state describes the data defect; the payment or credit decision has its own independent final status. A payment with an invalid feature must still be held, rejected or processed under the bank's approved fallback by its deadline. It cannot wait silently in a technical dead-letter queue while customer funds or obligations remain unresolved.

Route an unmatched beneficiary reference to the reference-data owner, an invalid currency scale to the source interface owner and a model timeout to service operations. The model risk owner assesses whether prior decisions are affected. A human reviewer needs the source evidence and allowed actions, not a generic "data error" notification. Record case ID, affected business IDs, earliest event, severity, owner, deadline and closure evidence. Escalate when queue age or impact exceeds a limit.

Do not equate replay with resolution. Reprocessing can update analytic history and feature state, but issuing a second payment release would be harmful. Use idempotent business commands and reconcile final ledger and hub states after repair. Compare original score inputs to corrected values as an impact analysis, and keep the original decision record immutable. Determine whether customer remediation is warranted from actual outcomes, not a score difference alone.

Test a broken nightly run

Inject a missing branch file, a duplicate file, one currency-unit switch and an ambiguous customer join. The run should identify the missing population through control totals, refuse to double-count the duplicate, fail the value check and quarantine the ambiguous relationship. Publish a diagnostic manifest but no normal model-ready dataset. After owners correct sources, create a versioned rerun and compare counts, balances and affected feature values. The approval record should name which downstream model runs were blocked or restarted.

Measure exception detection time, unresolved age, affected decisions, successful repairs and repeated source causes. A falling exception count is not proof of improvement if the checks stopped running or a new integration bypasses them. Periodically seed known faulty records and confirm detection and routing. The full control succeeds when a reviewer can account for every eligible item and explain both the technical repair and the final banking action.

A reconciliation break becomes a model-input decision

Consider a fictional outbound payment that the channel accepted at 09:01. The hub assigned instruction P-721, released it at 09:03 and sent a network message. The network acknowledgement arrived at 09:04, but the customer debit was absent from the account ledger at the expected booking checkpoint. A model that predicts payment completion should not turn these observations into a single "successful" flag. Channel acceptance, network acknowledgement, settlement evidence and booking are separate facts, each with its own identifier, timestamp and owner.

Validation first asks whether each event is structurally usable: required key, recognised status, valid currency and version. Reconciliation asks whether the expected events join to the same instruction and whether their amount and state agree at the relevant boundary. Exception handling decides what to do when a valid event has no match. The missing debit is not a parser defect; it may be timing, posting failure or a genuinely different processing path. Assign an exception ID and keep the unresolved state visible to the model consumer.

At 09:07 a duplicate network acknowledgement arrives with a new transport ID but the same business reference. The event ledger should retain the delivery for investigation while the reconciled instruction has one acknowledgement. At 09:10 the account debit appears with a posting time of 09:08. An as-of model score at 09:05 must still show booking unknown, even though a current dashboard can now show the debit. The exception closes only after the operations owner checks the expected posting, amount, account and any fee or FX difference under its rules. A late event should update the case through a versioned transition, not erase its earlier state.

Build tests for five variants: no network acknowledgement by deadline, acknowledgement with a mismatched instruction key, duplicate acknowledgement, valid acknowledgement with delayed booking, and booked debit with a subsequent reversal. For each, assert the observed source events, reconciled state, exception reason, feature value at a named score time, fallback action and final case transition. If a reversal follows the debit, the model outcome label must reflect its defined target rather than treating the earlier posting as permanent completion.

The business owner must set the deadline for calling a record late by product and rail. An instant-payment operational decision has a different window from overnight account reporting. A single global timeout can create false breaks for one product and hide genuine delay in another. Monitor unresolved cases by age, value, channel and model exposure. Report which decisions were made while reconciliation was incomplete, so validation can assess whether the model learned from uncertain inputs or outcomes.

An analyst should also distinguish a missing event from a missing interface. If the ledger feed itself stopped at 08:50, hundreds of payments can appear unbooked for the same technical reason. Group breaks by feed heartbeat and source partition before opening hundreds of independent payment investigations. Once delivery resumes, reconcile the backlog to the original instruction IDs and maintain the interval during which model features lacked booking evidence. The batch closure should not hide the period of degraded decision quality.

Reconcile a cross-border message journey

A cross-border payment may be accepted at initiation, enriched, screened, routed, sent to a correspondent and later returned. The hub assigns one stable instruction ID while each stage emits separate technical messages. A validation rule checks field format and permitted purpose codes; reconciliation checks that each accepted instruction has an accounted-for current and final state. An exception case records who investigates a missing or conflicting stage. Do not compare raw message count with original instruction count as though they should match one-to-one.

Take 1,000 accepted instructions. The stream shows 1,005 accepted messages because five were redelivered after an acknowledgment failure. The enrichment service has 999 distinct IDs; one was quarantined for an unknown currency code. Screening has 998 completed results and one pending candidate, while the remaining item is the quarantined instruction. The hub has released 970, held 29 and rejected one under approved rules. A good bridge explains every original ID and why stage totals differ. If the enrichment team simply drops the unknown code, later totals may look clean while a real customer payment has no decision.

Handle a reference exception

For the quarantined instruction, preserve raw code, source interface version, payment amount, currency and deadline. The reference-data owner determines whether the code is new, mistyped or mapped differently under a partner release. An AI classifier may suggest a purpose or country code, but the deterministic eligibility check must validate it before use. An analyst sees the source evidence and can approve a repair, request clarification or reject the instruction under policy. The case ID links the suggestion, human action, corrected message and eventual correspondent response.

If the payment is repaired and resubmitted, the hub must prevent a second release of the original instruction. A return days later remains linked to that instruction and should not be counted as a new outgoing transfer. Reconcile ledger entries, correspondent acknowledgment and customer status before marking the exception closed. A model's proposed fix is only the beginning of the operational path.

Validate a portfolio extract differently

For a monthly loan model, the source control is 50,000 facilities and 2 billion units of exposure. A feature join that returns 50,000 rows can still duplicate one facility and omit another; compare unique IDs and balances, not only row count. A missing high-value facility may be more consequential than many small accounts. Identify exclusions by product and status, and compare prior run cohorts after accounting for genuine new bookings, closures and repayments. An account migration can change identifiers without changing the underlying loan.

When a batch fails, publish an exception manifest and block normal model-ready output. A corrected rerun has a new version, changed-record report and approval. Downstream reports should not mix the old and corrected run. If a previous score is used under an approved contingency, mark its age and population coverage. It cannot be silently represented as a fresh score from the failed batch.

Exercise the controls

Inject one duplicate message, a delayed partition, an ambiguous customer relationship, a source code change and a balance-scale error. Validation catches code and relationship defects; reconciliation finds missing IDs and financial mismatch; exception workflow assigns owners and deadlines. Some defects may require a live fallback before correction. Verify that the bank can enumerate affected decisions, preserve their original inputs, repair analytic history idempotently and account for final customer actions. Test closure by handing another team the run and case IDs and asking it to reproduce every discrepancy.

Related learning paths

This application uses JavaScript for the full interactive experience. This text summary is served for accessibility and search indexing.

Validation, reconciliation, and exception handling · Malla Banking Academy