Core banking data as the starting point. A practical lesson in the data foundation for banking and payments practitioners.
Begin with the bank's recorded objects
A banking model may estimate fraud, credit risk, customer demand or operational delay. Before it can make a useful estimate, the bank must identify the object under observation. Is it one customer, one legal entity, an account, a lending facility, a card, a payment instruction, a posting, or a case? Core banking records give many of these objects their controlled identities and states. A model trained on a vague join of "customer activity" can be wrong even when its code runs perfectly.
Consider a fictional bank, Meridian. A customer has a current account, a credit card and a joint savings account. The customer initiates a transfer on Monday and disputes a card charge on Wednesday. An analyst wants a feature called "recent account activity" for a fraud model. Which accounts belong to the customer at each decision time? Does the joint account count in the same way as an individually owned account? Is the disputed charge a posted transaction, a pending authorization or a later adjustment? The answers come from source records and a documented business definition, not from a data scientist's guess about table names.
Core banking is not one universal software product. A bank may hold customer and account masters, deposits, loans, cards, payments and general-ledger information in different applications. A modern cloud ledger may coexist with older product systems. The lesson is about the authoritative business record and its movement into analytics. The source for an account's contractual status may differ from the source for a payment's settlement outcome. The bank must name the authority for each field and time period.
A practical inventory starts with six questions. What object is represented? Which system can change its authoritative state? Which identifier persists across systems? Which timestamp says when the fact was effective? What correction or reversal can follow? Which decisions consume the fact? An answer such as "the warehouse has a customer table" is incomplete. A warehouse may be a curated copy; its row can be stale, deduplicated or transformed for a particular report.
The Basel Committee's 2026 BCBS 239 implementation newsletter discusses governance, lineage and the difficulty of data aggregation across fragmented systems. It is expressly informational and does not create new supervisory expectations. The original BCBS 239 principles focus on risk data aggregation and reporting, particularly for systemically important banks. They provide a useful discipline for this lesson, but they are not a blanket rule that every AI feature in every bank is subject to the same BCBS 239 requirement.
The customer, account and product are different keys
A customer master identifies a person or organization within a bank's operating model. An account master describes an account and its attributes. A product record describes the contract or offer, such as a deposit account, card or loan. Those are related objects, not three labels for one row. One customer may hold several products and accounts; one account may have several holders or authorized users. A corporate group may contain several legal entities. A joint owner may not share the same rights as a signatory. An account number displayed to a customer may be changed, masked or reused under a migration policy, so the bank may also need an internal stable key.
Build relationships as dated, typed links. For each customer-account association, record the role, effective period, source, status and any correction. When a holder is removed on Thursday, a Tuesday decision must still be replayed with Tuesday's relationship. A current-state join can make a historical transaction appear to belong to a person who did not have authority at the time. In a credit application, incorrectly aggregating a joint account into an individual's balance history can also misstate available funds.
Customer identity is a governed matching problem. Two records with similar names might describe one person, or two relatives. An exact identity number might be mistyped or belong to an obsolete system key. Matching rules should produce a confidence and an exception path rather than silently joining ambiguous records. Human-approved merges need provenance; a later split must preserve prior decisions and identify affected outputs. The next lesson addresses Know Your Customer and customer master data in depth. Here, the core point is that a model must inherit the approved entity relationship at the relevant time.
Product terms determine how a field should be interpreted. A negative current-account balance may be an unauthorized overdraft or a contractual facility. A positive credit-card balance may represent debt rather than available cash. A mortgage repayment record has a different observation cycle from a current account's daily movements. A "limit" might be a credit line, an authorization ceiling, a transaction limit or a temporary control. A bank cannot safely put all columns called limit into one generic feature without defining product, unit and effective date.
For Meridian, an analyst defines a customer-level monthly inflow metric. The signed-off definition includes eligible account types, holder roles, currency conversion policy, posting statuses, internal transfers, reversals and the observation window. Joint-account inflows are not automatically counted twice for both holders. If another product needs a different definition, it gets a separately named measure. Reuse means preserving meaning, not forcing every customer decision into one ambiguous field.
Balances describe a state at a specified time
A balance needs its basis. Ledger or booked balance reflects recorded postings under an accounting convention. Available balance may reflect holds, pending items, overdraft availability and product rules. A value balance can reflect a different effective date. A statement balance is tied to a reporting cycle. None is automatically the customer's spendable amount at every moment. A model feature named "balance" without a basis and timestamp invites a false comparison.
Suppose Meridian displays an available balance of 900 units at 09:00. At 09:02 an outgoing payment places a 400-unit hold. A nightly posting at 23:00 books that payment. A data extract at 09:05 may show a booked balance unchanged from the previous day but an available balance reduced. A credit model run at 09:03 cannot use the later booking as if it was then known. A fraud model at 09:03 may legitimately use the hold if the feature feed had received it. Both require the source event time and the time the downstream system actually received it.
Currency and unit must travel with the number. A balance of 1,000 could mean EUR, SEK or an internal minor-unit representation. A consolidated view may convert it using a dated rate; the converted amount is a derived value, not the account's booked currency balance. A banking model should preserve original amount and currency, conversion method, source rate and as-of time if it compares cross-currency balances. It should avoid calling a translated value an account balance without qualification.
Balance history can be represented as snapshots, posting-derived intervals or a source-provided time series. Each has different failure modes. A daily snapshot misses intraday lows. A posting reconstruction can miss holds and source-specific available-balance rules. A source time series can contain late corrections. The feature owner should choose the representation based on the decision and measure whether the data is timely enough. "Average balance over 30 days" needs a calendar, cutoff, treatment of missing days and policy for account opening or closure within the window.
A negative or zero balance is not automatically a delinquency. It may represent an approved overdraft, a newly opened account, an end-of-day sweep, or a feed default. The analyst must distinguish an observed zero from "no record." Treating missing as zero can systematically depress a balance feature for channels with slower refresh. Before modeling, inspect the distribution by product and source and sample individual histories back to authoritative records.
Postings are not payment intentions
A customer can initiate a payment that is rejected before execution. An authorization can be reserved without a final posting. A card purchase can post days after authorization. A transfer can post and later reverse. A settlement event can occur in a separate rail or account from the customer's booking. Core posting data tells the bank about accounting entries, but it should not be used as a substitute for every earlier business event.
Consider a customer who attempts the same transfer twice because the first screen times out. The first instruction is accepted for processing, the second is rejected as a duplicate, and only one debit posts. A model trained solely on postings sees one debit; an operations model analyzing failed attempts needs both instructions. Conversely, a model that counts both requests as successful outflows overstates account conduct. The source event chain needs correlation identifiers and explicit statuses.
A posting ledger should carry entry identity, account, amount, currency, debit or credit direction, booking date, value date where relevant, source reference, product and reversal relationship. Some systems use different sign conventions. A negative amount in one feed may mean debit, while another feed stores a positive amount plus a debit flag. The data contract must test those conventions with real source examples. A dataset that balances in aggregate can still mislabel the direction of every individual feature.
Corrections are events, not an excuse to rewrite history invisibly. Suppose a fee posts in error on Monday and is reversed on Tuesday. A decision made on Monday may have seen the fee; an investigation on Wednesday should see both the original and the reversal. A current net amount of zero conceals the customer's temporary balance effect and the correction process. Store both the original decision snapshot and the later corrected view, linked by identifiers.
Payment-event detail belongs to the adjacent transaction and payment-event lessons. This lesson establishes the boundary: core account postings are the bank's booked record for the customer account, while initiation, validation, clearing, settlement and exception states require their respective systems of record. An AI model may combine them, but it must not collapse them into a single "payment completed" flag.
Point-in-time joins require two clocks
A record can have an effective time and an arrival time. A customer address change entered on Wednesday may be backdated to Monday. A card charge may be authorized on Monday, booked on Tuesday and corrected on Friday. For historical replay, the analyst needs to know both when a fact applies in business terms and when the decision system knew it. Using the final corrected state in a Monday model back-test creates look-ahead bias.
A useful feature contract therefore names the decision timestamp, source event timestamp, ingestion timestamp, effective interval, allowed lag and missing-data treatment. Some implementations use bitemporal tables to store valid time and system time. Others maintain immutable events and dated snapshots. The implementation can differ; the invariant is that a reviewer can reconstruct the value available at the original decision point and identify later amendments separately.
For Meridian's fraud feature, "number of posted debits in the previous 24 hours" sounds simple. Which time starts the 24-hour window, and in which time zone? Does a late posting with a prior value date enter the window? If it was absent at decision time, it cannot be retroactively added to the original score input. A back-test that uses the corrected ledger may be useful for studying outcomes, but it is not a faithful replay of the model's available evidence.
The bank can test this with four dates. Generate a decision at 10:00 Monday, a source event effective at 09:00 Monday, ingestion at 11:00 Monday, and a correction on Tuesday. The 10:00 score must not include the late-arriving event unless another approved real-time feed delivered it. A Tuesday investigation can display the correction, while the retained decision record keeps the Monday input. Write the expected values before running the feature job.
Point-in-time correctness also applies to product and relationship reference tables. A loan account can change repayment plan; a corporate entity can change ownership; an account can be closed. A historical event joined to today's product or relationship row may acquire the wrong terms. Effective-dated keys and explicit late-correction logic are safer than joining on a current flag. A model's apparent predictive power can be inflated by future information even when no one intentionally copied an outcome column.
Reconcile before attributing meaning
A data pipeline that successfully loads rows has not proved completeness. It needs a source-to-destination reconciliation at the unit relevant to its consumer. For account masters, compare eligible accounts, creations, closures and changes. For postings, compare counts and debit/credit amounts by source, date, currency, product and accepted/rejected state. For balances, compare dated snapshots and explain timing differences. An aggregate total can hide missing records offset by duplicates.
Suppose a source exports 100,000 postings for a day. The intake accepts 99,950, rejects 30 malformed records and detects 20 duplicates. Those categories account for all 100,000 received rows if they are mutually exclusive. A dashboard that reports 99,950 "processed" records without the 50 excluded rows is incomplete. If a repair later admits 25 of the 30 rejected records, the rerun must link those records to the original batch and avoid double counting. These invented figures illustrate a reconciliation contract, not a universal tolerance.
Reconciliation should separate transport, syntax, business and accounting checks. Transport asks whether each expected file, event partition or API response arrived. Syntax checks field formats and types. Business checks whether an account, currency or product reference is valid at the event time. Accounting reconciliation checks whether postings or control totals match the source and ledger under the defined period. A model should not score an account whose source record failed a critical business check merely because the pipeline has a non-null feature value.
Every mismatch needs a disposition. An operator can hold publication, publish a marked partial cohort under an approved policy, repair the source, or exclude affected records with traceable reasons. The choice depends on the model's use and consequence. An analytics report may tolerate a small identified lag with a clear completeness label. A credit decision that uses an incomplete income history may need referral or fresh evidence. A fraud service facing a missing account-status feed should use its approved fallback. A generic "quality passed 99%" metric cannot decide for all three.
A control total is only as good as its population. If a bank compares a source's booked-payment total with a warehouse table that also includes pending instructions, the difference is expected, not a defect. Analysts should align business event, currency, cutoff, time zone, reversals and treatment of internal movements. Record the equation and exceptions. For a daily posting feed, one candidate equation is beginning balance plus credits minus debits plus specified adjustments equals closing balance. The actual source's sign convention and cutoffs govern whether that equation is valid.
The Basel Committee's 2026 note identifies lineage and timely, accurate ad hoc reporting as continuing challenges for banks. Its scope is risk data aggregation, not a prescribed reconciliation formula for this fictional example. The useful design lesson is to connect a reported number to its inputs, transformations and exceptions so a reviewer can challenge it.
Lineage is a replayable path
A lineage diagram should answer more than "source table A feeds feature table B." A reviewer needs the source field and version, extraction method, mapping rule, quality checks, intermediate dataset, feature computation, model input and decision record. At each step, record an owner and an effective time. The path should also identify where an identifier changes or a field is aggregated. A customer-level total cannot be traced to one posting without a membership list or query definition.
For Meridian's balance feature, the lineage starts with account and balance records. A mapping chooses eligible products and customer-account roles. A job chooses a dated snapshot and converts currencies when policy requires it. A quality gate rejects stale or ambiguous values. A feature service publishes a value with its version and observation cutoff. A model consumes that version, and the decision log links the feature snapshot to an action. If the feature is corrected later, the original input remains recoverable.
Lineage can be tested from both ends. Forward tracing asks which models and reports consume a changed source field. Backward tracing starts with a disputed customer decision and finds the input records that shaped it. Both are needed for impact analysis. If a product code changes meaning next month, forward tracing shows whether an AML alert prioritization or credit model needs retesting. If a customer challenges a declined application, backward tracing identifies the dated account information and any missing values the model used.
A lineage catalog cannot automatically establish business correctness. It can show that a calculation took a field from a table, but not that the field meant "available" rather than "booked" balance or that a joint-account relationship was eligible for the use. The business owner must sign off the semantic mapping. The catalog should link to that definition and test evidence. A simple, accurate mapping for a high-impact feature is preferable to an impressive automated graph that hides an unresolved meaning.
The FFIEC IT Architecture, Infrastructure, and Operations booklet announcement describes examination attention to architecture, governance, operations, and interconnected assets in the U.S. supervisory context. It does not prescribe Meridian's particular data model. Its relevance is the need to make dependencies and ownership inspectable, including third-party services when they carry source or derived data.
Missing values have several causes
"Missing" can mean the source did not collect a field, an upstream feed failed, the account is too new, an event arrived late, a relationship is unresolved, the customer declined to provide information, or the value was suppressed under an access rule. These states should not automatically become one zero or "unknown" code. Their consequences differ. A new account with no 30-day history is not the same as a longstanding account whose 30-day feed is broken.
The feature contract should identify permissible null states and how each appears in training and live scoring. A product may not use a card-balance field; that is not a quality failure for a deposit-only account. An income field might be absent because the bank never collected it; a model should not infer income from a zero placeholder. If a critical source is temporarily unavailable, the service can return an explicit incomplete status and route according to policy. Imputation, when appropriate, is a documented model choice and must be tested on the population where it will be used.
Missingness itself can be predictive, but the bank should ask why. If a channel rarely collects optional data, a model may learn the channel rather than customer risk. If a legacy branch has delayed updates, a missing indicator may encode operational weakness and create unequal treatment. Segment the missing-value rate by source, product, channel and time. Investigate a step change after a release before treating it as customer behavior.
A quality dashboard should show completeness alongside validity, uniqueness, consistency and timeliness. Define each measure at the field and population level. "100% non-null" can coexist with every record carrying a default that hides a failed feed. Sample source documents or transactions to verify that plausible values are real. Assign an owner and action to threshold breaches; an alert with no correction path becomes a decorative metric.
Access is part of data meaning
A bank can possess a record without being entitled to reuse it for every model. A fraud investigator, credit underwriter, marketing analyst and customer-service agent have different purposes and permissions. The data owner must map each proposed use to the applicable law, consent or other lawful basis where required, contractual restrictions, internal policy, retention and access controls in the relevant jurisdiction. The existence of a convenient warehouse copy does not grant a new purpose.
For a feature design, document the minimum necessary fields, permitted recipients and retention period. A payment narrative may contain free text that reveals sensitive information unrelated to the model. A customer-service note may record an unverified allegation. A model should not absorb those fields merely because a broad extract contains them. When a restricted value is transformed into an aggregate feature, assess whether the feature itself remains sensitive or can reveal the original information. Pseudonymized identifiers can still permit linkage; they are not a blanket exemption from governance.
Access tests should reflect the journey. Can the feature developer see raw account identifiers? Can a case reviewer see only the facts needed for that case? Does a vendor receive data outside the approved region or purpose? Are logs and error traces carrying unmasked values? A feature service can enforce role- and purpose-aware access, but its controls must extend to cache, exports and model-monitoring datasets. Record who can correct a source value and who can only flag it for the source owner.
A correction request should not silently alter every historical decision. The bank may need to fix current customer data, assess affected past decisions and keep an immutable explanation of the value used before correction. A customer remedy process may involve re-evaluation, notice or other action under applicable law and bank policy. The exact obligation varies by product and jurisdiction. The data architecture should make the investigation possible rather than assume the corrected latest row tells the whole story.
The NIST AI Risk Management Framework is voluntary and organizes AI risk work around govern, map, measure and manage. It supports a use-specific assessment of data quality and impacts, but it does not create a universal permission to process customer data. Legal and privacy owners must identify actual obligations for the bank and use case.
Legacy migration can change a feature without changing its name
A bank migrating accounts from an old platform to a new one may change identifiers, product codes, balance cutoffs and posting conventions. A model can receive the same column names but different business values. Before cutover, map source-to-target fields with sample cases covering active, dormant, closed, joint, overdrawn, multi-currency and exception accounts. A migration record should preserve old and new identifiers with effective dates so historical activity remains attributable without accidental merges.
Parallel-run checks compare the old and new source views for a defined cohort. Some differences are expected: a source may show a pending hold that another does not, or a migrated product may have a new status taxonomy. Classify each mismatch as an intentional mapping, timing difference, defect or unresolved item. Do not force a zero difference by hiding exceptions in "other." The model owner needs to know which features change, how many records are affected and whether historical training data can be made comparable.
Suppose Meridian migrates a product that previously represented an overdraft as a negative current-account balance. The new system records a separate facility utilization. A feature called "days with negative balance" would fall sharply after migration, even though customer behavior is unchanged. A drift monitor might alert, and a model might shift its score distribution. The right response is to investigate the semantic change, define a comparable feature if valid, back-test affected cohorts and approve a controlled model change. Blind retraining would teach the model a false behavioral shift.
Dual feeds create a second hazard. During migration, the same posting may appear in both legacy and target extracts. Reconciliation must identify the authoritative source for each account and effective date, plus a deduplication key. An account count that matches the expected total may still contain one omitted account and one duplicated account. Sample traced account histories across cutover, including a reversal that arrives after an account moves. A source-ownership matrix should say who resolves that exception.
A complete data contract for one feature
Take a fictional credit-risk use that estimates the probability of a specified delinquency outcome over a stated horizon. It proposes a "monthly net account inflow" feature. The contract names the eligible customer and account population, accepted holder roles, included posting types, exclusions for internal transfers, treatment of reversals, calendar month and time zone, currency conversion method, source owner and refresh schedule. It names the earliest decision time at which the month's value can be used. It also defines missing and disputed records and links the value to a versioned feature calculation.
The bank should test the contract with cases that are easy to get wrong. A salary payment arrives at 23:59 local time on the last day of the month. A duplicate message is rejected. A legitimate credit posts and reverses two days later. The customer transfers between own accounts. A joint account changes holder role mid-month. One source file arrives after the feature job. A foreign-currency amount is converted using an approved dated rate. Each case has an expected contribution to the feature and a source record proving it.
An analyst can calculate a transparent example. Suppose eligible credits are 2,000 and 500 units, eligible debits are 900, and an internal transfer of 300 appears as both a debit and credit. If the definition excludes both internal legs, net inflow is 2,500 minus 900, or 1,600 units. If a 200-unit credit is reversed within the defined window, the contract may subtract it, making 1,400. Those values are illustrative. The definition must state whether a reversal after the cutoff revises a published feature or remains a later correction; it cannot silently change the historical input to an already completed decision.
The test pack also checks source failure. If the posting feed is incomplete for a day, a partial sum should not look like a valid low inflow. The feature returns a quality status and the decision workflow follows a documented fallback. If a source field changes name, mapping tests fail visibly. If the relationship link is ambiguous, the job excludes or refers the account under approved policy. If a revised record arrives, lineage shows the old and new derived values and any decisions that used the old one.
A completed contract has an owner, sign-off and review trigger. New products, channels, currencies, source versions or decision uses can change its scope. Reuse is safe only if the next use shares the definition and quality threshold. A reporting dashboard may accept a delayed final value while a live fraud decision needs a current provisional one. Give those separate names rather than promising one "enterprise feature" will meet every clock.
Conflicting source claims need a resolution rule
Two systems can each be authoritative for a different aspect of one customer event. The account system may own the booked balance. The card processor may own the authorization and merchant details. The payments platform may own the instruction and rail status. The customer master may own the current legal name, while a historical document retains the name used when an application was signed. The data contract should name field-level authority, not nominate one database as the universal truth.
A source-precedence table needs conditions. If a balance extract and the account API disagree, first compare their as-of times and balance bases. If one is end-of-day booked and the other is intraday available, there may be no contradiction. If both claim the same dated booked balance, the reconciliation owner investigates source versions, corrections and missing postings. An analytics engineer should not choose whichever number makes the model run. Record the mismatch and its operational disposition.
A bank may build a curated customer view that combines these sources. The view should show provenance for each field, a quality status and a last-observed time. A consumer can then decide whether a field is suitable for its purpose. A customer-service screen may display a provisional transaction with an explicit pending status. A month-end accounting report needs a reconciled booked view. A real-time fraud model can use a timely authorization signal that is not yet a posting. The source distinction should survive into the interface.
Source contracts also need change notices. A product owner introducing a new account status should notify feature owners before the new code reaches production. A data owner should publish schema, code-list and semantic changes with effective times and test records. Downstream teams should confirm whether the new state is eligible, excluded or referred, rather than allowing an unknown code to default to "active." Change detection can flag a previously unseen value, but only business review can establish its meaning.
During an incident, preserve evidence of the conflicting records. Keep source identifiers, timestamps, request and response versions, pipeline job and correction trail. A team can then determine which decisions were affected and whether remediation is needed. If the bank overwrites the earlier observation, it loses the distinction between a bad original source, a delayed correction and a downstream mapping error. That distinction changes where the fix belongs and which customers may need review.
A particularly useful acceptance check is a cross-source dispute. Give the same fictional account a posted debit, a pending authorization, a later reversal and a warehouse refresh that arrives out of order. Ask the account owner, payments owner and model owner to state the value each system should show at four times. If their answers differ, compare the definitions before changing code. The test should capture the source status, booked and available balance, derived feature and customer-facing description at each time. A tester can then find a semantic regression that an API schema test would miss. The resolved case becomes a reusable regression example for future migrations and feed changes.
The release test is a reconstructable decision
The team can review one model-assisted decision end to end. Start with a decision identifier and timestamp. Retrieve the customer or account identity as known then, the product and relationship versions, balance basis, relevant postings, source arrival times, feature calculation, quality flags, model input, policy action and customer outcome. Then identify any later correction. If the reviewer cannot explain a field's meaning or why a record joined, the feature is not ready for an irreversible action.
Use a small, deliberately difficult sample. Include a newly opened account, a joint holder, a migrated product, a reversed fee, a late posting, a missing feed and a multi-currency balance. The expected result for each should be written independently of the model score. Confirm that the data platform and case tools agree on identifiers and states. The reviewer should be able to replay the original input, not merely run today's query and obtain a plausible answer.
Separate three release outcomes. A valid, timely source record can be used under the approved policy. An incomplete but recoverable record can be routed for more evidence or human review if that path has capacity and meets the customer's required timing. A critical identity, entitlement or source-status failure stops the model-led action and invokes the approved fallback. The specific treatment belongs to the use case; a bank should not silently turn every unavailable field into a default and call the service resilient.
Core banking data earns the word "foundation" when its object identities, dated states and financial events survive such a replay. Better algorithms cannot repair an account joined to the wrong person, a booked balance presented as spendable cash or a late correction disguised as an original fact. The bank can build useful models on the records when it can state what each value meant, when it was known, which source owned it and how a customer-impacting error would be found and corrected.
This application uses JavaScript for the full interactive experience. This text summary is served for accessibility and search indexing.