Data lakes and curated data products

Data lakes and curated data products. A practical lesson in the data foundation for banking and payments practitioners.

How to study this topic

A bank data lake becomes useful for AI only when raw data is converted into governed, documented, access-controlled, quality-checked data products that business, risk, technology, and model teams can trust. Read this chapter as a banking operating model lesson, not as a technology sales note. A learner should be able to explain where the data starts, what controls touch it, what business meaning it carries, and what can go wrong if the bank feeds it into a model too quickly.

The important point is scope. A bank is not one channel and one system. A single customer can appear through raw landing zone, standardised zone, curated zone, data product ownership, metadata catalogues, and privacy classifications. The AI layer sees only data, but the bank has to remember the process behind the data: who captured it, whether it is final, whether it has been corrected, whether it is legally usable, and whether it matches the official book of record.

What data lakes and data products really means inside a bank

data lakes and data products is not just movement of records from one place to another. It is a controlled translation from operational reality into analytical evidence. In a bank, operational reality is messy. A customer changes address. A loan repayment is reversed. A collateral valuation is refreshed. A complaint is reopened. A balance is available for service but not yet final for accounting. A risk flag is valid today but expired tomorrow.

A mature bank treats data lakes and data products as part of the control environment. It defines owners, source systems, event times, business dates, cut-off rules, repair rules, enrichment rules, privacy tags, retention requirements, lineage, and exception ownership. That discipline is what separates a reliable banking AI foundation from a pile of interesting data.

Banking scope and source systems

The sources for this topic can include raw landing zone, standardised zone, curated zone, data product ownership, metadata catalogues, privacy classifications, feature-ready datasets, audit retention, unstructured documents, logs and events, cloud object storage, and lakehouse tables. Some are customer-facing, some are colleague-facing, and some are hidden operational engines. The learner should not assume that customer-facing channels are always the best source. A mobile screen may show an intent, a workflow system may show an action, but the core platform or ledger may show what became final.

Why this foundation matters for AI and ML

The common use cases are document intelligence, behavioural analytics, feature engineering, model training, case notes analysis, call transcript analytics, scenario simulation, and cross-domain risk insight. These are valuable, but they are also sensitive. A wrong score can refer a good customer, miss a stressed borrower, over-prioritise the wrong case, misstate risk, create unnecessary manual work, or give a relationship manager poor guidance. The bank must therefore treat data preparation as part of model risk management, not as a back-office technical chore.

Controls before the data reaches a model

Before data is used for a model, the bank should apply checks for data swamp, sensitive data leakage, unclear ownership, poor discoverability, unvalidated raw data use, and uncontrolled notebooks. These checks are not decorative. They prevent a model from learning the wrong lesson. A duplicate customer record can inflate behaviour. A stale risk rating can misclassify risk. A late file can make yesterday look safe when it was incomplete. A wrong join can attach one customer’s behaviour to another customer.

Controls should run at several levels: file or event control, schema control, field control, referential control, reconciliation control, privacy control, lineage control and business reasonableness control. Technical validation catches format and processing errors. Banking validation catches meaning errors. The strongest platforms use both.

Common mistakes in banks

The last mistake is over-automation. AI can support prioritisation, classification, prediction and explanation, but the bank must decide where human review remains mandatory. High-impact credit, compliance, customer harm, regulatory reporting and financial statement use cases need stronger controls than low-risk internal productivity use cases.

A simple bank-ready checklist

Before approving data lakes and data products for model use, ask whether the bank can answer these questions. What is the book of record? What is the event time and business date? What fields are mandatory? What quality thresholds apply? What reconciliation proves completeness? What privacy rules apply? What transformations are allowed? What happens when the data is late, partial or corrected?

Source anchors for further study

Use the Basel Committee's BCBS 239 principles to understand why accuracy, completeness, timeliness and adaptability matter for banking data and risk decisions.

Use the Basel Committee's digitalisation work to understand why APIs, AI, cloud, third parties and digital channels increase both opportunity and operational risk in banking.

For U.S. banking organisations within scope, the Federal Reserve's SR 26-2 revised model-risk guidance superseded SR 11-7 in April 2026. It calls for risk-based development, validation, monitoring and governance tailored to model use; other jurisdictions require their own assessment.

Raw storage and model-ready data are different promises

A data lake can retain source files, events, documents and semi-structured records that do not fit a traditional reporting table. That flexibility is useful for ML exploration and replay, but raw storage is not automatically governed data. A model-ready product needs a defined population, schema, source, quality state, timestamps, permitted use, retention and owner. A file in a bucket with a readable name is not an approved feature dataset. The bank should distinguish a raw immutable zone, a validated and standardized zone, and a curated product whose business definitions are fit for a specific consumer. The physical platform can vary; the control boundary cannot be skipped.

Consider a payment-event lake containing channel orders, hub decisions, Swift messages, case events and account reporting. A data scientist may want a journey table for repair prediction. They must map customer, instruction, message, UETR, case and ledger references without assuming one-to-one relationships. A raw file can contain several transactions, a payment can have multiple status updates, and a return has its own financial movement. The curated product states what one row represents, how events are ordered, which statuses are authoritative at each boundary, and how late corrections are handled. Otherwise a model may learn from events that occurred after its intended decision time.

Publish a data product with a contract

The product description should name its business owner, technical owner, consumers, source list, grain, field meanings, effective and availability times, update cadence, quality measures, access rules and change policy. It should include example rows and edge cases, not only a schema. A field called current_status must say which actor reported it and as of when. A derived repair_reason must distinguish a source reason code from an analyst's later interpretation. The contract can expose a stable version while a new mapping is tested. Consumers should know when a release changes field semantics even if the JSON type remains string.

A product is not validated by a successful pipeline alone. Reconcile it to source control totals and selected operational cases. Count accepted, rejected and quarantined events. Inspect a payment with a recall, one with a return, one still under review, and one with an uncorrelated external status. If the product cannot represent these states without overwriting history, it is too weak for training or customer-status AI. A model use approval should specify which version of the product was tested and what happens when a new source or message version enters it.

Lakehouse tables and point-in-time reads

Versioned tables can make it easier to reproduce a dataset, but a snapshot identifier alone is not enough. A table snapshot taken today can contain a backfilled record with last month's event time. An ML training row for last month's decision must use information that had reached the bank by then. Keep ingestion and publication time, source revision, effective time and table version. A point-in-time query should enforce both the business cutoff and availability cutoff. Test a record that arrives late, a corrected customer identity and a file reprocessed after a parser fix. The result should match the evidence used by the live service, not merely the latest cleaned historical view.

Schema evolution needs explicit rules. A new optional field may be safe for old consumers, but a code-set expansion can break a model that one-hot encodes known values. A renamed field might map to a different concept despite looking compatible. Validate downstream feature distributions and model outputs before promotion. Keep the raw source payload for investigation, subject to retention and privacy rules, and record the parser and transformation versions. If a downstream model sees an unknown category, its approved fallback should be tested rather than relying on an accidental library default.

Documents and unstructured inputs

A lake may hold policies, complaints, correspondence and investigation notes for RAG or classification. These are not automatically safe to index. Determine document authority, effective date, language, access level, retention and whether the text contains personal or confidential data. A RAG system should retrieve from an approved corpus and cite the exact version used in an answer. An old policy retained for audit may be searchable for historical replay but should not be presented as current guidance for a new case. A complaint summary model needs safeguards against mixing customers or inventing a conclusion unsupported by the source.

Unstructured text can carry hidden duplication and bias. The same adverse-media article may appear from several providers; treating each copy as independent evidence inflates a signal. A case note written after an investigator's decision cannot be used as an earlier input to predict that decision. A model trained on only escalated notes may learn the old escalation process rather than underlying risk. Curated products should mark source, author, creation time, availability time, deduplication group and permitted use. Reviewers need a route back to the original, not only a vector embedding.

Access, retention and removal

Data products often serve several teams. That does not mean every consumer can use every field. Apply purpose-based access to raw messages, personal identifiers, sanctions notes and model outputs. Anonymization or tokenization can reduce exposure but does not erase the need to assess linkage risk. The product owner defines retention and deletion behavior for source and derived records. A model trained on data that later becomes impermissible may require investigation of artifacts and retraining under applicable policy. The bank should know which models consumed which product versions.

If a vendor feed contract expires, a lake cannot continue distributing it merely because old files remain. The inventory should identify derived tables, features and models that depend on the feed. A replacement source may differ in coverage or meaning; compare distributions and performance before switching. If no replacement is approved, restrict the dependent model or invoke its fallback. The change should be visible to operations, risk and model validation, not hidden as a data engineering task.

Release scenario

At 10:00 a fraud model scores a payment using a curated event product. At 10:05 a late channel event arrives with an event time of 09:55. At 11:00 a customer master merge changes the linked identity. At day-end a status report corrects the payment's earlier state. The current lakehouse table may show all three facts together. The 10:00 score must be replayable with only the information available then. The team stores the product version and feature snapshot used for the score and appends later corrections with their own availability times. A validator compares online and offline values for this case and explains any difference.

The product passes release when a consumer can find its owner and contract, a data steward can reconcile it to sources, a model owner can use it within approved purpose, and an auditor can replay a selected decision. Test a missing file, an unknown code, a late event, a restricted document and a vendor contract change. A data lake becomes an AI asset when it makes these states controllable and observable; storage volume alone is no measure of readiness.

A product scorecard for actual use

Measure the data product against the decision it supports. For the payment journey product, report how many source instructions can be linked to a hub execution, external message, status, booking and customer notification, with legitimate gaps explained. Report unmatched or conflicting references by corridor and source version. Show event lag and corrections separately from ordinary missingness. A perfect schema-validity percentage cannot compensate for an incorrect join between two payments with the same amount. Sample those joins with operations and reconciliation evidence.

For a RAG policy product, report the percentage of indexed documents with an approved owner, effective date, access classification and supersession link. Test that a user lacking rights cannot retrieve a restricted investigation note through an embedding search or a generated answer. Remove a policy from the current corpus and verify that historical replay can still access the archived version under controlled authority. A search result that cites a file name without the exact clause and version is weak evidence for an operational decision.

For a credit feature product, compare online and offline values at sampled decision times, including late corrections. Track source freshness, population coverage and reasons for missing values. Review whether the product continues to represent the intended borrower group after a new channel or acquisition changes the population. Publish a version change with examples of affected cases and downstream models, then require consumers to acknowledge it. The strongest product scorecard combines technical health, business meaning and model consequence, so teams know what to do when a feed is green but the decisions are wrong.

Banking practice note on customer impact

For data lakes and data products, the customer impact angle matters because banking data is never neutral once it supports a recommendation, score, report or automated workflow. A bank may begin with a technical ingestion pattern, but the practical question is whether the data can safely support document intelligence, behavioural analytics, feature engineering, and model training. That means teams must inspect source reliability, product meaning, customer status, timing, lineage, access, consent, reconciliation and exception handling before calling the dataset model-ready. When this discipline is skipped, the model may still perform well in a narrow test while failing in real production conditions where customers change behaviour, products migrate, systems send corrections, and risk teams need evidence.

A useful control habit is to separate observed fact, derived feature, model assumption and business decision. Observed facts come from systems such as raw landing zone, standardised zone, curated zone, data product ownership, and metadata catalogues. Derived features transform those facts into signals. Model assumptions decide how signals are interpreted. Business decisions decide what action follows. Keeping these layers separate helps the bank explain the result without pretending that the model itself owns the banking decision.

Banking practice note on risk management

For data lakes and data products, the risk management angle matters because banking data is never neutral once it supports a recommendation, score, report or automated workflow. A bank may begin with a technical ingestion pattern, but the practical question is whether the data can safely support document intelligence, behavioural analytics, feature engineering, and model training. That means teams must inspect source reliability, product meaning, customer status, timing, lineage, access, consent, reconciliation and exception handling before calling the dataset model-ready. When this discipline is skipped, the model may still perform well in a narrow test while failing in real production conditions where customers change behaviour, products migrate, systems send corrections, and risk teams need evidence.

Banking practice note on operational resilience

For data lakes and data products, the operational resilience angle matters because banking data is never neutral once it supports a recommendation, score, report or automated workflow. A bank may begin with a technical ingestion pattern, but the practical question is whether the data can safely support document intelligence, behavioural analytics, feature engineering, and model training. That means teams must inspect source reliability, product meaning, customer status, timing, lineage, access, consent, reconciliation and exception handling before calling the dataset model-ready. When this discipline is skipped, the model may still perform well in a narrow test while failing in real production conditions where customers change behaviour, products migrate, systems send corrections, and risk teams need evidence.

Banking practice note on regulatory evidence

For data lakes and data products, the regulatory evidence angle matters because banking data is never neutral once it supports a recommendation, score, report or automated workflow. A bank may begin with a technical ingestion pattern, but the practical question is whether the data can safely support document intelligence, behavioural analytics, feature engineering, and model training. That means teams must inspect source reliability, product meaning, customer status, timing, lineage, access, consent, reconciliation and exception handling before calling the dataset model-ready. When this discipline is skipped, the model may still perform well in a narrow test while failing in real production conditions where customers change behaviour, products migrate, systems send corrections, and risk teams need evidence.

Banking practice note on data ownership

For data lakes and data products, the data ownership angle matters because banking data is never neutral once it supports a recommendation, score, report or automated workflow. A bank may begin with a technical ingestion pattern, but the practical question is whether the data can safely support document intelligence, behavioural analytics, feature engineering, and model training. That means teams must inspect source reliability, product meaning, customer status, timing, lineage, access, consent, reconciliation and exception handling before calling the dataset model-ready. When this discipline is skipped, the model may still perform well in a narrow test while failing in real production conditions where customers change behaviour, products migrate, systems send corrections, and risk teams need evidence.

Banking practice note on model limitation

For data lakes and data products, the model limitation angle matters because banking data is never neutral once it supports a recommendation, score, report or automated workflow. A bank may begin with a technical ingestion pattern, but the practical question is whether the data can safely support document intelligence, behavioural analytics, feature engineering, and model training. That means teams must inspect source reliability, product meaning, customer status, timing, lineage, access, consent, reconciliation and exception handling before calling the dataset model-ready. When this discipline is skipped, the model may still perform well in a narrow test while failing in real production conditions where customers change behaviour, products migrate, systems send corrections, and risk teams need evidence.

Banking practice note on business process design

For data lakes and data products, the business process design angle matters because banking data is never neutral once it supports a recommendation, score, report or automated workflow. A bank may begin with a technical ingestion pattern, but the practical question is whether the data can safely support document intelligence, behavioural analytics, feature engineering, and model training. That means teams must inspect source reliability, product meaning, customer status, timing, lineage, access, consent, reconciliation and exception handling before calling the dataset model-ready. When this discipline is skipped, the model may still perform well in a narrow test while failing in real production conditions where customers change behaviour, products migrate, systems send corrections, and risk teams need evidence.

Banking practice note on auditability

For data lakes and data products, the auditability angle matters because banking data is never neutral once it supports a recommendation, score, report or automated workflow. A bank may begin with a technical ingestion pattern, but the practical question is whether the data can safely support document intelligence, behavioural analytics, feature engineering, and model training. That means teams must inspect source reliability, product meaning, customer status, timing, lineage, access, consent, reconciliation and exception handling before calling the dataset model-ready. When this discipline is skipped, the model may still perform well in a narrow test while failing in real production conditions where customers change behaviour, products migrate, systems send corrections, and risk teams need evidence.

Banking practice note on privacy and access

For data lakes and data products, the privacy and access angle matters because banking data is never neutral once it supports a recommendation, score, report or automated workflow. A bank may begin with a technical ingestion pattern, but the practical question is whether the data can safely support document intelligence, behavioural analytics, feature engineering, and model training. That means teams must inspect source reliability, product meaning, customer status, timing, lineage, access, consent, reconciliation and exception handling before calling the dataset model-ready. When this discipline is skipped, the model may still perform well in a narrow test while failing in real production conditions where customers change behaviour, products migrate, systems send corrections, and risk teams need evidence.

Banking practice note on change management

For data lakes and data products, the change management angle matters because banking data is never neutral once it supports a recommendation, score, report or automated workflow. A bank may begin with a technical ingestion pattern, but the practical question is whether the data can safely support document intelligence, behavioural analytics, feature engineering, and model training. That means teams must inspect source reliability, product meaning, customer status, timing, lineage, access, consent, reconciliation and exception handling before calling the dataset model-ready. When this discipline is skipped, the model may still perform well in a narrow test while failing in real production conditions where customers change behaviour, products migrate, systems send corrections, and risk teams need evidence.

A lake is not a permission to mix meanings

A data lake can hold payment events, loan extracts, customer reference, documents and case records in their source forms. That flexibility supports AI experimentation, but it can also hide inconsistent units, clocks and permissions. A curated data product should publish a specific business definition, population, quality status, owner and access contract. A fraud training product and a credit risk product can draw from the same lake yet need different feature cutoffs and labels.

Separate landing from publication. Raw files retain source identity, schema, ingestion time and integrity checks under restricted access. Curated tables resolve status vocabularies, keys, reference mappings and corrections with documented rules. Model-ready products add point-in-time features and outcome labels for a defined use. Do not let a notebook read a raw bucket and call its result the bank's canonical transaction history.

Product contract

Define a product such as "eligible outbound payment events for fraud features." It should state which channels and statuses are included, what constitutes one business instruction, how retries and returns are represented, and how source control totals reconcile. Include event and availability timestamps, currency units, key scope, schema version and freshness watermark. List exclusions and a query to find affected decisions when a source partition is missing.

A customer feature product needs an entity definition and relationship version. Joint accounts, merged profiles and beneficial ownership cannot be reduced to an unqualified customer ID. A bureau product needs observation and receipt dates plus match and dispute states. A policy-document product needs authority, approval and effective dates. Each product should expose what its consumers may infer and when it is unsuitable.

Publishing a product includes quality gates. Check partition arrival, duplicate IDs, invalid values, join cardinality and coverage by critical product or channel. A global completeness percentage can hide a missing high-value corridor. Report source-to-publication delay and validity status. If a gate fails, block the product version or mark a narrow partial release with explicit consumer restrictions; do not silently substitute stale data.

Version and correction

Use immutable dataset or table versions for training and material reporting. A later source correction can create a new curated version linked to the old one. A model validation report should identify the exact version it used. If an analyst reruns a query against a mutable "latest" table, changed results may reflect data repair rather than model improvement. Preserve transformation code and manifests.

Time travel in a lakehouse can recover a table state, but it does not automatically reconstruct what a live model knew months earlier. A late-arriving event can have an old event time but a new ingestion time. A current customer mapping can be retroactively effective yet not have been available at the earlier decision. Build point-in-time training features using both business and availability clocks and compare sample rows with saved production requests.

An event correction should not erase the original banking action. If a payment later returns, retain instruction, acceptance, settlement and return as linked lifecycle events. If a customer relationship was wrong, keep a controlled original decision view and a corrected analytic view. Reverse lineage should enumerate decisions that consumed the old mapping and distinguish changed scores from changed final actions.

Document products for RAG

A lake can also hold policies, procedures, regulatory extracts and case documents. A retrieval assistant should not ingest every document in a bucket. Its curated corpus requires an owner, approval state, effective date, access scope, source version and extraction quality. Chunks and embeddings inherit those attributes. A vector index is a search representation, not the authority for policy content.

Test a superseded policy, a future-effective procedure, a restricted case note and an unanswerable question. The assistant should retrieve only authorized, current material or abstain. Preserve document IDs and passages supplied to the model, plus draft and human review. Rebuilding the index after a correction must not make an earlier unsupported answer appear to have used the new corpus.

OCR and parsing can change meaning. A scanned table can lose column alignment; a digit in a limit can be misread. Check source spans and critical values, not just text extraction success. A generated summary is not source evidence. A curator should be able to withdraw or correct a document and identify answers affected by its previous version.

Training product

A credit model training product might contain application-time cash-flow and bureau features with a later default label. It should define eligible applications, funded population, feature cutoff, outcome horizon and maturity. Rejected applicants lack a repayment outcome on the proposed loan; they should not be coded as good accounts. The product records source snapshots, joins, missingness and label revisions.

A fraud training product uses authorization-time features and later confirmed outcomes. Blocked transactions are selectively observed. Keep holds, investigation dispositions and mature labels separate. A lake can store them all, but a model-ready product must not join a later case note into the earlier feature vector. Independent validation should hand-replay examples from source availability timestamps.

Data-product versioning helps compare candidate models. If two models use different source coverage or label definitions, a raw metric comparison is misleading. Hold the product version and cohort constant when possible, and disclose differences otherwise. A versioned product does not guarantee fairness or suitability; it makes the assumptions inspectable.

Access and retention

Raw banking data can be highly sensitive. Restrict landing zones, curated tables, feature products, training exports and backups according to purpose. A developer may need derived values; an investigator may need a specific case; a compliance assistant may need approved policies but no customer documents. Tokenized identifiers can remain linkable, and vector indexes can reveal document content. Log access and exports.

Set retention with legal, privacy and business owners. Temporary experiment copies should not persist indefinitely. At the same time, a bank needs enough protected evidence to explain material past decisions and correct source defects. A manifest and source reference can reduce duplication if the underlying archive is dependable. Test retrieval after schema migration and disaster recovery.

Vendor access requires a defined data flow. A model vendor should not receive the entire raw lake because it serves one classifier. Limit fields, contract handling, monitor transfers and plan an exit. An external observability tool can receive sensitive payloads through verbose errors even when the primary model API is scoped; inspect actual traces.

Incident exercise

A new payment-hub version starts writing amounts in minor units while the curated product expects major units. Schema validation passes because both are numbers. Distribution checks and hand-reviewed examples detect the scale shift. The product owner blocks publication, identifies affected partitions and prevents invalid values from reaching fraud training and serving. If decisions were already made, reverse lineage lists requests that consumed them and their actual actions.

The corrected product receives a new version and reconciliation report. Training datasets built from the defective version are flagged; candidate models may need re-evaluation. Historical decision journals retain the original features and a separate corrected scenario. Customer remediation depends on final effects, not merely a changed numeric score. The incident is closed after source, product, model and business owners verify their respective paths.

Acceptance exercise

Choose one raw event, one curated row and one model-ready feature. Trace field meaning, key, unit, event and ingestion times, reference version, quality gate and authorized use. Inject a duplicate retry, a missing partition, a changed code meaning and a document withdrawal. Verify that consumers receive validity signals, versions remain reproducible and reverse lineage can identify affected decisions.

A lake becomes useful for banking AI when its curated products make meaning and quality explicit. The bank should know which source facts were available at a decision, which transformations produced a model input and what action followed when a product failed its contract.

From raw payment events to a governed product

Imagine a bank landing channel requests, payment-hub events, screening outcomes and ledger confirmations in a lake. The raw zone retains the source payload and receipt metadata under restricted access. A curated product called payment-decision-history is proposed for fraud-model training. Its owner must define more than file location: the grain is one model decision, the key identifies that decision and its instruction, and the published columns carry explicit types, timing and allowed purpose. A raw lake can store many versions of an event; a curated product must explain which version a consumer receives.

Start with a fictional transfer submitted at 11:00. The channel captures an amount and beneficiary. The hub enriches party identifiers at 11:01, fraud scoring occurs at 11:02, and screening returns a referral at 11:03. A later payment status and account posting arrive after noon. The fraud-training product should preserve the 11:02 feature view and link later outcomes separately. If the curator simply selects the latest record per instruction, the training row can include the screening referral or final posting before the model could have known either. A table partitioned by business date is not a point-in-time dataset merely because its rows have timestamps.

A schema change that preserves history

The hub changes its beneficiary type vocabulary from two broad values to five more precise values. Keep the raw code, its source schema version and the mapping used for the curated field. New values must not silently fall into an old default group. Validate counts by original and mapped value, test a historical sample, and publish a new product version with a migration note. If the model still expects the old two-value feature, the producer can support a documented compatibility mapping for a limited period. The risk owner must assess whether that mapping changes scores or excludes a material customer group.

Correction events need a clear contract. A source may send the original instruction, then a correction to a party identifier and a reversal of an erroneous posting. In the raw zone, retain all three events with their source IDs and receipt order. In the curated decision view, select only facts available at each score cutoff. In a current operational view, apply corrections according to a defined precedence rule. Mixing the two views in one table without an as-of parameter makes it easy for a data scientist to train on information that arrived later.

Ownership and consumer promises

Assign a producer owner for each feed, a curator for the transformation and a business owner for the product's intended use. The product contract should say whether data arrives continuously or after a batch close, the allowed lag, how duplicates are identified, how missing values are represented, and how a breaking schema change is announced. A consumer may tolerate a late outcome label in retrospective evaluation but cannot tolerate stale beneficiary velocity at a pre-release fraud decision. Publish freshness at the dataset and partition level so one delayed source does not hide behind a recent timestamp from another source.

Quality checks should span source and business meaning. Verify unique decision IDs, referential integrity to instructions, valid amount and currency pairs, recognised status values and reconciliation of counts to the source system. Then sample a payment that was held, one that was returned and one that completed normally. Trace each through raw records, curated decision features, eventual outcome and any exclusion. A dashboard showing 99.9% complete rows does not reveal whether the missing 0.1% are concentrated in held transactions, which could change the model's learned pattern.

A controlled rebuild

Suppose a source correction changes the mapping for one day of payments. The curator creates a new immutable product version, runs the same quality and reconciliation checks, and produces a difference report: added decisions, removed decisions, changed features and changed labels. Model validation compares affected populations and score outcomes using a frozen candidate model before deciding whether retraining or reapproval is needed. Historical decisions retain their original source and feature versions even if the new product corrects today's analytical view.

Access must follow purpose. An operations analyst may need case IDs and status, while a model training job may need pseudonymous entities and defined outcomes. A policy-retrieval index may require document access labels that are irrelevant to payments. Curated products should not copy every raw field simply because storage is inexpensive. Record the transformation and access grant so an auditor can identify who could see sensitive material and why. A usable lake product is a versioned promise about meaning, timing, quality and permission, backed by enough raw evidence to challenge it.

Keep the outcome contract apart from the event product

A confirmed fraud outcome is rarely available at payment initiation. It can follow customer contact, investigation, dispute or recovery, and its status may later change. The lake should publish a separate outcome product with label definition, evidence source, confirmation state and observation end date. A training join can then choose a sufficiently mature cohort without rewriting the earlier feature product. A missing confirmed label is not the same as a negative fraud result. Likewise, a payment that was held by a control does not reveal what would have happened if it had been released. The product should identify such interventions so model evaluation does not treat prevention as proof of a false alarm.

Give a consumer three sample instructions: one released and later confirmed fraudulent, one held and never settled, and one released with no mature investigation outcome. Require the consumer to show its feature cutoff, label eligibility and exclusion rule for each. Then change a late case disposition and verify that a new outcome version appears while the original decision features remain unchanged. This test makes product boundaries visible to model developers, validators and operations owners. It is more informative than a schema check because it exposes whether the dataset preserves the causal order of the bank's decisions.

Related learning paths

This application uses JavaScript for the full interactive experience. This text summary is served for accessibility and search indexing.

Data lakes and curated data products · Malla Banking Academy