Feature versioning and reproducibility

Feature versioning and reproducibility. A practical lesson in the feature store for banking and payments practitioners.

How to study this topic

Feature versioning and reproducibility allow a bank to explain exactly which data, definition, calculation, time window, quality rule, and feature value supported a model score or decision at a specific point in time. Read this chapter as a practical banking lesson. The subject is not only machine learning terminology. It is about how a bank converts controlled evidence into a signal that can support a model, dashboard, rule, human decision, operational queue or risk review.

The banking meaning

feature versioning and reproducibility matters because a feature is never just a column in a model table. In banking, a feature may affect whether a credit application is referred, whether a fraud case is escalated, whether a customer is treated as vulnerable, whether a compliance alert is prioritised, whether a portfolio looks stable, or whether a manager trusts a risk dashboard.

Source systems and evidence

Useful source areas include feature definitions, calculation code, source data versions, business calendars, time windows, training datasets, scoring runs, model versions, quality rules, exception exclusions, approval records, and audit logs. These sources do not have equal strength. Some are books of record. Some are event logs. Some are case-management evidence. Some are customer-declared information. Some are external signals. Some are derived from earlier analytics. The bank must know which source is authoritative for each feature.

Feature definition

For feature versioning and reproducibility, the definition must be specific enough to stop accidental reinterpretation. A vague feature name can create real damage. "Income stability" means different things if it uses salary credits, declared income, bureau income, employer data or account turnover. "Customer risk" means different things if it uses conduct risk, credit risk, fraud risk, AML risk or relationship value. Feature names must not pretend simplicity where the bank has complexity.

Data quality controls

Controls for this topic include version identifier, immutable training snapshot, point-in-time join, lineage record, change approval, and rollback path. These checks should run before a feature is used for training, scoring, monitoring or reporting. The point is not to create a heavy process for every experiment. The point is to match the control to the business impact of the feature.

Model use and decision boundaries

Common uses include model validation, audit review, regulatory evidence, incident investigation, model rollback, performance comparison, champion challenger testing, and customer complaint review. Each use has a different risk level. Internal exploration is different from model validation. Portfolio monitoring is different from direct customer decisioning. Fraud triage is different from automatic blocking. Compliance prioritisation is different from a final suspicious activity decision. Credit support is different from final credit approval.

Good AI governance is not anti-innovation. It prevents uncontrolled use. If feature use is clear, teams can build faster because they do not keep reopening the same argument. They know the approved source, approved definition, approved model use and approved control path.

Customer impact and fairness

A strong bank does not assume fairness because the model is mathematical. It tests, challenges and documents. It also recognises that a feature may be statistically useful and still inappropriate for a particular decision. Banking needs both predictive discipline and judgement.

Operational controls

Operationally, feature versioning and reproducibility needs monitoring. The bank should monitor freshness, volume, distribution, missing values, outliers, source outages, calculation failures, feature drift, downstream model impact and exception ageing. A feature that changes suddenly should trigger investigation before model users accept the result as a business trend.

The operating model should define who responds when the feature fails. Is it a source-system owner, data engineering team, model owner, product owner, risk team, compliance team, fraud operations, credit policy team or feature-store team? Without ownership, feature quality issues become slow and expensive.

Governance and audit

Audit evidence should include the feature definition, owner, source lineage, transformation logic, quality results, approval, version, access controls, monitoring history, issue history and model linkage. This evidence is especially important for credit, financial crime, regulatory reporting, model validation and customer-impacting processes.

Practical example

Imagine a bank wants to use feature versioning and reproducibility in a model. The weak approach is to pull fields from a warehouse, calculate a signal, test model performance and push it to production. The stronger banking approach starts earlier. It asks what the signal means, who owns the source, which customers are covered, which exclusions apply, what the time window is, and whether the result is suitable for the intended decision.

Only after that should the feature become reusable. Reuse is powerful, but it multiplies both value and error. A reusable feature must be easier to trust than a local one, not merely easier to consume.

Bank-ready checklist

If the answers are strong, the feature can support banking AI responsibly. If the answers are weak, the model may still run, but the bank is carrying hidden risk.

What a feature version means

A version is a claim that the same defined inputs under the same conditions produce the same feature value. Versioning only a code package is insufficient. A banking feature also depends on source schema, event meaning, customer relationships, classification rules, time windows and corrections. A salary-count feature may change because a payroll classifier recognizes another payment code, while the transformation code stays the same. A version record should identify the business definition, implementation, source contract, effective date and consumer approvals. A model decision should retain the value served and versions of the feature and model, as well as the policy that turned the score into an action.

Different changes need different treatment. Correcting a typo in a description may leave values unchanged. Changing an observation window from ninety days to three calendar months may alter boundary cases. Replacing an account mapping can alter a population of customers. Adding a source can improve coverage but make a model trained on old missingness behave differently. Classify whether the change is semantic, technical, source-related or a backfill, then test actual value differences. A familiar API field name cannot guarantee compatibility.

Reproduce a real decision

Consider a loan application at 14:00 on a Tuesday. The feature service returned a three-month income estimate, a debt-obligation value and a missingness reason. The model returned a score, then an affordability rule referred the application. A reproducibility pack contains application ID, source IDs and availability timestamps, relationship mapping version, feature definition and code versions, exact feature response, model artifact, score, policy version and human action. A reviewer should be able to trace the decision without depending on the current warehouse state. If the income feed was corrected on Wednesday, that update belongs in the investigation, not in a silent rewrite of Tuesday's input.

Store the actual online response where appropriate. A recomputation can help verify the calculation but might differ because a source event arrived late or an identity link was repaired. Both original and corrected views are valuable when explicitly labelled. The original shows what the bank used; the corrected view helps assess what should have happened and which customers may need remediation. A reproducibility test compares them and explains the difference. Merely rerunning today's code on today's data is neither a replay nor proof that a historical model decision was justified.

Four version dimensions

Source version identifies the feed's schema and business event interpretation. Transformation version identifies the calculation and feature semantics. Model version identifies the fitted artifact and preprocessing. Policy version identifies thresholds, deterministic checks and fallback actions. A change to any dimension may alter the customer result. A deployment manifest should pin compatible combinations and state the eligible population. If the model expects a count of accepted instructions but the feature API now returns settled payments, schema validation may pass while business meaning breaks.

An entity-resolution version can be a fifth material dependency. Merging two customer profiles changes the account set used for transaction counts even if source events and calculation code are unchanged. Capture effective and observed dates of links. For a historical decision, use the relationship known then. A current corrected relationship can be used for impact analysis. Tests should include merges, splits, joint holders and corporate groups, because those cases are likely to expose hidden dependency changes.

Change proposal and impact analysis

Before promotion, describe why the change is needed and which consumers will receive it. On a fixed dated sample, calculate old and candidate feature values side by side. Count changed, missing and out-of-range values by product and channel. Compare model scores, policy actions and relevant customer groups. A small average difference can conceal a large change for a thin-file segment. Review whether model validation, threshold approval, customer notices or operations training must change. Name a rollback owner and a date for post-release monitoring.

Suppose the bank revises its repayment-lateness feature to respect an agreed grace period. Many values may improve from one missed payment to zero, but some accounts near the boundary could move across a credit referral threshold. The reviewer checks the servicing schedule, posting rules and sample cases rather than approving based on the total number of changed records. A historical backfill can create a restated training dataset; label it separately from values actually served. If a challenger model is trained on the new definition, it needs validation against production-like availability and an approved deployment combination.

Online and offline parity

An offline pipeline may recompute across complete historical files, while an online service consumes a delayed event stream. They can share a semantic contract yet differ in availability. For sampled live decisions, compare both implementations at the archived decision cutoff. Check value, null state, entity key, source coverage and last update. If the offline result uses a later correction, do not classify it as a parity defect without first specifying the expected historical view. The test should expose the difference between as-served and restated values.

Parity failures can arise from rounding, time zones, boundary inclusion, duplicate suppression or customer mapping. A counter that uses calendar days offline and rolling twenty-four-hour windows online may agree often and diverge near midnight. Write tests for those boundaries. If the online feature misses an entire channel, a global average comparison may hide it; segment by source and product. Retain discrepancies, responsible owner, fix and impact assessment rather than simply forcing the values to match in a later refresh.

Reproducible training

A training dataset should record the exact eligible population, observation cutoffs, label horizon, source snapshots and feature versions. Store code and configuration, random seeds where stochastic steps exist, dependency versions and a manifest of included and excluded records. A model can be reproduced technically yet still be invalid if labels use future information or if development records were selected by a previous policy. The training record therefore includes business assumptions and known selection limitations, not just hashes.

When a label changes after a fraud investigation or loan recovery update, keep its earlier state and correction time. State which label snapshot was used for a training run. A later outcome revision may warrant a new evaluation, but it should not retroactively alter the evidence supporting an earlier model approval. A validator can take one sampled observation, reconstruct its features as of decision time and trace the outcome over the defined horizon. If that cannot be done, a metric in a report is difficult to challenge.

Rollback and incident evidence

Rollback is more than redeploying old code. An earlier feature transformation may require an older source mapping or online cache schema; a model may expect the newer feature value. A rollback plan identifies compatible versions across source, feature, model and policy, plus the state of in-flight decisions. Test it before a material change. If a defective feature was served for several hours, preserve logs before refreshing caches. Identify affected decisions, compare original and corrected values and assess any customer or regulatory consequence under bank policy.

Versioned logs should identify a decision ID and include request and response times. For batch scores, record snapshot date, run ID, input manifest, rejected records and publication status. A partial batch should not be overwritten with a corrected file under the same run ID. A superseding run can be clearly linked to the earlier one. This enables risk and operations teams to reconcile how many accounts were eligible, scored, skipped or acted on.

A concrete boundary test

Define a prior-seven-day distinct-beneficiary feature. One payment occurs exactly seven days before the score, a second is submitted twice after a timeout, and a third is corrected after the score. The expected count depends on the approved inclusive boundary, deduplication and destination identity rules. Freeze the raw event sequence and availability times. Run it through each feature version and compare as-served online, historic offline and current restated results. Record why the values differ. Add a customer-account merge and a source outage to expose identity and freshness dependencies.

The same test should feed an approved model and policy version. Does a changed count cross a hold threshold? Does an unavailable source trigger referral rather than zero? Can an investigator see the underlying events without unrestricted access? A versioning policy is credible when these questions can be answered from a concrete decision trace, not only from a release tag.

A classifier update through the version chain

Imagine a payroll classifier that previously treated a particular employer code as an ordinary transfer. The data owner verifies that it is a salary code for one product and proposes a mapping update. The feature owner tests whether the change affects income averages for applicants with three complete months of history, for new applicants with only one month, and for customers receiving mixed salary and contract income. The model owner compares score and referral movements, while the credit policy owner examines affordability decisions. This is not a harmless data-cleaning edit merely because the new classification is more accurate.

The bank can deploy a new classifier version behind a feature version and pin current models until validation is complete. A shadow calculation shows both values on live applications without changing the decision. If the new value moves a case across a threshold, an analyst checks the original payment evidence and whether the proposed rule matches the approved income definition. The release record lists source code, mapping version, feature version, model version and policy version. A sampled decision after release should have the new combination; one before release should retain the old value and configuration. A correction to an old payment remains a separate restatement event.

Reproducibility under external data

Bureau, market and third-party identity feeds can change without the bank changing its own code. Store the response or a permitted immutable reference with request time, provider version, status and legal-use conditions. A later bureau file can include new accounts or corrected balances; it should not replace the response used at application. For a market-rate feature, record the publication vintage and the bank's receipt time. A value effective for Monday but published Tuesday is not available for a Monday decision. Tests should include a delayed update and a provider outage.

When external data cannot be retained verbatim because of contract or law, preserve sufficient governed evidence to explain the feature and arrange a controlled replay mechanism within those constraints. An unavailable third-party response is different from a response that explicitly reports no match. The fallback rule must preserve this distinction. If a provider silently changes a field meaning, contract monitoring, distribution checks and sampled reconciliation should detect the shift before models absorb it as a customer behavior change. A version manifest should include external dependencies as well as local code.

Model comparison across feature versions

An incumbent and challenger model should be compared using the feature versions each was designed to consume, on a common eligible cohort and decision clock. Running an old model with new inputs without testing compatibility can make it appear better or worse for the wrong reason. Separate the effects of a new feature definition, a newly trained model and a changed policy threshold. A controlled experiment or offline comparison can hold two dimensions fixed while varying the third. Report population coverage, missingness, scores, final actions and mature outcomes for each candidate path.

If a model was trained on a backfilled dataset with corrected customer links, assess whether those corrections are available in the live service. A challenger can perform well on reconstructed history and fail on live incomplete data. A shadow release measures actual coverage, latency and output distribution without changing customer actions. Validation should inspect disagreements by product and relevant segment, especially where a definition change moves many records from missing to populated. Production approval states the exact feature versions and accepted fallbacks rather than a vague statement that the model works with the feature store.

Governance of deprecation

When a feature version is retired, list its live consumers and confirm migration or deliberate removal. Keep past values and definitions for the required audit period. An old training dataset may still refer to the retired version, so its catalogue entry remains interpretable even after the online endpoint is disabled. If a model continues consuming it under an exception, set an owner and sunset review. The bank should avoid automatically upgrading models to a new major feature version just because a shared service changed its default.

A release review should sign off semantic tests, backward compatibility, access rules, performance and monitoring. After release, inspect value distributions, missingness, model outputs, actions and complaints against the expected change. If results differ, restrict use or roll back using the tested combination. Reproducibility means a bank can explain and rebuild the evidence for an actual decision across time; versioning supplies the controlled map of the dependencies that made that decision possible.

For a final audit exercise, select one accepted loan, one referred loan and one fraud hold from before and after a feature change. Have an independent reviewer locate the original input response, identify the source and definition versions, and calculate the feature from archived events available at that time. Compare the reproduced score and policy action with the recorded customer outcome. Where a correction was later applied, show both views and the assessed customer effect. Record failures explicitly: missing source evidence, incompatible model versions or an ambiguous cutoff reveal gaps that a release test alone might not catch. This exercise makes reproducibility an observable property of the decision process.

The reviewer should also test a deliberately failed request. A timeout, missing source and incompatible feature version must have distinct recorded statuses and policy fallbacks. Replaying only successful scores can hide a period when automation silently bypassed a control. Reconcile the number of eligible decisions, successful scores, referrals and failures for the release window. If records cannot be joined, note the uncertain scope and repair the logging path before relying on it for incident analysis. The audit is complete when the bank can reconstruct both ordinary and degraded decisions from their original evidence.

Banking practice note on definition clarity

For feature versioning and reproducibility, definition clarity matters because the bank is not simply creating a useful variable. It is creating a signal that may influence people, portfolios, controls, workloads, provisions, risk reports or regulatory evidence. A feature that looks harmless inside a model table can become very material once it is reused across decisions.

The practical approach is to separate observed fact, derived feature, model interpretation and business action. Observed facts come from systems such as feature definitions, calculation code, source data versions, business calendars, and time windows. Derived features apply definitions and time windows. The model interprets those features within its approved purpose. The business action decides what happens next. Keeping these layers separate makes the result easier to challenge and easier to explain.

This is also why the feature owner should maintain definition, lineage, version, quality threshold, permitted-use tag, monitoring rule and retirement rule. Without those basics, the bank may save time during build and lose far more time during validation, incident review, audit or customer challenge.

Banking practice note on source ownership

For feature versioning and reproducibility, source ownership matters because the bank is not simply creating a useful variable. It is creating a signal that may influence people, portfolios, controls, workloads, provisions, risk reports or regulatory evidence. A feature that looks harmless inside a model table can become very material once it is reused across decisions.

Banking practice note on quality monitoring

For feature versioning and reproducibility, quality monitoring matters because the bank is not simply creating a useful variable. It is creating a signal that may influence people, portfolios, controls, workloads, provisions, risk reports or regulatory evidence. A feature that looks harmless inside a model table can become very material once it is reused across decisions.

Banking practice note on permitted use

For feature versioning and reproducibility, permitted use matters because the bank is not simply creating a useful variable. It is creating a signal that may influence people, portfolios, controls, workloads, provisions, risk reports or regulatory evidence. A feature that looks harmless inside a model table can become very material once it is reused across decisions.

Banking practice note on model validation

For feature versioning and reproducibility, model validation matters because the bank is not simply creating a useful variable. It is creating a signal that may influence people, portfolios, controls, workloads, provisions, risk reports or regulatory evidence. A feature that looks harmless inside a model table can become very material once it is reused across decisions.

Banking practice note on customer outcome

For feature versioning and reproducibility, customer outcome matters because the bank is not simply creating a useful variable. It is creating a signal that may influence people, portfolios, controls, workloads, provisions, risk reports or regulatory evidence. A feature that looks harmless inside a model table can become very material once it is reused across decisions.

Banking practice note on audit trail

For feature versioning and reproducibility, audit trail matters because the bank is not simply creating a useful variable. It is creating a signal that may influence people, portfolios, controls, workloads, provisions, risk reports or regulatory evidence. A feature that looks harmless inside a model table can become very material once it is reused across decisions.

Banking practice note on operational fallback

For feature versioning and reproducibility, operational fallback matters because the bank is not simply creating a useful variable. It is creating a signal that may influence people, portfolios, controls, workloads, provisions, risk reports or regulatory evidence. A feature that looks harmless inside a model table can become very material once it is reused across decisions.

Banking practice note on fairness review

For feature versioning and reproducibility, fairness review matters because the bank is not simply creating a useful variable. It is creating a signal that may influence people, portfolios, controls, workloads, provisions, risk reports or regulatory evidence. A feature that looks harmless inside a model table can become very material once it is reused across decisions.

Banking practice note on change control

For feature versioning and reproducibility, change control matters because the bank is not simply creating a useful variable. It is creating a signal that may influence people, portfolios, controls, workloads, provisions, risk reports or regulatory evidence. A feature that looks harmless inside a model table can become very material once it is reused across decisions.

Reproduce a decision, not only a formula

A feature version names the transformation and its dependencies: source schema, code lists, reference relationships, time window, units, missingness and eligibility. A commit hash for a computation function alone cannot reproduce a fraud score if a beneficiary table changed after the decision. Record the feature definition, source record references and versions, availability cutoff, served value, validity state, model artifact and policy version. The original business action and any later correction must remain distinguishable.

Imagine a seven-day payment count that excludes returns. Version 1 counts accepted instructions and subtracts returned ones retroactively; version 2 records attempts as known at each decision and treats later returns as outcomes. The same label "seven-day count" now has a different meaning. Give the new definition a version, backfill analytically under a new dataset ID and evaluate the model trained on each version. Do not replace version 1 in historic decision records with version 2 values.

Point-in-time replay

Select a card authorization at 09:10. Collect its source events, ingest timestamps, customer mapping valid then, feature window and online-store publication revision. Recreate the vector using only information available before the authorization deadline. A later source correction may produce a more accurate picture of what happened, but a distinct corrected reconstruction is needed. The reviewer should be able to compare the two and explain which policy action was actually taken.

Training reproducibility is broader than replaying one row. Freeze eligible population, observation cutoffs, label definition and maturity, source manifests, transformation versions, splits and exclusions. A new run with the same code can differ if a mutable table contains updated outcomes or an account crosswalk has changed. Use immutable snapshots or manifests and compare counts, balances, feature distributions and representative rows. An experiment tracking entry without underlying dataset evidence is insufficient.

Controlled change

Treat changes to feature semantics as model-impacting, even if API shape stays constant. For a currency conversion update, calculate differences by currency and amount band, then assess changes in scores, thresholds and customer actions. For a new beneficiary mapping, test fresh and migrated payees. A backward-compatible schema may still change meaning. Assign version and effective date, run parallel comparison and seek approval at the level required by the model's use.

Rollback should select a consistent combination of model, feature and policy versions. Reverting code while leaving a new reference dictionary can produce a hybrid never validated. Test that a rollback restores both serving behavior and evidence capture. For delayed event repair, restore online counters carefully to avoid double-counting; do not replay downstream payment commands. Document how long old versions can be retrieved and who approves expiry.

Reproducibility drill

Give an independent reviewer a production decision ID, then ask for exact input values and the source and transformation versions that produced them. Next ask them to rebuild a training sample linked to the deployed model and compare it with the registered dataset. Finally introduce a backdated KYC correction and ask for an impact list without overwriting either historic evidence set. Success means the reviewer can explain the original action, the correction and the current feature separately.

Replay a score after a definition changes

A bank's fraud model uses a feature counting accepted transfers to new beneficiaries in the last twenty-four hours. On Monday, the implementation identifies a beneficiary by account number. On Tuesday, a new version adds bank and country to the canonical key to avoid false matches. Both versions may be reasonable for different data conditions, but the change can alter a customer's score and referral outcome. Version the feature definition and its transformation, and retain which version the live model consumed for every decision.

At 09:00 on Monday, instruction P-88 was scored with feature version 1, value two, model version M4 and policy version R7. On Wednesday a reviewer replays it. The bank should recover the source events available by 09:00 Monday, the customer and beneficiary identity maps then active, the version 1 calculation, score and action. Running Wednesday's feature code against today's corrected source data is a new simulation, not a reproduction of Monday's decision. Store both results with distinct run IDs if an impact analysis needs them.

Versioning must cover more than source code. A model artifact can remain M4 while a reference dictionary, currency conversion, missing-value rule or online materialisation changes its inputs. An end-to-end release manifest identifies feature definitions, source schema, transformation image or commit, lookup versions, model artifact and decision policy. The bank does not have to duplicate every raw record inside the model log, but it must retain a reliable path to the evidence under its retention and access rules.

A controlled rollout and rollback

Before promotion, compute versions 1 and 2 on a dated sample. Compare feature distributions, changed model scores, policy actions, customer segments and referral queue capacity. Include a late event, duplicate instruction, corrected beneficiary key, missing reference value and new channel. A shadow run can expose differences without changing live actions. The owner approves a rollout scope and monitoring window; the operations team knows which version is live and which fallback applies if the new mapping fails.

If version 2 produces an unexpected surge of referrals, rollback to version 1 requires more than redeploying code. The online store may already contain version 2 materialised values, and a consumer may cache them. Separate feature namespaces or explicit versioned keys make rollback safer. Record the exact cutoff for affected decisions, invalidate or expire stale cached values under the approved process, and verify that model requests receive the intended version. Retain both the failed rollout and its remediation for audit.

An acceptance test should reconstruct three decisions: one before promotion, one during version 2 use, and one after rollback. Each should return the logged feature version, value, source cutoff, model score, policy action and evidence reference. Then run a counterfactual comparison using the alternative feature version, clearly marked as analysis rather than the historical fact. Reproducibility is the ability to explain what actually happened; versioning makes a controlled alternative comparison possible without confusing it with the original customer decision.

Version the source interpretation

A seven-day transfer count may change when the hub reclassifies pending payments as accepted. The function's source code can be identical while its input semantics move. A reproducible feature version therefore includes source interface and code-list versions, event status, entity key, window and missingness rule. Store the actual served value, validity and publication time alongside the score. A later source correction creates a new analytic calculation; it must not rewrite the original decision journal.

For a payment held at 10:03, an independent reviewer retrieves instruction and earlier event IDs, the reference snapshot, stream watermark and transformation version. They compute the count and compare it with the served vector. A late 10:01 event arriving at 10:05 produces a different complete-history value. Both can be correct for their respective questions. A backtest of live performance must use the earlier availability cutoff, while an incident analysis may need the corrected history too.

Keep combinations compatible

Changing the feature schema may require a new model artifact and policy threshold. A rollback that restores an old model while leaving a new beneficiary mapping can create a combination never validated. Register compatible feature, model and policy versions, then test rollback on a small set of boundary requests. Preserve dataset manifests and label definitions so an offline candidate can be rebuilt from the same eligible cohort and maturity window.

When a training row changes after a backdated account correction, retain the prior dataset version and a changed-record manifest. Compare model metrics on equivalent cohorts rather than announcing a lift from different labels or exclusions. A reproducibility drill should demonstrate the original customer action, a corrected analytic scenario and the exact source change linking them.

Rebuild after a reference correction

A beneficiary record is corrected on Friday with an effective date of Tuesday. On Wednesday the model used the earlier reference version and treated a transfer as a new payee. A training query run next month might use Friday's corrected record for Wednesday unless it reconstructs publication history. Preserve both effective and availability time, the original served feature and the corrected analytic value. The historic customer action follows the original evidence, while model evaluation can assess whether the correction would have changed it.

Build two test datasets: the original as-of decision replay and the corrected source-history analysis. Compare row counts, feature values and affected scores, naming each manifest and transformation version. A model artifact must link to the dataset version actually used for training. If a later rollback restores the model but not the reference mapping, the system may produce a hybrid never validated. Test compatibility with controlled requests and retain versions long enough for authorized audit and incident review.

Evidence retention check

Take a production decision from a prior quarter and retrieve its exact feature vector, input validity, source manifest or protected record references, transformation version, model version and policy action. Rebuild a sample under the original cutoff, then run a separate corrected reconstruction using a later reference change. The reviewer should explain both values without replacing the historic decision.

Repeat after a source migration and a model rollback. Ensure the old schema and code list remain interpretable and that the rollback selected a validated combination of components. If the bank retains only today's reference tables, a version number without reconstructable data is insufficient. Record gaps and owners before declaring the feature reproducible.

For a batch model, compare a manifest's facility IDs and balances with the source control at the original cutoff. A rerun may have the same row count but different facilities after a customer merge. Record changed IDs and source revisions as well as aggregate totals. A reproducible score depends on the same business population, not just the same executable code.

Primary sources for further study

Related learning paths

This application uses JavaScript for the full interactive experience. This text summary is served for accessibility and search indexing.

Feature versioning and reproducibility · Malla Banking Academy