What a feature means in banking

What a feature means in banking. A practical lesson in the feature store for banking and payments practitioners.

How to study this topic

A feature in banking is a controlled signal derived from banking data and used by a model or decision process, such as income stability, repayment behaviour, balance volatility, exposure trend, case history, or customer relationship depth. Study this chapter as a banking control topic first and a data science topic second. The model is not the beginning of the story. The beginning is the bank's duty to know what happened, where the data came from, whether it is complete, whether it is reconciled, whether it is allowed to be used, and whether the result can be explained later.

The banking meaning

banking feature meaning is important because banking data carries legal, financial, customer and regulatory meaning. A balance is not just a number. A balance may be available, booked, ledger, cleared, unsettled, blocked, earmarked, overdrawn or adjusted. A customer status is not just a label. It may affect onboarding, credit treatment, servicing, conduct obligations, complaints, vulnerable customer handling and regulatory reporting.

Sources and boundaries

Typical sources include raw customer facts, account behaviour, credit repayment history, collateral values, income evidence, deposit flows, channel usage, case outcomes, risk ratings, product holdings, treasury exposure, and external bureau or market data. These sources do not have equal authority. A digital channel may show customer intent, but a book-of-record platform may show final outcome. A CRM note may show relationship context, but it may not be structured enough for automated model use. A risk system may hold a rating, but the rating may have a valid-from date, expiry date, override reason or review cycle.

Why AI depends on this discipline

The practical AI use cases include application scoring, behavioural scoring, early warning, customer analytics, fraud models, complaint triage, portfolio risk, and relationship management. These can help a bank act faster and with more consistency, but only if the model input has banking quality. A model trained on weak data can still produce a neat score, rank, category or recommendation. The output may look professional while the underlying evidence is damaged.

AI does not remove the need for banking controls. It makes those controls more important because one bad input can influence many downstream decisions at speed. The safest banks treat model input preparation as part of governance, not as background plumbing.

Validation controls

Validation asks whether the data is acceptable for its intended use. For this topic, useful controls include definition ownership, calculation logic, permitted use, refresh frequency, and point-in-time correctness. These controls should run before data becomes a model feature, dashboard metric, risk indicator or operational priority. The controls must check both technical shape and banking meaning.

Validation should produce actionable exceptions. It should say what failed, why it matters, which downstream consumers are affected, whether the model must stop, whether degraded use is allowed, and who owns correction. A generic red status is not enough for production banking.

Operational design

Operational resilience also means fallback. If the feature feed is unavailable, should the bank stop the model, use the last good value, route to manual review, switch to a rule-based control, or degrade the service? The answer must be designed before failure.

Governance and audit

BCBS 239, model-risk guidance and AI-risk frameworks all point in the same practical direction: banks need controlled data, clear limitations, documented governance and monitoring. The details differ by jurisdiction and model type, but the core discipline is stable.

Common mistakes

The final mistake is rushing feature reuse. Centralised features are powerful, but reuse without purpose control can create hidden risk. The same feature may be safe for portfolio monitoring and unsafe for direct customer decisioning. A bank must know the difference.

Bank-ready checklist

If these questions are answered well, AI adoption becomes much safer. The bank is not merely feeding data into a model. It is turning banking evidence into controlled decision support.

Source anchors for further study

BCBS 239 supports the banking discipline of accurate, complete, timely and adaptable risk data aggregation and reporting.

Model-risk guidance such as the Federal Reserve supervisory guidance emphasises input quality, data constraints, limitations, validation, monitoring, documentation and governance.

NIST AI RMF is useful as a general AI risk framing for governance, mapping, measuring and managing AI risk, but banking still needs bank-specific controls around customers, products, books of record, risk and regulatory evidence.

Practical banking note on business meaning

For banking feature meaning, the business meaning point matters because banking AI does not live in isolation. It sits beside customer journeys, product ledgers, risk controls, finance numbers, operations queues, compliance obligations, audit evidence and management decisions. A weak data point can move quietly through all of these areas unless the bank designs controls that stop, flag or limit it.

The practical discipline is to keep four layers separate: observed banking fact, derived data signal, model interpretation and business action. Observed facts come from systems such as raw customer facts, account behaviour, credit repayment history, collateral values, and income evidence. Derived signals reshape those facts into features or indicators. Model interpretation produces a score, classification, ranking or recommendation. Business action decides what the bank will actually do. When these layers are blurred, nobody can explain the result properly.

Practical banking note on lineage

For banking feature meaning, the lineage point matters because banking AI does not live in isolation. It sits beside customer journeys, product ledgers, risk controls, finance numbers, operations queues, compliance obligations, audit evidence and management decisions. A weak data point can move quietly through all of these areas unless the bank designs controls that stop, flag or limit it.

Practical banking note on customer impact

For banking feature meaning, the customer impact point matters because banking AI does not live in isolation. It sits beside customer journeys, product ledgers, risk controls, finance numbers, operations queues, compliance obligations, audit evidence and management decisions. A weak data point can move quietly through all of these areas unless the bank designs controls that stop, flag or limit it.

Practical banking note on risk control

For banking feature meaning, the risk control point matters because banking AI does not live in isolation. It sits beside customer journeys, product ledgers, risk controls, finance numbers, operations queues, compliance obligations, audit evidence and management decisions. A weak data point can move quietly through all of these areas unless the bank designs controls that stop, flag or limit it.

Practical banking note on model limitation

For banking feature meaning, the model limitation point matters because banking AI does not live in isolation. It sits beside customer journeys, product ledgers, risk controls, finance numbers, operations queues, compliance obligations, audit evidence and management decisions. A weak data point can move quietly through all of these areas unless the bank designs controls that stop, flag or limit it.

Practical banking note on operational fallback

For banking feature meaning, the operational fallback point matters because banking AI does not live in isolation. It sits beside customer journeys, product ledgers, risk controls, finance numbers, operations queues, compliance obligations, audit evidence and management decisions. A weak data point can move quietly through all of these areas unless the bank designs controls that stop, flag or limit it.

Practical banking note on audit evidence

For banking feature meaning, the audit evidence point matters because banking AI does not live in isolation. It sits beside customer journeys, product ledgers, risk controls, finance numbers, operations queues, compliance obligations, audit evidence and management decisions. A weak data point can move quietly through all of these areas unless the bank designs controls that stop, flag or limit it.

Practical banking note on privacy and permitted use

For banking feature meaning, the privacy and permitted use point matters because banking AI does not live in isolation. It sits beside customer journeys, product ledgers, risk controls, finance numbers, operations queues, compliance obligations, audit evidence and management decisions. A weak data point can move quietly through all of these areas unless the bank designs controls that stop, flag or limit it.

Practical banking note on feature reuse

For banking feature meaning, the feature reuse point matters because banking AI does not live in isolation. It sits beside customer journeys, product ledgers, risk controls, finance numbers, operations queues, compliance obligations, audit evidence and management decisions. A weak data point can move quietly through all of these areas unless the bank designs controls that stop, flag or limit it.

Practical banking note on monitoring

For banking feature meaning, the monitoring point matters because banking AI does not live in isolation. It sits beside customer journeys, product ledgers, risk controls, finance numbers, operations queues, compliance obligations, audit evidence and management decisions. A weak data point can move quietly through all of these areas unless the bank designs controls that stop, flag or limit it.

From source fact to decision input

A feature is a defined value presented to a model at a specified decision time. It is not simply a column copied from a database. For a credit application, prior six-month missed-payment count needs a population, account relationship, observation window, definition of missed, treatment of reversals and a cutoff that precedes the decision. For a payment fraud score, new-beneficiary indicator needs a rule for beneficiary identity, first-seen time, channel and whether a repaired instruction changes the value. Two services can both expose a field named risk_score while referring to different models and outcomes. A feature contract states exactly what is measured and what action it supports.

Separate source observation, derived feature, model estimate and policy action. A ledger posting is an observation. A count of qualifying late repayments is derived from postings and a repayment schedule. A probability of default is a model estimate with a horizon and calibration. A refer decision is a policy action that may also use affordability and legal checks. If a training table stores all four without names or lineage, a later analyst may use a prior decision as an input and create a self-reinforcing model. The feature catalogue should classify each field and prohibit future outcomes from entering an earlier score.

A precise feature contract

For each feature record title, business purpose, model uses, owner, source systems, grain, entity key, unit, allowed values, calculation, null reason, observation period, event-time rule, availability cutoff, refresh cadence, maximum age, privacy classification and fallback. Store definition and code versions. Include a normal example and difficult cases. A customer may have two accounts, one loan transferred to a new platform, a payment reversal and a master-data merge. The contract should say how those cases affect the value. A technical schema can check numeric type; it cannot decide whether the feature represents this borrower at this decision time.

The entity key is critical. A customer ID may change on merge, an account can be shared, and a corporate customer can have nested legal entities. A model scoring one applicant must not accidentally aggregate a spouse, another company in a group, or a same-name person unless the approved purpose and relationship definition require it. Keep relationship versions with effective and availability time. Test joins by tracing raw evidence to a sampled feature value and confirming the model's population. An ambiguous identity should become a controlled exception, not a confident numeric feature.

Point-in-time behavior

Suppose a loan is scored at 10:00. A salary credit has an effective date of yesterday but arrives at 10:05. It cannot be part of the 10:00 feature snapshot. A transaction posted at 09:50 but only ingested into the model store at 10:02 was likewise unavailable. Historical training must use the same availability rule as production, not only a business date. Record event time, source publication time where relevant, ingestion time, feature calculation time and decision time. Corrections received later should remain linked but not overwrite the earlier input used for audit.

The observation window must specify whether it uses calendar days, business days, complete months or rolling hours. A six-month repayment feature calculated on 1 July may cover January through June, while a rolling 180-day window begins in early January on a different date. The choice can change values at boundaries. Test an event exactly at the cutoff and one just after it. For a real-time fraud feature, decide whether the current payment is excluded from prior velocity and how a retry is counted. A one-minute clock error or duplicate publication can have a material effect when the threshold is close.

Online and offline parity

An offline warehouse can calculate a feature across years of history. An online store serves it at decision time with a latency and freshness requirement. They need the same semantic definition even if their engineering implementations differ. Sample live decisions and reconstruct each input from frozen source and code versions. Compare value, null state, age and entity key. A numerical match is not enough if the online service used a different customer ID that happened to produce the same count. Keep the actual online snapshot with the score; a later recomputation from cleaned data is not proof of what was used.

If parity fails, classify the reason: source latency, transformation difference, identity change, backfill, rounding or code-version drift. Fix the responsible layer and assess affected decisions. A daily credit feature may tolerate a longer update interval than a pre-release fraud feature, but that tolerance must be approved by use. A feature store outage should trigger a documented referral or fallback, not a silent zero. A model validated on complete offline features may be unsafe if the online path frequently serves stale values.

Reuse with boundaries

Centralising features can save effort and improve consistency, yet one definition may not fit every decision. Prior payment count for fraud might include submitted and held orders to capture attempted velocity. Prior payment count for reconciliation workload might include only released messages. A credit repayment feature counts obligations under a different legal and business relationship. Label these features separately and explain their grain and stage. Reuse the source event and common identity service where appropriate, but do not force different questions into one ambiguous number.

Approval for a new model use should check purpose, source rights, customer impact and validation population. A feature that was acceptable for portfolio monitoring may be too stale or too broad for an individual credit decision. A sanctions case flag can be sensitive and should not become a general-purpose customer-risk feature without compliance review. The feature catalogue should list consumers and alert them when a source or calculation changes. A version change can require a model retest even if the API field name stays the same.

A worked feature example

Define prior three completed months of qualifying salary credits for a fictional loan applicant. The source is the applicant's owned current accounts as linked by the customer master at the application time. A qualifying credit follows an approved payroll classification, excludes internal transfers and reversals, and is counted by posting date when it was available to the model. The value is a count from zero to three plus a missingness reason if account history is incomplete. The calculation stores source transaction IDs, account relationship version, classification version and cutoff. It does not infer that every regular credit is salary or that three credits establish affordability.

Test a payroll credit reversed the next day, two salaries in one month, a joint account, a newly opened account and a late transaction. The expected value for each case is documented before implementation. If the bank learns after a decision that a credit was misclassified, it can investigate and correct the current feature while retaining the original snapshot. The model's score and policy action remain separate. This example turns the abstract word feature into a reproducible banking claim that a validator can challenge and an analyst can explain.

Release and monitoring

The feature owner signs the definition; source owners attest to data meaning; model owners validate use; product and control owners approve the action and fallback. Release evidence includes tests for calculation, point-in-time cutoff, identity joins, online/offline parity, access controls and outage behavior. Monitor missingness, freshness, distribution and unexpected category changes by product and channel. A sudden stable zero can indicate a failed source rather than lower risk. Investigate model outcome changes alongside feature health and decision policy. The feature is fit for production when the bank can replay its value and explain why it was available, lawful and relevant at the moment it influenced an action.

Another feature, another clock

Consider a payment-fraud feature called new beneficiary in the prior 24 hours. The bank must define beneficiary identity from the actual submitted account and party data, not only a saved-template label. A customer might edit a template, use two aliases for the same account or send through a channel with no template concept. The first-seen timestamp comes from a controlled event at the chosen boundary, such as an accepted customer instruction. A held payment can establish that the beneficiary was attempted, even if no external message was sent. The next fraud decision must know whether that attempt counts. The feature should retain the chosen identity rule, event set, window and update time.

Suppose the first attempt at 09:00 is held, then the customer retries at 09:03 after a browser timeout. The second submission may be a duplicate transport attempt, not a new behavioral choice. The idempotency relationship should prevent it from inflating the count. At 09:10, an analyst changes the beneficiary account after verifying a source error. A later payment to the corrected account could be first-seen under the approved rule even though the original attempt was older. Model and policy owners must decide which state matters for fraud risk and test the correction path. A generic SQL distinct count cannot settle that business choice.

This feature also raises access and privacy questions. Raw beneficiary names and account numbers may be restricted, while a derived first-seen indicator can be served to a specific fraud model. The mapping still needs a secure way to reconstruct source evidence in an investigation. If the identifier is hashed, key rotation and cross-channel consistency need testing. If a customer disputes a hold, operations should be able to explain the actual event pattern without exposing confidential model weights or another customer's data. The feature contract bridges those technical and conduct boundaries.

Feature importance does not prove correctness

A model explanation may rank a feature highly. That tells the reviewer something about the fitted model's behavior under an explanation method; it does not prove the feature's source was accurate, its use was lawful or the decision fair. A wrongly merged customer history can be highly predictive in a development sample because the same bad join appears in training and test. A validation team should challenge feature provenance, stability, proxies and sensitivity to missing data. It should compare model performance with and without a material feature, examine the effect by relevant population, and review whether its business meaning is plausible.

For a credit feature, check whether a proxy for geography or account tenure creates unacceptable disparities under applicable law and policy. For a fraud feature, check whether a new device indicator flags customers whose access circumstances changed for legitimate reasons. A business owner may retain a feature with a controlled human-review action while rejecting its use for an automatic decline. That is a model-use decision, not a property encoded in a feature table. The record should show the rationale, tested population and monitoring needed after release.

Change control example

The bank updates its payroll classifier to recognise a new employer payment code. The salary-credit count changes for a subset of applicants. Before promotion, compare old and new values on a fixed historical sample and on current applications, including customers previously referred for missing income evidence. The model owner assesses score changes and whether validation or threshold approval is needed. The customer process owner assesses any explanation or documentation change. Release both classifier and feature versions with an effective date; do not overwrite past snapshots. A rollback must restore compatible calculation and model inputs, not just switch a code repository tag.

The release test should include a salary from the new code, a transfer that resembles payroll but is not, a reversal and a customer with no account history. Compare expected counts and downstream actions. If a model uses the feature in both onboarding and portfolio monitoring, approve those uses separately because their decision clocks and customer effects differ. Keep a dependency list so a later classifier fix reaches all consumers. A well-managed feature has a meaningful definition, a tested calculation, an approved use and an evidence trail through every version.

Definition before computation

A feature is a measured fact prepared for a specific decision, observation time and eligible entity. "Number of payments" is incomplete: does it count initiated, accepted or settled instructions; inbound or outbound; retries; reversals; and which clock bounds the window? For a fraud decision at 10:03, define a distinct outbound instruction count over the preceding hour using business IDs, accepted status, customer account, and events available before the score deadline. The definition determines both its meaning and whether the number can be reproduced.

Write a feature contract with entity key, source events, event-time window, ingestion cutoff, aggregation, missingness semantics, units, version, permitted population and owner. A zero count means the bank observed eligible activity and found none. An unavailable feed means it does not know the count. A separate validity indicator prevents the model from interpreting failure as ordinary inactivity. A branch channel that never supplies a device ID differs from a mobile event where the normally present ID has vanished. The model's training and live serving should apply the same distinction.

A hand calculation

Suppose account A initiates payments at 09:15, 09:45 and 10:01. The 09:45 message is retried twice under the same instruction ID, and the 10:01 event does not arrive until 10:05. At the 10:03 decision, the one-hour observed distinct-instruction count is two, assuming both earlier instructions reached the feature service. A corrected historical event-time count can later be three. Record the original two and the data watermark; never replace it in the decision journal with three. Test that retries do not inflate the count and that the late event changes only later eligible scores.

A transaction's amount feature might be copied from the instruction, converted to a reporting currency or expressed relative to the customer's recent spend. Each choice serves a different question. For a foreign-currency payment, retain original amount, currency, conversion source and rate time. A rate from tomorrow would introduce information unavailable to the live decision. Relative amount requires a baseline with its own observation window and validity. If history is insufficient, report that condition explicitly rather than dividing by a fabricated average.

Entity and label boundaries

An account feature, customer feature, household feature and beneficiary feature can show different activity. A customer with two accounts might make five payments in total but only two from the account being scored. A household link can be ambiguous or privacy-restricted. Choose the entity appropriate to the risk hypothesis and policy permission, then test merges, splits and reference corrections. An identity change should be versioned; a current master record is not automatically the correct historical key.

A feature is not the outcome label. "Prior confirmed fraud" can be a feature if confirmation existed before the decision, but a later case disposition cannot be backfilled into that decision's input. A blocked transaction's lack of settled loss reflects the bank's action, not proof that the attempted payment was safe. Document the difference between observation, action and mature outcome when constructing training rows. Test one example where a case changed status after the score.

From analysis to production

An analyst may calculate a feature in a warehouse using complete history, while a production stream works with partial arrivals. Compare values at the same cutoff across representative entities and failure cases. Use the source's availability time, not merely event time, for a historical simulation of what the bank could know. Compare missingness and distributions by channel, segment and model version. A numerical match on normal examples is insufficient if the online path treats reversals or timeouts differently.

Feature importance does not establish a causal relationship or permission to use a field. A strong predictive proxy may reflect prior policy or unequal source coverage. Examine purpose, access, fairness implications and operational reliability before release. A feature that cannot be explained or produced within the decision deadline should not be silently substituted. The approval pack should state what the value means, what happens when it fails and how changes are controlled.

Reviewer exercise

Give a reviewer the source events, timestamps and IDs for one transfer, its actual feature vector and the final banking action. Ask them to calculate each critical feature and identify missing inputs. Repeat after a channel migration that changes status codes and a reference update that merges customer IDs. The test passes when the original vector remains explainable and the new mapping produces an explicitly versioned result. A useful banking feature is an auditable measurement of available evidence at a named decision, not just a numeric column correlated with an outcome.

One beneficiary flag, several possible meanings

Suppose a bank proposes a feature named new_beneficiary for pre-release payment fraud scoring. That name is not a specification. Does new mean the first payment to an account, first creation of a saved payee, first payment by this customer, or first payment by any account in a corporate group? Does the comparison include failed attempts? How are account identifiers normalised across local and cross-border rails? An analyst should write a feature contract before a developer implements a boolean.

For an illustrative retail transfer at 10:05, define new_beneficiary as true when the resolved customer had no previously accepted payment instruction to the same canonical beneficiary account in the prior ninety days, using events available before the current score. The current instruction is excluded. Store the customer-resolution version, beneficiary-key rule, cutoff, observation window and source-event IDs. A payment initiated earlier but only published after 10:05 cannot appear in the original score, although a later analytical rebuild may know about it. The feature value needs an availability state as well as true or false: source unavailable is not the same as a confirmed lack of history.

Now test four cases. A customer sends two instructions to the same payee within seconds after a channel timeout. If the second is a transport retry of the first business instruction, it should not become historical evidence of an earlier beneficiary. If it is a genuinely new instruction accepted after the first, the value may differ, depending on when the first acceptance reached the feature service. A beneficiary saved in the address book but never paid should remain new under the payment-history definition. A payee account whose display name changed should remain the same beneficiary if its canonical account key has not changed. Each result follows from the written contract rather than a model developer's intuition.

The source of the entity link matters. A joint account can be associated with two customers, while a corporate parent can have several subsidiaries. The bank must choose whether history belongs to the initiating legal entity, account, user or group, and should not silently apply a retail identity rule to a corporate payment. A customer merge performed next week can improve the current master record but must not make last week's logged feature value appear to have used the merged identity. Version the relationship and retain the decision-time key.

Use the feature in a controlled action

The model may combine this flag with amount, device, session, recent accepted-payment velocity and other permitted inputs. The flag alone does not establish fraud. A policy could refer a high-risk score for additional verification or investigation, while a low-risk score proceeds under the ordinary controls. The decision journal should retain the feature definition version, actual value, missingness, model version, score and policy outcome. If a customer asks why a transfer was delayed, the bank's explanation needs an approved, accurate reason rather than a raw feature label that may expose internal security logic.

Assess the feature before promotion. Compare its distribution by channel, customer segment and product. Test whether a new digital channel produces more missing beneficiary history because its keys do not match the legacy directory. Check precision and workload only with mature, intervention-aware outcomes; a held transfer has no observed released-payment loss. Measure whether the feature adds useful signal beyond existing rules, and whether it creates unnecessary friction for customers who routinely pay new recipients. The business owner decides acceptable trade-offs and fallback when the history service is stale.

A change to the ninety-day window is a new feature version, not a harmless parameter tweak. Recompute a dated sample under both definitions, compare changed decisions and examine peak times and customer segments. Record who approved the change and how the bank will monitor it in production. This worked flag shows what a banking feature really is: a dated, permissioned measurement of an explicitly defined population at a decision boundary, with enough evidence to reproduce and challenge the resulting action.

Interpret the number for a real decision

An applicant's "monthly income" feature could mean declared salary, verified salary, qualifying credits or a mixture. The model cannot choose the meaning from the column name. For an application at 10:00, identify the approved source and observation window, excluding internal transfers and reversals under a documented rule. Keep the source evidence and a validity state. A customer with no accounts in the bank is not a customer with a verified income of zero. A self-employed applicant whose deposits vary across months may require a different assessment path; a three-month average alone is not a universal income measure.

Hand-calculate a simple example. In January, February and March, the account receives 3,000 salary credits each month. In March, it also receives a 5,000 transfer from the customer's own savings account and a 500 credit later reversed. At an April application, qualifying salary could be 3,000 per complete month if classification and ownership evidence are valid. An unguarded sum of all incoming credits would materially overstate the figure. Record the transaction IDs and dates that support the 3,000, and test what happens if the payroll classification is absent. The feature response should convey uncertainty rather than silently replacing it with a precise but unsupported value.

Define a behavioral window

For card fraud, "number of transactions in the last hour" needs a reference clock and status rule. At 10:03, include distinct accepted authorizations that reached the feature service before its deadline. Exclude retries with the same business ID and define whether the current authorization is included. A delayed 10:01 event reaching the bank at 10:05 cannot have informed the 10:03 action. The original counter and later corrected analytic counter are both useful but answer different questions. A reviewer needs to see which one the model actually used.

Two same-amount purchases at one merchant can be distinct transactions, while a single authorization can have multiple status messages. Test both. If merchant category changes under a new network code list, version the mapping; a broad "other" bucket may alter fraud patterns. A numeric value without event semantics can pass type and range checks while measuring the wrong activity.

Granularity and reference versions

A borrower can own two accounts and a business account, and can be linked to a household. An account-level outflow count differs from a customer-level total. An identity service may merge records after the decision; replay must use the relationship version available then. A source correction may identify an earlier wrong link and trigger impact review, but should not overwrite original score evidence. List entity key, relationship validity, join cardinality and any unresolved conflicts in the feature contract.

For a beneficiary age feature, choose creation, first use or verification time. A destination created a week ago but first paid today has different ages under each definition. Select one based on the fraud hypothesis and the allowed action. Test a payee migrated from a legacy system that received a new technical ID; treating it as newly created can flood a review queue. An explicitly unknown age may be more honest than a fabricated zero.

Evidence, use and change

The catalog entry for a feature states purpose, source owner, transformation owner, unit, window, freshness limit, missingness meaning, permitted consumers and version. An investigator can use it to reconstruct a customer action, and an engineer can test online and offline parity. A credit product may permit a weekly repayment history while a live payment needs a seconds-old velocity count. One infrastructure platform can host both; their contracts cannot be reduced to the same service-availability metric.

Before promoting a new feature, compare old and new score vectors in shadow on the target population. Inspect exclusions, missingness by channel, source rights, potential proxies and business outcomes. An offline correlation with a label is a starting hypothesis, not permission to use a field or evidence of improved decisions. When the source changes, identify every model consuming the feature and evaluate action impact. The feature is a named, time-bounded measurement with a controlled interpretation.

Feature approval exercise

Propose a feature called recent account activity for a loan application. The analyst first asks whether it counts transfers, card payments, deposits or all postings, and whether the customer owns every included account. The data team prepares two versions on the same 100 applications: one counts raw messages, the other distinct posted transactions after retries and reversals. Compare feature values, missingness and score bands, then inspect the ten largest differences with source evidence. A version that looks more predictive in a backtest may have used later settlement statuses; replay it at application cutoffs before drawing a conclusion.

Approval names the chosen definition, its permissible purpose, source owner, transformation version, freshness rule and response when identity resolution is ambiguous. Check whether limited-history applicants follow a validated route rather than receiving a fabricated zero. Record the model and policy versions that consume the feature. On a later interface change, a contract test with ordinary, duplicate, returned and delayed transactions should fail if meaning changed. A good feature catalog supports this exact challenge, not just discovery by column name.

Primary sources for further study

Related learning paths

This application uses JavaScript for the full interactive experience. This text summary is served for accessibility and search indexing.

What a feature means in banking · Malla Banking Academy