Market data and external risk signals

Market data and external risk signals. A practical lesson in the data foundation for banking and payments practitioners.

The model question determines the external data

External data is useful in banking AI only when it changes a defined decision or forecast. An interest-rate curve may help a treasury model estimate repricing exposure. A dated unemployment series may help a credit-risk team test how expected losses behave under different economic conditions. A bureau response may supply an applicant's observed obligations at a lending decision. These sources do not form one generic risk feed. They have different populations, publication clocks, rights of use and errors. Combining them in a feature store without a business contract can give a model impressive correlations that disappear in production.

Begin with the decision point. Who uses the output? Is the model predicting a payment repair, estimating a borrower outcome, forecasting liquidity or detecting an unusual market move? When must it act, and what information could the bank actually have known then? A daily central-bank reference rate published later in the day cannot justify an earlier customer quote. A revised unemployment estimate cannot be used as a feature in a back-test that claims to reproduce decisions made before the revision. An external source is not automatically authoritative for every purpose simply because a respected institution publishes it.

This lesson focuses on how external signals become controlled ML inputs. It does not teach trading strategies or imply that market data predicts individual customer behavior. Each example is illustrative, with local legal, contractual and model-governance review required for a real bank. ECB Data Portal documents its statistical API and historical versions; the Federal Reserve Bank of St. Louis FRED API and its vintage-date endpoint show why a data value and the date it was available must be treated separately.

Classify the source before engineering a feature

Market observations include reference rates, FX rates, instrument prices, spreads, curves and volatility measures. Macro series include inflation, employment, output and property indicators. Credit bureaus and approved third parties provide borrower-level information subject to distinct legal and contractual conditions. Sanctions lists, company registers and adverse-media services have compliance and identity uses; they should not be casually treated as interchangeable numerical credit features. Internal bank prices and positions are not external signals merely because they are stored in a vendor platform. The source catalogue should record the producer, instrument or population, geography, unit, frequency, time zone, publication lag, revision practice, licence and permitted model use.

A price is incomplete without its market convention. A rate may be a mid, bid, ask, fixing, reference or executed customer rate. A curve point needs a currency, tenor, day-count and observation time. A spread needs a defined benchmark and instrument population. An FX quotation needs a base and quote currency: EUR/USD and USD/EUR are related mathematically but not the same stored value. If a data engineer silently inverts a rate or converts a percentage into a decimal inconsistently, every downstream feature can look plausible while being wrong. Unit and convention checks belong in ingestion, not in a model analyst's memory.

The ECB's exchange-rate methodology explains that its reference rates are published on working days on a schedule. A bank should not use such a reference series as if it were a continuously executable customer price. A liquidity forecast might use the rate as a contextual daily variable; an actual payment's booked FX rate comes from the bank's trade or pricing evidence. The feature definition must say which one it uses and why. Mixing reference, quoted and executed rates in one column creates false volatility and can distort both training and operational explanations.

Four clocks for every observation

Record the observation period, the source publication time, the bank ingestion time and the model decision time. They answer different questions. Monthly unemployment for April describes a period, but its first public release may occur later. A bank may ingest it several hours after publication. An online model at 09:00 can only use the version available to its service by 09:00. A later revision may change the reported value for April; that revised number belongs to a later vintage. Back-testing on today's final series as if it were available at every historical decision creates look-ahead bias.

The data contract should store value, unit, period, publication timestamp, source revision or vintage, ingestion timestamp, effective or valid date and quality state. The same series may have a first estimate, a scheduled revision and a methodology break. Use the version that was available at each historical decision when training or replaying. Keep later values for current reporting and analysis, but do not overwrite the earlier snapshot. The ECB API's history option and FRED vintage-date functionality are examples of source capabilities that can help construct this evidence. The bank still needs to test how its own feed captured those versions at the time.

For a daily FX series, distinguish a market holiday, delayed publication, a source outage and a stale cache. A missing value is not zero. A forward-fill policy may be acceptable for a specific forecast with a documented maximum age; it is dangerous for a quote or a risk limit that requires current data. At a weekend decision, the last published reference may be several calendar days old. The model must know that age, and the business owner must decide whether to use it, restrict the decision or fall back to another approved method.

Joining a macro signal to a borrower

Suppose a bank estimates the risk of arrears for a portfolio of mortgages. It considers regional unemployment as a contextual variable. The borrower record has an address, but that address may change, be missing or represent a mailing location rather than the property or employment region. The external series has its own regional boundaries and release dates. A join on a current postcode can incorrectly assign today's geography to a historical loan decision. The feature contract must define which geography matters, what source establishes it at the observation date, how boundaries are mapped and when the series was available.

Do not infer causality from a correlation between regional unemployment and arrears. Loan origination policy, product mix, repayment holidays, interest-rate changes and collections actions can move together. A model may use the macro variable for a limited forecast or scenario analysis, but it needs out-of-time testing, segment review and clear limitations. A bank should compare it with a simpler baseline and test performance under periods that differ from the development sample. A highly predictive regional variable can also act as a proxy for sensitive or protected characteristics in some settings, requiring fairness and legal review before individual-level use.

For expected credit loss, forward-looking economic information may be relevant under the applicable accounting framework, but accounting measurement is not the same as a single borrower score. Scenario selection, weighting, staging, governance and overlays require their own methods and approval. An ML model can estimate components or challenge assumptions; finance owns the financial statement process. The model lineage should show how the economic series and scenarios entered the estimate, with effective dates and versions. A regulatory or accounting reference must be scoped to the bank's jurisdiction and product rather than copied from a generic source list.

External credit data is not just another market series

A credit bureau response contains borrower-level observations and sometimes scores, inquiries or derived attributes. Its purpose, consent or lawful basis, retention and dispute mechanisms differ from public macro data. The bank must keep the request and response identifiers, provider timestamp, applicant match confidence, report version, fields used and any adverse-action implications under local law. A missing bureau record can mean thin history, identity mismatch, service outage or a jurisdiction where coverage differs. Treating all four as a low-risk score or a zero obligation is a model and conduct error.

Before training, define the population that was actually queried and the reason for selection. A bank may request bureau data only for applicants who reached a particular stage, creating selection bias. A model trained on accepted loans alone cannot assume that declined applicants would have behaved the same way. Keep the acquisition and decision policies in the dataset, and evaluate the model against the population where it will be used. If a provider corrects a dispute after the decision, retain the original response for replay and the correction for remediation. The two records answer different questions.

A liquidity forecast using market context

A treasury team forecasts end-of-day cash needs across currencies. Internal payment queues, opening balances, expected settlements and correspondent account movements are the primary operational evidence. FX and rate context may improve a forecast or explain movement, but an external series cannot replace the bank's own cash and payment events. The forecast horizon matters: a next-hour liquidity alert, an end-of-day position and a monthly planning forecast need different refresh cycles and error measures. Define the action first, such as whether treasury should investigate a projected shortfall or adjust funding within its approved controls.

Suppose the model uses a daily reference rate and intraday market quotes. Their timestamps and conventions differ. The feature store should retain each separately, with a flag for stale or missing observations. An analyst may calculate a currency-normalized amount for comparison, but must keep the original amount and rate source. If a payment's actual conversion was executed at another rate, reconciliation and P&L need that executed evidence. A model's contextual FX feature does not change the booked transaction or authorise a trade. Back-testing should align the forecast cut-off with the data actually available at that time, including payments that arrived late and market quotes that were revised or cancelled.

Evaluate the forecast at the level where a decision is made: currency, account, settlement window or legal entity. A low aggregate error can hide a material shortfall in a small currency. Record directional bias, tail errors and occasions when a fallback forecast was used because the market feed failed. The operational response should name an owner and required approval. If market volatility rises, the model may become less reliable precisely when treasury most needs it; stress tests should examine that condition rather than relying on average historical days.

Rates, curves and credit models

Interest-rate movements can affect borrower affordability, product pricing and portfolio behavior, but the mapping depends on contract terms. A fixed-rate loan does not reprice like a floating-rate facility. A benchmark change may affect a specific cohort only after a reset date. A model feature called current rate should therefore identify the applicable benchmark, reset mechanics, observation date and customer contract. A central-bank policy rate is a macro context variable, not necessarily the rate charged to a borrower. Using one as a substitute for the other can produce an attractive but misleading explanation of repayment risk.

For a loan portfolio, construct features such as remaining months to reset, payment-to-income change under a defined rate scenario and current contractual rate relative to the relevant benchmark. Each requires internal contract data joined to external rate observations. The joining logic must handle missing reset dates, caps, floors and rate types. A projected future rate belongs in a scenario, not in a historical observed feature. If the bank uses a forward curve, record the date and market source of the curve so an older forecast can be replayed. A model validation should test that a feature moves in the expected direction for a fixed-rate and floating-rate account under a controlled rate shock.

Do not conflate predictive performance with safe pricing. A credit-risk model may help estimate default likelihood, but product pricing also depends on funding cost, capital, expected loss, competition, customer treatment and approved policy. A model using market data to adjust an offer must have an explicit action boundary and fairness review where relevant. The bank should explain which values were observed, which were forecast, and which were chosen by policy. A model output that estimates vulnerability to a rate increase should not be turned into a higher price without a separate business and conduct decision.

Sanctions, company and adverse-media sources have different duties

An external list of designated persons is a compliance reference with legal and operational consequences. It is not an ordinary market signal. A list update has a publication time, ingestion time, effective scope and matching rules. A potential name match requires investigation under the bank's sanctions programme, with a case record and authorised disposition. A machine-learning ranking of likely false positives can help allocate work, but must not bypass a required screen or clear a hit merely because historical cases with similar names were cleared. The screened message and list versions must be retained for replay.

Company registers can help establish legal entity identity and beneficial ownership, but coverage, timeliness and jurisdiction vary. A registered address may be stale; an apparent ownership change may need verification. Adverse-media services may deliver articles that are duplicated, mistranslated, about a same-name person or irrelevant to the bank's risk purpose. A model that converts a media count into a customer-risk score without source review can penalise the wrong person. A compliant workflow distinguishes an observed publication, a verified entity match, a relevant allegation, a case assessment and a final action. Access and retention follow the bank's purpose and applicable law.

If external compliance data is included in a model, specify whether the feature represents an unresolved match, a confirmed disposition, a prior alert or a verified legal status. Those are different facts. A case opened after a lending decision cannot be an earlier input. An unresolved alert should not be coded as confirmed wrongdoing. Where a source is incomplete or inaccessible, the model should report the limitation rather than infer that the customer is clean. This boundary keeps external data from becoming a hidden proxy for an unreviewed compliance conclusion.

Licence, privacy and permitted use

The ability to download a dataset does not establish permission to train a model on it, store it indefinitely or expose it to every team. Market-data vendors often distinguish display, non-display, derived-data, redistribution and historical usage rights. Bureau contracts and privacy rules can impose different purposes and retention conditions. The source catalogue should record the licence, lawful basis, permitted users, storage location, transformation limits, expiry and renewal owner. Legal and procurement review should happen before a data scientist builds a dependency that the bank cannot lawfully deploy.

A derived feature can still reveal protected source data. A model that uses a proprietary price series to create an embedding or a summary may trigger licence questions even if it does not show raw prices. A credit feature inferred from bureau data may remain personal data and need access restriction, correction handling and auditability. Define what leaves a secure environment in training extracts and whether vendor or foundation-model services may receive it. Use the minimum fields needed for the approved decision. The governance record should include a way to remove or replace a source if rights change, without silently serving stale features.

The operational fallback must be planned. If a vendor feed stops at noon, should the model refer cases, use a documented last-good value within a maximum age, switch to an approved source or pause an automated action? The answer depends on the use. A treasury limit may require a current price; a longer-horizon forecast may tolerate a lagged macro release. Record the fallback state as an input and monitor how often it occurs. A model trained mostly on complete data can behave unexpectedly when missingness becomes widespread, so failure tests must cover source outages and partial feeds.

The ingest contract and quality gates

For each feed, the data owner should maintain a record of provider, delivery method, series or instrument identifier, source publication schedule, version, entitlement, unit, precision, allowed values, timezone, holiday calendar, expected count and consumer. Ingestion validates syntax and business meaning separately. A numeric field that parses successfully may still be wrong because a percentage was scaled by 100, a currency pair was inverted, a curve tenor changed, or a vendor published a revised methodology. Record accepted, rejected, late and corrected observations, then alert an owner when a defined threshold is breached.

Cross-source checks can be useful but require care. Two sources may publish different rates legitimately because they represent different times, markets or conventions. A discrepancy rule should compare like with like and state tolerance and escalation. A large move may be a real event; blindly clipping it as an outlier could hide the very signal a risk model needs. Quarantine a suspicious value while retaining its raw form and provenance. An authorised analyst can compare the provider's correction notice, alternative observations and internal trade evidence before deciding whether to accept it. The decision and the features affected should be recorded.

Quality reporting should show timeliness, completeness, accuracy checks, revision frequency and unresolved exceptions by series and model use. A green daily file count is weak if the same stale data was resent. Test the bank's ingestion time against the source's publication time, not merely the file arrival time. If an observation is backfilled, its availability time remains when it entered the bank, not the earlier economic period it describes. An offline training table must reproduce the same availability rule as the online store. Independent review should sample raw provider records through transformation, feature publication and a model decision.

Feature design: a value is not the whole feature

Consider a seven-day exchange-rate change used to contextualise payment behavior. Define the currency pair, source, fixing time, business-day calendar, observation window, missing-day treatment and calculation. If the model scores at 10:00, the latest daily reference may be yesterday's observation; do not leak today's later fixing into the feature. Store both the computed change and its age. A rate movement may explain an aggregate pattern of international payment volume, but it should not be used to accuse a specific customer of fraud without validation against legitimate travel, commercial activity and product context.

For a macro feature, define whether a value is a level, change, surprise relative to a forecast, rolling average or scenario assumption. A release surprise requires a pre-release expectation available at the time, not today's revised consensus. A rolling average must specify the vintages used at each point. For a credit model, the right horizon and geography matter more than the number of indicators. Adding ten highly correlated series can make a model harder to explain without improving decisions. Compare candidate features with a baseline and test stability across release revisions, economic regimes and borrower segments.

External signals are often shared across many borrowers on the same date. Randomly splitting borrower rows into train and test can place the same macro period in both sets and overstate future performance. Use time-based holdouts and, where relevant, geographic or product stress slices. Clustered outcomes and common shocks require careful confidence assessment. A model may appear to learn from unemployment when it is really learning the calendar period of an earlier policy. Validation should test whether the feature adds information beyond time and internal portfolio changes, and whether the gain persists in a later period.

A point-in-time replay example

Take a fictional loan decision at 09:00 on 15 May. The model uses a monthly employment series, a daily benchmark rate and the applicant's bureau response. The employment series describes April but was first published at 08:30 that day; the bank's feed ingested it at 09:10. The bank's 09:00 model could not have used it, even though it was public for thirty minutes. It must use the earlier available vintage or invoke its documented missing-data treatment. The rate observed on the previous business day is available at 09:00, but today's later fixing is not. The bureau response arrived at 08:58 and can be used if it passed identity and quality checks by the scoring boundary.

Six months later, the employment series has been revised and the bureau provider has corrected an account. A current portfolio analysis may use those corrected facts, but a replay of the 15 May decision must recover the old employment vintage, old bureau response, feature code and model version. If the bank cannot do that, it cannot confidently explain what influenced the original outcome. A validator should construct this case in test data, include a late feed and a later correction, and compare the production score with a reconstructed score at the same decision time. Differences need investigation, not an assumption that the latest data is better.

The same example shows why source names are insufficient. Employment series A from a statistical agency may be available to the public at 08:30, while the bank's licensed feed was delayed until 09:10. A market reference rate B may have a publication time but no executable quote at that value. A bureau response C may be current but matched to the wrong person. The feature store needs availability and validity evidence for each. An ML system that ingests more sources without those distinctions can become less trustworthy.

Evaluation that matches the action

For a liquidity forecast, measure error at the currency and settlement window where treasury acts, including large misses. For a credit-risk feature, assess incremental discrimination or calibration and decision outcomes after controlling for development-period effects. For a repair predictor, an external market variable may be irrelevant even if it correlates with volume. A feature is useful only if it improves a defined intervention at acceptable cost and risk. Validate on a period after development, compare against a simpler baseline and review the effect of missing, revised and stale inputs.

Watch for confounding. Suppose payment delays increased during a volatile FX week because a correspondent had an outage. A model that sees exchange-rate volatility may predict delay, but changing a payment route based on the rate would not address the outage and might worsen liquidity. The analyst should trace the causal chain and ask whether the feature identifies a condition the bank can act on. If the model output is a queue priority rather than an automatic decision, evaluate whether operations actually resolve cases faster and whether low-priority customers are harmed. Report both model metrics and business consequences.

Monitoring after release should separate feed health from model behavior. Track missingness, age, revisions, schema changes and licensing state for each external source. Then examine score distribution, outcome performance and overrides by relevant segment. If a provider changes a series definition, the model may drift even if the economy did not. A change-control record should identify affected features, models, decisions, customer cohorts and remediation steps. Pausing or restricting use may be safer than retraining immediately on a small, biased sample. The owner should approve any material change and preserve the previous version for audit.

A delivery contract a business analyst can test

For one proposed external feature, write an acceptance table before engineering starts. The first row states the model use and action: for example, a weekly portfolio forecast that helps a risk committee assess an approved scenario. The next rows state source series and owner, geographic population, units, publication calendar, initial and revised values, availability cutoff, transformation, maximum age, permitted use, retention and fallback. Then list consumers, the quality alert owner and the decision if the feed is unavailable. This is not paperwork after a model is built; it determines whether the historical dataset and live service answer the same question.

Test the contract with a normal publication, a revised value, a missed release, a holiday, a duplicate record, an unexpected unit change and a source cancellation. For each case, specify the accepted value, quality state and model action. A stale signal may cause a referral, a reduced-confidence forecast or no score, according to the approved use. Verify that the user interface indicates a restricted or fallback result. The test evidence should include the raw source item, transformed record, feature value, score and final action. A developer can check the pipeline ran; a business analyst must check that the bank acted on the right interpretation.

For a price feed, simulate bid and ask reversed, a currency pair inverted and a price observed outside its normal trading window. For a macro feed, simulate a release revised twice and a series rebased to a new index. For bureau data, simulate no hit, multiple candidate matches and a correction after a dispute. The same generic missing-value rule cannot handle all these cases. The control owner signs off the source-specific behavior, and model validation confirms that the resulting feature distribution is within the tested range. If not, the service needs a documented restriction or fallback.

A cross-team incident scenario

At 08:00 on a Monday, the bank's feed presents a sharp fall in a market rate. A treasury forecasting model changes its projected funding need, and a credit monitoring model flags a portfolio shift. Operations first sees that the new observation has a different unit and an unchanged source timestamp from Friday. The data team quarantines the value, identifies a provider format change and restores the last valid observation within the approved age limit for the weekly forecast. The intraday treasury control, which requires a current value, switches to its approved manual process. Neither model should silently continue as if the suspect rate were a genuine market shock.

The incident record identifies the raw file, affected features, model versions, decisions made before detection, customer or financial impact, fallback duration and correction. Treasury reviews any funding action it took; risk reviews whether credit cases were referred or restricted; technology fixes parsing and tests the provider's new format. Independent validation asks whether pre-release checks should have caught the unit change. This scenario illustrates a core AI/ML lesson: model performance is a property of the data path and action, not just fitted parameters. The correct response is controlled investigation and remediation, not an automatic retraining job on a bad observation.

Governance and source notes

The bank should keep an inventory linking each external dataset to the models and reports that use it. An owner reviews entitlements, data quality, vendor changes, approved purposes and downstream dependencies. A model owner documents which signals are necessary, the limitations of each, the point-in-time construction and the expected behavior under missingness. A validator challenges whether the signal adds reliable value and whether the historical test used the vintage that was actually available. Operations and finance own the actions and accounting consequences. No team can transfer its decision responsibility to an external feed or a model provider.

These sources are starting points for the evidence chain, not substitute approvals for a bank implementation:

The practical boundary

An external signal should enter a bank model only with a defined purpose, verified source, point-in-time availability, lawful use, tested transformation, fallback and owner. Its value may be highly predictive in one historical sample and still be unsuitable for a live decision because it arrives too late, changes after publication, measures a different population or cannot legally be reused. The strongest design is often smaller: a few well-defined signals connected to an observable action, with a clean baseline and a replayable decision record. That gives a reviewer a way to ask what the model knew, why the feature mattered, what the bank did and whether the customer or balance sheet benefited.

Questions that reveal a weak external feature

Ask the team to demonstrate the exact value a model saw on a randomly chosen historical decision date. If it can show only the latest series, the feature is not ready for point-in-time validation. Ask what happens when the source revises a prior month and which version remains in the audit snapshot. Ask whether the feed's timestamp means market observation, provider publication, bank receipt or feature calculation. A single column named date is rarely enough to answer all four. Ask whether the licence permits this specific model use and whether the data can be shown to a customer, validator or regulator if needed.

Ask the model owner to explain what action the feature changes. If removing it does not alter a decision, forecast quality or control outcome in a meaningful test, its operational cost and risk may not be justified. If it appears to improve an aggregate measure, inspect the gain by period and segment and check for look-ahead bias. A monthly macro release repeated on every customer row can make a random test split look convincing while providing little evidence of future performance. Demand a time-based test and a simple comparator. If the feature is sensitive to source revisions, report both first-vintage and latest-vintage results rather than hiding the difference.

Ask the operations owner what they will do when the external value conflicts with internal evidence. A bureau attribute might indicate an obligation that the applicant disputes; the bank needs a correction and review path. A market quote might diverge from the rate at which a transaction executed; accounting must use execution evidence for that transaction. A sanctions list alert might arise from a same-name person; compliance must investigate identity rather than average the alert into a generic score. These are different decisions, and a common data lake does not make their authority the same.

Finally, ask whether the feature can be stopped safely. A source outage, contract termination or provider methodology change should have an inventory of dependent models and a tested fallback. The bank may continue some forecasts with a clearly marked last-good value, refer customer decisions, or suspend an automated limit. The action needs a named owner and a communication path to affected teams. If no one can say which live decisions depend on a source, the bank has an integration risk that a good offline model metric will not reveal.

Related learning paths

This application uses JavaScript for the full interactive experience. This text summary is served for accessibility and search indexing.

Market data and external risk signals · Malla Banking Academy