Decision thresholds and outcome handling

Decision thresholds and outcome handling. A practical lesson in ai data and model operations for banking and payments practitioners.

Plain language meaning

Decision thresholds and outcome handling explain how banks convert model scores into approve, decline, refer, hold, monitor, investigate or escalate outcomes, while controlling customer impact, risk appetite, false positives, false negatives, override authority and evidence.

This topic is about bank decision design after an AI or ML score. It is not about assuming the highest score automatically becomes the correct action.

In a real bank, this topic cannot be handled as a loose data-science or technology idea. It affects customer outcomes, fraud and AML control, operational queues, service continuity, privacy, security, model governance, audit replay, management reporting and regulatory confidence. AI should improve speed and quality, but the bank must still prove source data, permitted use, approved logic, human accountability, fallback handling and retained evidence.

Where it sits in the banking AI journey

This card belongs to AI Data and Model Operations. The working flow is Model score, Threshold policy, Outcome route, Human or system action, and Evidence and monitoring.

Read the flow as a bank operating model. Each stage needs a source system, a data owner, a timing rule, a quality gate, a model or rule boundary, an exception path, a customer-impact view, a fallback option, a monitoring requirement and a retained record. That is what separates useful AI adoption from uncontrolled automation.

Banking data and evidence

The important data points are score, confidence, threshold band, reason code, risk appetite, customer segment, override reason, and final outcome. These items matter because they can influence risk scoring, operational repair, fraud action, AML triage, customer treatment, reporting, model monitoring and management decisions.

The evidence pack should include threshold paper, decision trace, reason code, override log, customer notice, monitoring report, and recalibration record. A strong bank can replay the journey from source data to transformed input, AI output, rule result, human action, system outcome and monitoring result. A weak bank only knows that a process ran and hopes the process was right.

Controls that make AI adoption safe

The core controls are threshold approval, reason-code mapping, impact testing, override governance, four-eyes review, outcome monitoring, and periodic recalibration. These controls keep the topic anchored to banking purpose, approved policy, data governance, model-risk expectations, operational resilience, customer fairness, privacy, security and auditability.

The practical design should define what AI may recommend, what it must never decide alone, which deterministic rule remains authoritative, who owns thresholds and overrides, how degraded service is handled, how customer harm is detected and what evidence is retained. Without that control design, faster AI can simply make weak processes fail faster.

Architecture and data-operation lens

Banking AI depends on the architecture around it. Storage, streams, feature definitions, training sets, model versions, thresholds, feedback labels and rollback paths must be governed before the bank relies on AI output. The model is only one part of the control chain.

A bank-grade design connects channels, source systems, core records, payment hubs where relevant, fraud systems, AML platforms, case tools, data platforms, feature stores, model-serving endpoints, policy engines, audit logs and management dashboards. It also records degraded operation, recovery actions and lessons learned.

Regulatory and governance lens

Federal Reserve SR 26-2, dated 17 April 2026, gives revised model-risk guidance for traditional models and non-generative AI models used by banking organisations, including development, validation, monitoring, change control and governance.

The Federal Reserve's 2026 model-risk guidance states that generative and agentic AI are outside that guidance, while broader bank risk-management and governance practices still need to control tools and processes not covered by the guidance.

NIST AI RMF 1.0 uses Govern, Map, Measure and Manage functions for AI risk management, and NIST AI 600-1 adds generative-AI risk actions for source grounding, content provenance, data protection, cybersecurity and human oversight.

FFIEC Architecture, Infrastructure and Operations guidance expects financial-institution technology environments to be governed, resilient, secure, monitored, documented and aligned to business risk, including emerging technologies such as artificial intelligence and machine learning.

BCBS 239 remains current for effective risk data aggregation and risk reporting, and the Basel Committee's January 2026 newsletter reiterates the importance of accurate, comprehensive and timely data capabilities in banks.

FFIEC BSA/AML examination guidance expects suspicious activity monitoring systems and supporting technology to be risk-based, explainable by management, independently tested where appropriate and aligned to the bank's risk profile.

FinCEN's 12 June 2026 Section 314(b) materials clarify information sharing for possible terrorist activity, money laundering and fraud-related specified unlawful activity within the statutory safe-harbor framework for participating financial institutions.

OFAC's Framework for Compliance Commitments describes sanctions compliance programme components including management commitment, risk assessment, internal controls, testing and auditing, and training.

Diagram walkthrough

Read the diagram from left to right as Model score, Threshold policy, Outcome route, Human or system action, and Evidence and monitoring. It is a banking control map. The point is to show how data, AI or ML output, rules, human action, operational routing and audit evidence should connect.

Use it as a 30-minute study method. For each box, ask which system creates the data, which definition is used, which model or rule acts, what can go wrong, who can override it, how a fallback works, which customer or regulatory impact exists and what record proves the final state.

Most important mistake to avoid

The common failure is treating a threshold as a technical parameter. In a bank, a threshold is a business, risk, compliance, customer and operational decision that must be approved, monitored and explainable.

The correction is disciplined scope. Keep the chapter anchored to banking purpose, prove the data path, make ownership visible, test failure behaviour, record the evidence and make the final outcome explainable without relying on memory, assumptions or developer-only knowledge.

A threshold is a policy choice

A fraud score ranks transactions, but it does not by itself say which payment to hold. A bank might choose a threshold that sends cases to investigators, constrained by review capacity, customer friction, loss exposure and mandatory controls. The acceptable trade-off differs by product, channel and consequence. A threshold calibrated on last year's labels may not work after fraud patterns, customer mix or investigation procedures change.

Create a dated decision table: score band, hard-rule results, action, authority, customer status and escalation. A sanctions hit, for example, should not be converted into an ordinary fraud risk score or bypassed because a model is confident. A missing score is not a low score. A timeout may trigger a predetermined fallback, and an investigator's release should be recorded separately from the initial recommendation. The bank's approved policy, not the model developer's convenient default, owns these outcomes.

For evaluation, compare true and false positives, missed cases, queue volume, investigation time and customer impact on an appropriately labelled cohort. Labels may be delayed or revised. Test scores exactly on the boundary, contradictory control results, retry after timeout and an unauthorised override. Preserve model version, threshold version, rule path and final disposition so the outcome can be reconstructed. The NIST AI RMF Measure playbook emphasizes contextual testing and monitoring; it does not prescribe one optimal banking threshold.

A score needs a policy

A model score ranks or estimates an outcome; it does not by itself decide what a bank should do. A fraud score can lead to release, extra authentication or review depending on payment value and policy. A credit probability can lead to approval, referral or decline only in combination with affordability, eligibility and human processes. An AML priority score can order a queue without closing underlying alerts. The policy that maps a score to an action needs its own owner, version, evidence and monitoring.

Define the unit and time of the decision. A payment authorization has a tight deadline; a loan application can have multiple review stages; an alert may be reopened. The policy must state which model output was available, whether critical features were valid and what other rules applied. A timeout or abstention is not a score of zero. The decision journal should retain the actual path and final outcome.

Thresholds should reflect business consequences. A lower fraud threshold can catch more suspicious transfers but also hold legitimate customers and fill analyst queues. A higher threshold can reduce friction while allowing more losses. The tradeoff changes with transaction amount, prevalence, review capacity and delayed outcomes. A single global threshold may be inappropriate, but segment-specific policies need fairness, legal and governance review.

Understand score meaning

Some scores estimate probabilities; others are arbitrary rankings. A value of 0.8 from one model is not necessarily an 80 percent chance of fraud. Check calibration on a representative, mature population before using probabilistic language. Prevalence changes can alter precision even if ranking quality remains stable. Compare scores across model versions only after validating their scales.

A threshold chosen on a balanced training dataset can fail on rare-event production traffic. If one in a thousand transactions is confirmed fraud, even a useful classifier can send many legitimate payments to review at a low threshold. Evaluate on real prevalence and report expected number of cases per day, losses caught and false holds. Include uncertainty in rare and delayed labels.

Score distributions also change when inputs fail. A stale velocity feature may push many payments toward low risk; a customer-master merge may raise scores for an entire cohort. Monitoring an average score without input quality and action rates can miss the cause. Threshold governance includes data validity, not just a numeric cutoff.

Multiple action bands

A three-band policy may release low-risk payments, challenge or review middle cases, and hold high-risk cases. The actual bands must be validated for the bank's product, latency and obligations. There can be independent rules for value, account status or mandatory screening. A model-based low-risk band must not bypass a legal hold. Record which rule took precedence and why.

For credit, a middle band can route uncertain cases to underwriters with source evidence and permitted override reasons. A manual queue has finite capacity and can cause delays. The policy should specify service times, customer communication and when a case is escalated. A high score may still be overridden by a deterministic affordability rule; a low score should not automatically trigger an adverse action without the approved decision process.

For AML alert ranking, thresholds may define priority tiers rather than clearance. Cases below a ranking cutoff can still require investigation under the bank's monitoring program. Validate that lower-priority cases are not indefinitely deferred and sample them for missed important outcomes. An improved queue metric is not permission to suppress mandatory alerts.

Optimize the full workflow

Evaluate thresholds against final decisions, not only model labels. A payment held by the model may later be released after customer verification. The bank should measure both prevented losses and false friction, including time held. A credit referral can lead to approval after additional documents. An AML case may consume analyst time without a confirmed finding. The model threshold changes workload and customer experience as well as statistical classification.

Use explicit cost assumptions and sensitivity analysis. Fraud loss amounts are skewed; preventing one large loss can dominate averages. False holds can harm customers even when they are brief. An analyst hour has opportunity cost, and a backlog can delay genuinely high-risk cases. Show results across plausible ranges rather than a single invented monetary optimum. Have business owners approve the objective and limits.

Capacity can change during incidents. A threshold producing 500 reviews per day may be workable with ten trained analysts but not when half the team is unavailable. The fallback plan should specify a controlled mode, escalation and customer treatment. Changing a threshold in an emergency requires authorization, scope, expiry and post-incident impact review.

Policy precedence

Write a decision table that orders mandatory controls, data validity, model score, other rules and human action. A payment can have a low fraud score and a possible sanctions match; the screening hold remains governed by its separate control. A model can recommend credit approval while an affordability rule prevents it. A case worker may resolve a fraud alert but not override a compliance restriction without authority.

Avoid hidden defaults in code. If the model API times out, the orchestrator should use an approved fallback, not treat the absent response as zero risk. If a score is malformed or outside range, reject or refer it. If a feature is stale, the model may need to abstain. Log the reason and policy version. Test each branch with representative requests.

The final customer-facing status should reflect the actual transaction or application state. A payment sent for manual review is not settled. A referred loan is not declined unless the process records that outcome. A generated assistant recommendation is not a final policy action. Reconcile the policy journal to payment, credit and case systems.

Calibration and review

Calibrate a probability model on mature outcomes for the intended population. Assess reliability by score band, period and relevant segment. A global calibration curve can hide poor calibration for new customers or a migrated product. If source or policy changes alter who is observed, report that limitation. A decline or hold changes which outcomes can later be seen.

Threshold review should use out-of-time evidence and include uncertainty. Fraud labels can take weeks; credit defaults months. Recent action rates and source quality are leading signals, not definitive performance outcomes. Schedule later mature-cohort analysis. A temporary drop in observed loss can result from more holds or an incomplete label feed.

If calibration changes, determine whether the model, inputs, label process or population shifted. A source mapping defect should be repaired at source. A new policy threshold should be versioned separately. Retraining can be appropriate after genuine behavior change, but not as a substitute for investigating data quality.

Segment effects and fairness

Analyze referral, hold, approval and error rates across relevant groups under legal and privacy controls. A feature missing for thin-file customers can push them into manual review more often. A channel with slower data delivery can have more fallbacks. The threshold may be numerically the same while the experienced outcome differs. Report both model and policy effects.

Review override patterns. If underwriters frequently reverse a particular group's referrals, investigate whether source data, model calibration or review practice is responsible. Do not automatically treat overrides as ground truth. Record reasons and new evidence. A fairness assessment should include the full workflow, including delays and access to meaningful human review.

Segment-specific thresholds can address different base rates or objectives but can also create legal or ethical problems. The bank's policy and legal owners should decide what is permitted and justified. Document the rationale, test alternatives and monitor outcomes. Do not hide a segment policy inside an undocumented feature or code branch.

Outcome definitions

Keep model outputs, policy decisions and later outcomes separate. A fraud score is an estimate; a hold is an action; a confirmed fraudulent transfer is a later outcome. A credit default label is not the same as an initial decline. An AML alert disposition is not proof of absence or presence of crime. Training and monitoring should join them with IDs and time windows.

Outcomes can be revised. A customer may dispute a transaction, an investigation may reopen, a loan may restructure. Define maturation and correction rules. Preserve the initial decision and later events. A performance report should identify which label version and cohort it uses. Overwriting historic outcomes can make model comparisons impossible.

Selection bias arises because policy controls observation. A blocked payment may never generate a chargeback; a declined applicant cannot default on the unfunded loan. Measuring only approved cases can make a strict threshold look artificially safe. Use controlled sampling or other justified evaluation methods where appropriate and disclose unobservable outcomes. Do not label every prevented transaction fraudulent.

Fraud example

Suppose a model ranks transfers by risk. On a representative mature cohort, the bank evaluates several thresholds against case capacity and losses. At one setting, 200 transfers per day enter review; at another, 800 do. Compare incremental confirmed fraud caught, legitimate payments delayed, average hold duration and segment distribution. Include high-value tails and scenarios where labels are incomplete. The policy owner chooses a setting within risk appetite and service capacity, with documented reasons.

During a feature-store outage, 5 percent of requests lack valid velocity. The score route is unavailable for them. A separate approved fallback handles those instructions and records the path. Monitoring compares their actual payment outcomes and customer friction with ordinary decisions, acknowledging that outage traffic may differ. The threshold itself is not secretly lowered to compensate for missing data.

Credit example

A credit model estimates default risk for an application. A low-risk score can support automated processing only if identity, eligibility and affordability checks pass. A middle-risk score may request human review and additional evidence. A high-risk score may lead to an adverse action under approved policy and applicable obligations. The decision log identifies which factors actually governed the outcome.

A bureau correction arrives after a decline. The bank recomputes features and score, then applies the appropriate review and remediation process. A changed probability does not automatically mean the final decision changes; another rule may still govern it. Preserve both original and corrected records and communicate accurately. Monitor how often source corrections materially change actions.

AML example

A model ranks transaction-monitoring alerts by expected investigative value. Priority bands direct staffing, but underlying alerts remain in the governed case population. Track oldest case age, high-severity coverage, analyst workload and outcomes from sampled low-ranked cases. If the model or feed fails, revert to an approved queue order and measure backlog. A threshold that improves apparent precision by leaving difficult cases unreviewed is not a successful control.

If a new screening list causes many possible matches, do not use a low model score to auto-clear them. Screening requirements and dispositions follow their own policy. Separate model priority, mandatory screening and final case outcomes in reports.

Version and test the decision table

Store model artifact, feature and policy versions for each request. A threshold change can alter actions without changing model scores. A new model can change score scale while policy configuration stays constant. Roll out them together only after testing the combination. Shadow comparison should evaluate final action differences and workload, not merely score correlation.

Create tests for score just below and above each boundary, missing score, stale feature, mandatory hold, human override, duplicate retry and late model response. Check final state in authoritative systems. A threshold test that only asserts a numeric comparison misses precedence and outcome handling.

Have an independent reviewer sample actual decisions after release and reconstruct score, data validity, policy branch, human action and final outcome. Investigate gaps and unexpected action distributions. A bank can manage a model threshold only when it knows precisely which decisions it changed and what happened to customers afterward.

Decision table worked example

Consider a payment with four inputs: mandatory screening status, feature validity, fraud score and payment value. First, an open screening hold routes to compliance regardless of fraud score. Second, if a critical fraud feature is stale, the approved degraded policy applies rather than interpreting a score calculated from it as normal. Third, if inputs are valid, the fraud policy may release low-risk payments, challenge a middle band and refer a high band. Fourth, a value limit may impose an additional hold. Each branch produces a reason code and an action recorded against the payment instruction.

Test combinations, not only individual rules. A high fraud score and screening hold should yield the correct combined case state without duplicate customer messages. A model timeout with a high-value instruction should follow the fallback value rule. A low score with an invalid account should still be rejected under account controls. A human analyst may release a fraud referral but cannot silently remove an independent compliance hold. The authoritative payment hub status should agree with the decision journal.

Now change the fraud threshold version. Replay a representative historic population through the candidate policy without executing payments. Count action changes by product, value, channel and customer group. Estimate additional review load and possible loss tradeoffs using mature labels and explicit uncertainty. Compare this with a model replacement at the old threshold; the two changes should be evaluated separately. If both deploy together, retain enough evidence to attribute an unexpected shift.

Queue capacity and service times

A threshold is constrained by downstream review. If a fraud team can handle 300 cases per day and the selected band generates 500, backlog grows even when model classification is statistically strong. Queue age then changes the value of detection: a suspicious instant payment cannot wait until tomorrow for a decision. Calculate arrivals by hour, priority, staff availability and service duration. Include weekends and incident peaks. A daily average obscures short periods of overload.

For AML, rank ordering can improve attention to higher-risk cases but must maintain coverage and escalation for the full governed alert population. Measure oldest case, distribution of ages and outcomes of sampled lower-priority cases. A drop in false-positive closures might indicate better triage or merely fewer reviews. Report both denominator and disposition maturity. A model that reduces workload by suppressing alerts without authorization changes the control, not just the ranking.

For credit referrals, delay can affect customer decisions and fairness. A threshold that sends thin-file applicants to human review may be reasonable if reviewers have time and evidence, but harmful if those cases routinely wait longer or are abandoned. Measure completed applications, withdrawals, review times and final outcomes by segment. Treat the referral experience as part of the policy's effect.

Economic evaluation with uncertainty

Avoid a single "cost per error" assumption. A fraud loss includes direct amount, recovery, investigation and customer effects; a false hold includes inconvenience, lost trust and possible downstream payments. Different products have different values and reversibility. Show threshold tradeoffs over a range of plausible costs and prevalence. State which assumptions materially change the preferred action.

For credit, expected loss estimates depend on default probability, exposure and recovery assumptions. A model score may inform risk pricing or review, but other rules and legal obligations govern the final action. Examine how calibration error affects outcomes near a policy boundary. A small score shift can move many applications if the distribution is concentrated near a cutoff. Stress those boundaries under changed economic conditions.

For AML, monetary loss alone is an inadequate objective. Missed serious cases, compliance obligations, investigative capacity and customer friction matter. Use domain-defined severity and coverage constraints, plus sampled case review. A statistical threshold optimization cannot override mandatory controls.

Selective labels and exploration

The bank sees detailed outcomes for transactions it allowed and accounts it funded, but not equivalent outcomes for those it blocked or declined. This limits threshold simulations. A prevented fraud loss cannot be measured directly for every held payment. A credit applicant who was rejected cannot be assumed to have defaulted or repaid. Report observed outcomes, assumptions and uncertainty separately.

Where lawful and appropriate, carefully designed review samples can improve knowledge about lower-ranked cases. Human analysts may inspect a random subset of low-priority alerts to detect missed patterns. A bank can monitor outcomes under controlled policy changes, but experiments involving consequential customer decisions need governance and safeguards. Do not casually randomize mandatory controls.

Historical data reflects previous thresholds. If old rules blocked many high-risk payments, a new model trained on settled transactions may have little direct evidence about that region. Validate with domain experts, synthetic stress cases and monitored rollout. A strong aggregate metric on observed cases does not prove the safety of a newly released population.

Human override governance

Define who may override each branch and what evidence is required. A fraud analyst can resolve a case based on verified customer contact; a credit underwriter may consider additional documented income; a compliance officer handles a screening match under a separate process. Record the original recommendation, new information, reason and final action. A free-text "approved by manager" field without context is weak evidence.

Overrides should feed monitoring but not automatically become labels. A reviewer can be mistaken; an override can reflect policy or customer circumstances rather than model error. Analyze frequency and direction by product, model version and segment, and sample underlying cases. A spike after a source migration may indicate data quality; a spike after threshold change may indicate poor policy calibration or insufficient capacity.

Review authority during outages. If the case tool is unavailable, a manual spreadsheet process may not meet privacy, audit or reconciliation needs. The approved fallback should specify capacity, access and later integration. A technical model fallback is incomplete if the people who must decide cannot see valid evidence.

Outcome reconciliation

Map each unique eligible instruction or application to exactly one current business state and a history of transitions. Model requests can be multiple because of retries; policy evaluations can be repeated after new information; the business action must remain coherent. Reconcile decision journal to payment hub, ledger, application system or case tool. Investigate orphan model scores and actions with no score or fallback reason.

For payments, distinguish instruction accepted, held, settled, returned and canceled. For credit, distinguish submitted, referred, approved, declined, withdrawn and funded. For AML, distinguish alert created, assigned, reviewed, escalated and closed under approved taxonomy. A model performance report that treats all terminal statuses as the same outcome cannot measure actual effect.

Preserve corrections. If an application is reconsidered after a bureau update, keep the original decision and a linked new assessment. If a fraud hold is released, measure delay and customer communication. If an AML case is reopened, retain its prior disposition. Later outcomes should join to the decision version that actually governed the initial action.

Monitoring triggers

Watch score distribution, input validity, action rate, queue age, override rate, mature outcomes and segment effects together. An increase in referrals can come from model drift, a threshold change, source missingness or a new channel. Diagnose by version and source before responding. A threshold change can be reversed quickly if a preapproved rollback exists, but restoring an old policy without compatible model scores is unsafe.

Define quantitative triggers with responsible owners and review windows. A sudden feature-staleness spike may require immediate fallback. A slow calibration change needs mature labels and analysis. A growing manual queue may require staffing or a controlled policy change. Record the evidence reviewed and the decision to continue, restrict or change the use.

Communicate limits to operations. A model score may be unavailable; a middle band may mean review rather than rejection; a compliance hold may supersede all fraud paths. Customer-facing messages should accurately describe the business status without exposing confidential detection details. Training staff on these distinctions prevents a correct model-policy combination from failing in the human workflow.

Threshold change drill

Take a candidate threshold change for a representative month. Rebuild the eligible population and original scores from decision records. Apply the proposed policy in simulation, including mandatory rules and feature-invalid cases. Count changed actions, new case volume by hour, possible fraud losses and legitimate customer delays. Stratify by product, channel and relevant groups. State which labels are mature and which blocked outcomes cannot be observed.

In a controlled rollout, monitor actual action and queue rates immediately and mature outcomes later. Define rollback conditions before launch. Sample boundary cases and inspect whether source data, score and reason codes agree with the decision table. If a rollback occurs, retain both policy versions and identify every affected instruction. That is the evidence needed to revise a threshold responsibly.

Worked boundary review

Suppose a fraud threshold refers scores at or above 0.70 and releases lower scores, subject to other controls. A payment scores 0.6997. If the interface rounds to 0.70, an operator may believe the referral rule was violated. Store the underlying numeric output and specify comparison precision. Test exact equality, just below, just above, missing, nonfinite and out-of-range values. A malformed score should never be treated as low risk through a default comparison.

A second payment scores 0.72 but has a stale beneficiary feature. The correct path may be a data-validity fallback rather than ordinary high-score referral. A third scores 0.20 but has a separate mandatory hold. The policy must preserve the hold. A fourth is retried after timeout and receives a late score; the earlier fallback action remains the actual historical decision. These cases show why a single threshold test does not validate outcome handling.

For each, compare the orchestrator record with the payment hub's terminal state and case queue. A referral that never opened a case is an operational defect. A held payment that the ledger shows settled requires immediate investigation. A duplicate retry with two case IDs may inflate workload and confuse customer communication. Measure these exceptions, not merely whether the threshold code returned the expected string.

Policy ownership and expiry

Name the person or committee authorized to approve thresholds and the evidence required for normal and emergency changes. Record the model artifact and data definitions for which a threshold is valid. An old cutoff may be inappropriate after recalibration or a new feature version, even if scores still range from zero to one. Require a limited rollout and monitoring plan for material changes.

Emergency thresholds need an expiry, incident reference, scope and review. A temporary policy introduced during a fraud spike can remain indefinitely if no one owns the sunset. Compare its effects on losses, holds, queue capacity and segments when mature outcomes arrive. Either approve a durable policy through normal governance or restore the prior version with evidence.

Human reviewers need the current decision table and authority boundaries. If policy changes faster than case tools or training, employees may act on obsolete instructions. Version operational guidance with the policy and test a sample of live cases after deployment. The final decision depends on people and systems interpreting the same rule.

Audit question

For a particular declined or held customer, can the bank show the score as returned, validity of inputs, threshold and precedence rules, human action, actual business outcome and any later correction? Can it identify everyone affected by a threshold defect? Can it explain why a monitoring chart changed without guessing whether model, data or policy changed? If any link is missing, the threshold is not yet controlled as part of the AI decision system.

Boundary and precedence exercise

Test fraud scores immediately below, equal to and above a proposed review threshold using the unrounded model output. Then test a missing score, stale feature, screening hold and invalid account. The policy should record which condition governed each final action. A low fraud score cannot release a mandatory screening hold; a model timeout is not a zero score.

Replay a representative period under the candidate policy without executing transfers. Count changed holds, analyst workload by hour, legitimate delays and mature confirmed loss, with uncertain blocked outcomes labeled as such. In a bounded rollout, reconcile decision journal to payment hub and case queue. A threshold version should be separable from model and feature versions so an unexpected referral spike can be diagnosed. Test an analyst release when a separate screening hold remains open; the payment must stay held. A score rounded for display should not alter the comparison at the policy boundary. Retain unrounded output, precedence reason, human authority and final hub status for an independent review.

Related learning paths

This application uses JavaScript for the full interactive experience. This text summary is served for accessibility and search indexing.

Decision thresholds and outcome handling · Malla Banking Academy