Network Analytics and Entity Resolution

Financial crime is rarely organised as neatly as a bank's customer master. One person can appear under several customer records, one company can use several accounts and payment channels, several apparently unrelated customers can share a director or device, and one beneficial owner can sit behind a chain of companies that crosses jurisdictions. Criminal networks exploit those gaps deliberately, but ordinary customers also create messy data for innocent reasons. Families share addresses. Corporate groups share directors. Outsourced accountants use the same contact details for many clients. Mobile devices are replaced. Names are transliterated differently. Legal entities merge, split, rebrand and change ownership. A useful financial-crime control therefore needs to connect information without pretending that every connection proves common ownership, control or wrongdoing.

That is the role of entity resolution and network analytics.

Entity resolution asks a foundational question: which records most likely refer to the same real-world person, organisation, account, device, address, wallet or other subject? Network analytics starts after, or alongside, that work and asks: how are those resolved subjects connected, how do those relationships change over time, and which patterns deserve investigation? The distinction is important. If the bank joins two different people incorrectly, every downstream relationship built on that join may be wrong. If it fails to join two records that belong to the same person, a network may be fragmented and the bank may miss a pattern entirely.

A mature operating model therefore separates four things that are often collapsed into one score:

  1. Identity confidence — how confident are we that two records represent the same subject?
  2. Relationship confidence — how confident are we that a link between two subjects is real and meaningful?
  3. Financial-crime risk signal — does the resolved network contain behaviour or structure that warrants review?
  4. Legal or case conclusion — what does the applicable law, policy and investigated evidence require the bank to do?

A graph can reveal a path. It does not automatically prove beneficial ownership, sanctions applicability, conspiracy, money laundering or suspicion. Those are separate conclusions that need appropriate evidence and, where relevant, jurisdiction-specific legal analysis.

Records from customer, account, payment, device, corporate and digital-asset systems are resolved into entities, connected through evidence-backed relationships, analysed for risk patterns and then investigated before any legal or suspicious-activity conclusion is made.

The mental model: records are not entities, and links are not conclusions

Banks hold records because systems need records. A current-account platform needs a customer identifier. A card platform needs a cardholder record. A corporate onboarding platform needs directors, authorised signatories and beneficial owners. A sanctions engine needs screened names and identifiers. A payment hub needs debtor, creditor and agent information. A fraud platform may hold device, session and behavioural signals. A virtual-asset control may hold wallet addresses and analytics labels. These records are useful, but they are not automatically a clean representation of the real world.

Imagine that the same natural person appears as Mohammed Al Rashid, Muhammad Al-Rashid and M. Alrashid in three systems. One record contains a date of birth, one contains a passport number, one contains only a phone number and address. A simplistic exact-name join may treat them as three people. A simplistic fuzzy-name rule may merge them even if the date of birth differs materially. Entity resolution exists to make that decision more disciplined.

The same problem appears with legal entities. A company may have a registered name, trading name, former name, local-script name and abbreviated payment name. Registration numbers can change format. Branches can be represented as separate operational records even when they are not separate legal persons. A group can contain subsidiaries with nearly identical names. An acquired company may continue to exist legally while using the acquirer's brand. Resolution logic must therefore understand identifiers, entity type, effective dates and source quality rather than relying on string similarity alone.

After resolution comes the network. The network can contain many kinds of nodes: natural persons, legal entities, accounts, cards, merchants, counterparties, devices, addresses, phone numbers, email addresses, IP ranges, vessels, trusts, directors, beneficial owners, wallets and service providers. The edges between them also have different meanings: owns, controls, directs, signs for, pays, receives from, shares a device with, uses an address with, is a beneficiary of, is a trustee of, is parent of, is correspondent for, or transacts with.

Those relationships are not interchangeable. Sharing an address is weaker evidence of common control than a verified corporate-registry ownership link. A payment is evidence of value movement, not ownership. A device link may indicate one person controlling several accounts, but it can also reflect a family tablet, a branch kiosk or an outsourced bookkeeping service. A bank should therefore store the type, source, effective period and confidence of a relationship rather than reducing everything to connected = true.

Why network thinking matters in a bank

Many financial-crime controls are naturally record-centric. Customer due diligence is often performed per customer. Screening compares a party against a list. Transaction monitoring evaluates transactions or accounts. Fraud systems assess sessions and payments. Investigators receive an alert and open a case. This operating model works reasonably well when risk is concentrated in one customer or one transaction. It becomes weaker when risk is distributed across a network.

A mule operation is a simple example. Ten accounts may each receive modest credits that look ordinary in isolation. If the ten accounts were opened within days of each other, use three devices, transfer to the same two beneficiaries and rapidly pass funds onward, the network tells a stronger story than any one account. The relationship pattern becomes part of the evidence.

Complex ownership creates a different problem. A company may appear low risk if the bank looks only at its immediate shareholder. When the ownership chain is resolved through intermediate entities, a natural person, trust or sanctioned owner may become relevant several layers away. FATF's revised Recommendation 24 and related guidance emphasise adequate, accurate and up-to-date beneficial ownership information for legal persons, while Recommendation 25 addresses trusts and similar legal arrangements. These are global standards for transparency; they do not prescribe one graph algorithm or a universal ownership threshold for every bank decision. The institution still has to apply the law and policy relevant to the customer and jurisdiction.

Network analysis can also improve monitoring. The Wolfsberg Group's work on monitoring for suspicious activity encourages institutions to look beyond narrow transaction monitoring and use customer behaviour and attributes to identify risk more effectively. Relationship data can be one part of that broader view. A sudden new beneficiary may be more concerning when the beneficiary is already connected to known mule accounts. A corporate payment to a new supplier may deserve more scrutiny if the supplier shares directors and registration addresses with companies already linked to suspicious trade activity. None of these links is determinative on its own, but they improve prioritisation.

Entity resolution: how a bank decides whether records belong together

A practical resolution process normally has several stages.

1. Standardise without destroying the original evidence

Names, addresses, dates and identifiers arrive in inconsistent formats. The bank may remove punctuation for comparison, standardise date formats, normalise casing, parse addresses, convert scripts through transliteration and standardise legal-entity suffixes. The original value must still be preserved. An investigator may need to know exactly what the customer entered, what a payment message contained or what a registry reported at the time.

Normalisation is not truth. Converting St. to Street, for example, can help matching, but it can also create false equivalence in multilingual data. Transliteration is particularly sensitive: multiple valid Latin representations can exist for the same original-script name. The resolution service should record the transformation used rather than silently replacing the source value.

2. Generate plausible candidates

Comparing every record with every other record becomes computationally expensive and creates unnecessary false matches. Candidate generation, sometimes called blocking, narrows the population. It may use exact or partial identifiers such as passport number, company-registration number, date of birth plus surname, phone number, address key, LEI or another stable attribute.

Blocking must not be so strict that it guarantees false negatives. If records are considered only when the name is identical, transliteration differences will be missed. If phone number alone is used, recycled numbers or shared business contacts may create false candidates. Mature systems normally use several candidate strategies and measure which populations they fail to bring together.

3. Compare attributes with different evidential weight

Not every field deserves equal weight. A verified national identifier is generally stronger than a common surname. A company-registration number is stronger than a shared registered office. Date of birth may be highly discriminating for natural persons, but source reliability matters. An email address can be useful, but aliases and shared mailboxes are common. A device identifier may be strong behavioural evidence but is rarely appropriate as a legal identity key by itself.

The resolution engine can use deterministic rules, probabilistic matching, machine learning or a combination. The important control principle is not the specific technology. It is whether the bank can explain what evidence drove the match, how thresholds were chosen, how uncertain cases are handled and how errors are measured.

4. Decide whether to merge, link for review or keep separate

A strong system does not force every candidate pair into yes or no. High-confidence matches may be automatically resolved. Low-confidence candidates may remain separate. Borderline cases may be sent for review. In some architectures, records are never physically merged; instead, a canonical entity identifier links them while preserving the source records. This can be safer for auditability because the bank can reconstruct which systems contributed which facts.

The decision should also be reversible. If later evidence proves that two customers were incorrectly joined, the bank needs a controlled way to split the entity and assess downstream impact. That means identifying which alerts, risk scores, cases, sanctions decisions or reports used the incorrect link.

Entity resolution should distinguish high-confidence matches, uncertain candidates requiring review and records that remain separate; only after identity is resolved should the bank treat downstream network links as belonging to the same subject.

False merges and missed matches: both are financial-crime risks

Entity resolution creates two fundamental error types.

A false merge occurs when different real-world subjects are treated as the same entity. This can be severe. An innocent customer might inherit another customer's adverse-media association, sanctions alert, high-risk network or fraud history. A case could pull unrelated transactions into an investigation. Customer restrictions may be applied unfairly. Model feedback can then reinforce the error because future training data assumes the joined identity was correct.

A missed match occurs when records belonging to the same subject remain separate. The risk is different. The bank may fail to aggregate exposure, miss repeated onboarding attempts, underestimate network centrality, fail to see total transaction activity or overlook an ownership chain.

For that reason, overall match accuracy is not enough. Testing should consider precision and recall, but operational severity matters too. A small number of false merges involving sanctions or account closure can be more damaging than many low-risk duplicate records. Thresholds may therefore differ by use case. A marketing system might accept a looser household match than a sanctions decisioning system. Financial-crime architecture should not reuse a single enterprise identity score blindly across every control.

Building a graph that can be defended later

A financial-crime graph is useful only when its relationships are evidence-backed. The graph should not simply contain A connected to B. A defensible edge normally includes at least:

  • the relationship type;
  • the source system or authoritative source;
  • source record identifier;
  • when the relationship was observed;
  • when it was effective, if known;
  • whether it was customer-declared, independently verified, inferred or analytically derived;
  • confidence or quality where the relationship is inferred;
  • who or what process created the edge; and
  • version or model information if analytics produced it.

This matters because relationships change. Directors resign. Beneficial owners sell shares. Devices change. Accounts close. Addresses are reused. A company may have been owned by a person on the date of a transaction but not today. The bank needs a temporal graph, or at least historical relationship versions, so an investigator can reconstruct what was known and true at the relevant time.

Point-in-time reconstruction is particularly important in sanctions and ownership analysis. A current corporate registry cannot automatically prove historic ownership. Likewise, a vendor's current entity label does not prove what its data showed when a payment was released six months earlier. Case evidence should preserve the relevant source and timestamp.

Beneficial ownership and control: graph structure helps, but law decides

Ownership graphs are one of the most valuable uses of entity resolution because multi-layer structures are difficult to understand in flat tables. FATF's 2023 Recommendation 24 guidance focuses on transparency of legal persons and the need for adequate, accurate and up-to-date beneficial ownership information. FATF's 2024 Recommendation 25 guidance performs a similar role for legal arrangements such as trusts. A bank can use graph modelling to represent shareholders, controllers, trustees, settlors, beneficiaries, directors and intermediate entities while retaining the different legal meanings of those roles.

The graph must not invent a universal definition of beneficial ownership. Jurisdictions define ownership and control differently, and the same relationship may matter differently for CDD, tax, prudential reporting and sanctions. The EU's Anti-Money Laundering Regulation, Regulation (EU) 2024/1624, contains detailed rules for multi-layered ownership and control structures and is part of the EU's new AML framework. Those rules should be applied in their legal scope and effective timeline, not exported as a global algorithm.

The United States provides another useful caution. OFAC's 50 Percent Rule is a U.S. sanctions rule and concerns ownership by blocked persons. OFAC states that ownership interests of blocked persons are aggregated and that indirect ownership can matter. OFAC also states that control without 50 percent or more ownership does not by itself automatically block an entity under that specific rule, although other sanctions provisions can still apply. A global bank must therefore store ownership and control facts separately and let the applicable sanctions regime determine their legal consequence. A graph engine should not hard-code “control = blocked” or “25% = beneficial owner” as universal truth.

The same principle applies to U.S. CDD. FinCEN's February 2026 exceptive relief changed when covered financial institutions must identify and verify beneficial owners for repeated account openings, while ongoing risk-based CDD remains. That is current U.S.-specific regulation, not a reason to rewrite global KYC policy everywhere.

Relationship data beyond ownership

Ownership is only one type of network. Banks can gain useful insight from other relationships when they are governed carefully.

Accounts and beneficiaries. A receiving account that appears across many unrelated scam victims can become a high-value investigative node. Fan-in, rapid onward transfer and repeated links to already-confirmed mule accounts may strengthen the case for review.

Devices and sessions. Several customers using one device can indicate account control by one operator, but shared devices are common in families, small businesses and assisted-service contexts. Device linkage should therefore be contextual and should not become identity proof without corroboration.

Contact information. Reused telephone numbers, email addresses and postal addresses can connect customers, but data quality is critical. Corporate service providers, accountants and registered offices legitimately serve many companies. An address can be a useful lead while still being a weak ownership fact.

Corporate roles. Directors, authorised signatories, company secretaries, trustees and agents can reveal clusters. The same professional intermediary appearing in hundreds of companies may be entirely legitimate; risk depends on the wider context, business model and activity.

Payments. Transaction edges can show flows, circular movement, common counterparties, rapid dispersal and convergence. A payment link demonstrates a transaction, not a personal relationship. Repeated patterns and surrounding facts determine its significance.

Digital assets. Wallet addresses, VASP relationships and blockchain-analytics labels can connect fiat and virtual-asset investigations. This chapter does not replace dedicated blockchain tracing. The important network principle is to separate customer identity, wallet attribution, service-provider labels and transaction-path evidence. Vendor attribution should include provider, confidence and observation date because labels can change.

Legal-entity identifiers. Where relevant, the Legal Entity Identifier can help with entity identity. GLEIF distinguishes Level 1 data — broadly “who is who” — from Level 2 relationship data that includes direct and ultimate accounting consolidating parents. That data is useful, but it is not a complete beneficial-ownership register and should not be treated as one. Its relationship definitions and reporting exceptions have to be understood before using it in AML or sanctions logic.

What graph analytics can tell an investigator

Once entities and relationships are modelled, graph techniques can help investigators navigate large populations. The terms can sound mathematical, but the operational ideas are straightforward.

A connected component is a group of nodes linked to one another. It can help identify a cluster of customers, accounts and counterparties that would otherwise appear separate. The component is not automatically a criminal network; it is simply a connected population.

Degree measures how many direct links a node has. A beneficiary receiving from hundreds of customers may have high degree. So might a utility company, payroll processor or major merchant. High degree therefore needs contextual comparison.

Centrality methods try to identify nodes that occupy structurally important positions. A person or account that connects several sub-groups may be operationally important to investigate. But centrality does not prove leadership or criminal control.

Community detection groups nodes that are more densely connected to each other than to the rest of the graph. This can help find mule rings or related corporate clusters, but results depend on data, algorithm and parameters. The group boundaries should be treated as analytical hypotheses.

Path analysis finds how two subjects are connected. This is valuable when a screened party, suspicious beneficiary or high-risk entity is several relationships away from the customer. The investigator should be able to see every edge on the path and the evidence supporting it.

Motifs and patterns identify recurring relationship shapes. Examples include many senders converging on one account, one account rapidly dispersing to many others, circular flows, common-device clusters or repeated corporate-director patterns. A useful detection model combines these shapes with amount, time, customer profile and known typologies rather than treating the shape alone as suspicious.

Temporal analytics looks at how networks form and change. Ten accounts created over two years are different from ten accounts opened within forty-eight hours and immediately used in a coordinated flow. The timing of edge creation can be as informative as the relationship itself.

From signal to case: the investigator still has to prove the story

The purpose of network analytics is not to replace investigation. It is to create better questions and prioritise evidence.

A network-driven alert should tell the investigator why the cluster was surfaced. It should identify the entities and edges that contributed materially, the data sources, time window and relevant scores. The investigator should then test alternative explanations. Are the customers members of one household? Is the shared address a corporate-service provider? Is the common device a legitimate business terminal? Does the payment beneficiary operate a marketplace? Is the supposed ownership link current and verified?

The case should distinguish observed facts from analytical inferences. “Accounts A, B and C logged in from Device X” is an observed or system-derived fact if the device data is reliable. “The same person controls all three accounts” is an inference. “The accounts are being used to launder criminal proceeds” is a much stronger conclusion requiring additional evidence and, in many jurisdictions, a suspicion decision under local law and policy.

If the investigated facts create suspicion, the SAR or STR process follows the applicable jurisdiction's rules. If the graph exposes a sanctions ownership issue, the sanctions decision process follows the applicable regime. If the network suggests fraud, customer-protection and fraud operations may need immediate action. These outcomes can overlap, but they should remain separately governed.

The technology stack should preserve source records and lineage, resolve entities with measurable confidence, build versioned relationship graphs, run explainable network analytics, pass evidence to case management and feed reviewed outcomes back into controlled tuning.

Privacy, data protection and the temptation to collect everything

Graph analytics becomes more powerful as more data is connected. That creates an obvious governance risk: a technically useful link may not be legally permitted or proportionate to use for every purpose.

FATF's work on data pooling and collaborative analytics recognises the value of combining information to reveal activity that no institution can see alone, while also stressing that information sharing must comply with applicable data-protection and privacy frameworks. FATF's July 2026 report on public-private partnerships similarly emphasises a robust legal basis, governance and secure information exchange. These principles matter even inside one banking group. Cross-border transfers, local secrecy rules, purpose limitation, access controls and data-retention rules can constrain what a global analytics platform may centralise or expose.

The design question is therefore not “what data can we technically ingest?” but “what data are we authorised to process for this use, for how long, in which jurisdiction, and who may see it?” A global graph may need logical or physical segmentation. Some attributes may be available only as privacy-preserving features. Some jurisdictions may permit collaborative analytics but not raw-record sharing. Legal and privacy teams should be involved before deployment, not after the graph has become operationally indispensable.

A further fairness issue is guilt by association. Network tools make connections visually persuasive. A dense graph can look incriminating even when the edges are weak. Investigators need training to assess evidential strength, and user interfaces should expose edge type and confidence instead of showing every connection with the same visual weight.

Data quality is a control, not housekeeping

Network analytics magnifies both good and bad data. One wrong passport number can merge unrelated people. A stale ownership record can create an incorrect sanctions path. A reused phone number can link hundreds of unrelated customers. Missing source timestamps can make historic reconstruction impossible.

For that reason, the graph platform should capture lineage and support quality metrics. Teams should know which source systems contribute records, how often they refresh, where transformations occur and which fields are considered authoritative for which entity type. Data-quality issues should feed operational queues with owners and service levels, not disappear into model-performance discussions.

A bank should also distinguish unknown from negative. If no beneficial owner edge is present, that may mean no owner exists above the relevant threshold, or it may mean the source does not contain ownership data. If an LEI parent is absent, GLEIF reporting exceptions may explain why. If no device link exists, the channel may not collect device information. Absence of an edge should never be interpreted automatically as evidence that no relationship exists.

Designing a useful investigator experience

Graph tools can fail even when the analytics are technically strong. An investigator faced with a thousand-node “hairball” cannot work efficiently. The user interface should start with the alerting subgraph — the smallest set of nodes and relationships that explains why the case was raised — and allow controlled expansion.

Useful functions include filtering by relationship type, date and confidence; distinguishing verified from inferred edges; showing source provenance; collapsing low-value service-provider nodes; highlighting material paths; comparing the network at two dates; and exporting a case-ready evidence view. Search should support identifiers, aliases and local scripts where available. Accessibility and mobile or small-screen behaviour matter because the diagram or graph must remain interpretable outside a large analyst workstation.

The case record should preserve the relevant network snapshot or evidence references rather than relying on a live graph that may change. A future reviewer must be able to reconstruct what the investigator saw when the decision was made.

Network evidence is time-dependent: customer identity, ownership, device use and transaction relationships can appear, change or end; the case must preserve what was known and effective at the decision date rather than substituting today's graph.

Governance: who owns the truth?

No single team owns every network fact. Customer data may be owned by onboarding or product teams. Corporate ownership may come from KYC operations and external registries. Device data may be owned by fraud technology. Payment data comes from multiple rails. Sanctions policy defines legal application. Data engineering builds pipelines. Model or analytics teams develop resolution and graph logic. Investigators consume the output. Privacy and legal teams constrain use. Internal audit challenges the end-to-end control.

A workable governance model therefore assigns ownership by layer.

Data owners are responsible for source quality and meaning. The entity-resolution owner is responsible for matching logic, thresholds, performance and remediation of false merges. The graph platform owner is responsible for relationship modelling, lineage and availability. Financial-crime control owners decide how graph signals are used in monitoring and cases. Sanctions or legal teams own regime-specific interpretation. Model-risk or validation functions provide independent challenge where analytics fall within their scope. Operations own human-review procedures and evidence standards.

Changes require control. A new device feature, matching algorithm or external vendor can change thousands of relationships overnight. Release governance should include impact analysis, backtesting, versioning, approval and a rollback plan. If a threshold is changed, the bank should understand which populations are newly merged or separated and which downstream controls will be affected.

Management information should measure outcomes, not just graph size. Useful measures include unresolved candidate volumes, false-merge rates, missed-match findings, human-review disagreement, alert uplift attributable to network features, confirmed network cases, customer-impact errors, data-quality ageing, source coverage and model drift. Metrics should be segmented where relevant because performance can vary by language, geography, customer type and data richness.

Governance keeps identity confidence, relationship confidence, financial-crime risk and legal conclusion separate, with clear owners for source data, resolution logic, graph analytics, investigation, sanctions interpretation, privacy and independent validation.

What good looks like

A mature network-analytics capability does not produce the largest graph or the most alerts. It produces more defensible understanding.

The bank can explain why records were linked. It can reconstruct the relationship as it existed at the relevant time. It distinguishes verified ownership from inferred association. It uses graph patterns to prioritise investigations without treating connectivity as guilt. It routes legal questions to the correct jurisdictional framework. It preserves source lineage. It measures false merges and missed matches. It controls access to sensitive network data. It gives investigators a usable view rather than a visual spectacle. It learns from reviewed cases without converting every analyst judgement into unquestioned training truth.

For a business analyst, architect or product owner, the core design principle is simple: make every important node and edge explainable enough to survive challenge. For an investigator, the principle is equally simple: use the network to ask better questions, then prove the answer with evidence.

The next sections deepen that operating model through architecture, testing, a realistic case and implementation controls.

Operational deep dive: making relationship evidence usable

The base chapter separates identity, relationships, risk signals and legal conclusions. This section goes deeper into the operating mechanics: how the bank builds the graph, how an investigator interprets it and how the institution avoids turning a useful analytical tool into a source of false certainty.

A graph model should start with business meaning

Graph technology is attractive because it can connect data that is awkward to represent in flat tables. The technology is not the starting point, however. The starting point is the relationship vocabulary the bank needs for financial-crime decisions.

For natural persons, the graph may need roles such as customer, beneficial owner, director, trustee, settlor, beneficiary, authorised signatory, payee and payer. For legal entities it may need direct shareholder, indirect shareholder, parent, subsidiary, branch, fund manager, merchant, respondent bank or service provider. Operational objects such as accounts, cards, devices, addresses and wallet addresses need their own entity types because they are not people or companies even when investigators use them to infer control.

This distinction prevents a common design failure: storing every connection as the same generic edge. A payment edge answers “value moved between these endpoints.” An ownership edge answers “this source reported an ownership relationship.” A device edge answers “these sessions were observed from the same or related device identifier.” They can support one investigation, but they are different evidence and should remain different in the data model.

A well-designed relationship record therefore carries context. At minimum, the investigator should be able to determine the two endpoints, relationship type, source, observation or effective date, confidence where inferred, and whether the link is verified, declared or analytically derived. If a vendor supplied the relationship, the provider and version should be preserved. If an analyst confirmed it, the case or review record should be traceable.

Canonical identity without destroying source truth

Many banks create a canonical customer or party identifier so records from several systems can be viewed together. That is useful, but a canonical identifier should not erase the differences between source records.

Suppose a corporate customer is represented in the onboarding platform, core banking system and trade-finance platform. One system contains the registered name, another still holds a former name and the third contains an abbreviated payment name. The entity-resolution service can associate the records with one canonical legal entity, but the graph should retain the source-system identifiers and historic values. When an investigator asks why a payment screened differently from the KYC record, the source-specific values may explain the difference.

The same principle applies when two records are probabilistically matched. Rather than overwriting one customer with another, the platform can preserve both records and store the resolution decision. This makes later split or correction possible. It also allows the bank to distinguish “two source records resolved to one entity” from “one source system manually merged the customer records,” which may have different audit implications.

Deterministic and probabilistic matching belong together

Deterministic matching works well where reliable identifiers exist. Two legal-entity records with the same verified company-registration number in the same jurisdiction are strong candidates for resolution. Two natural-person records with the same high-quality government identifier may be similarly strong, subject to local data-handling rules and known identifier quality.

Probabilistic matching is valuable when identifiers are absent or inconsistent. It can combine name similarity, date of birth, address, contact details and other attributes to estimate whether records refer to the same subject. The system should expose the evidence contributing to the result. Investigators should not be asked to trust a score without understanding what drove it.

The bank should also recognise negative evidence. A close name match with materially different dates of birth and distinct verified identifiers may be evidence that records should stay separate. Poor resolution engines focus only on similarities and underuse contradictions.

For multilingual populations, transliteration needs special testing. Arabic, Cyrillic, Chinese and other scripts can create multiple valid representations. Name order can differ. Patronymics and compound family names can be handled inconsistently. A model trained mainly on one language may perform unevenly in another. Performance testing should therefore be segmented by relevant language and customer population rather than reporting one global accuracy figure.

Ownership and control graphs need effective dates

Corporate structures are not static. A shareholder can sell an interest, a trust can change trustees, a director can resign and a company can be acquired. Network analysis used for customer risk, sanctions or investigation must be able to answer when a relationship applied.

A useful ownership edge contains an effective-from date and, where known, effective-to date. It also records when the bank learned the fact. These are not always the same. A registry filing might state that a director changed on 1 March, while the bank receives the updated record on 15 March. For a payment released on 10 March, the difference matters operationally.

That leads to two distinct questions in retrospective review:

  • What was the real-world relationship at the transaction date, based on evidence now available?
  • What did the bank know, or reasonably have available, when it made the original decision?

Both can be important. The first helps assess current risk and suspicious activity. The second helps assess whether the control operated as designed at the time.

Do not turn graph traversal into a universal ownership calculator

It is tempting to multiply percentages through every corporate chain and treat the result as a universal beneficial-ownership answer. That can be useful in specific legal frameworks, but it is not universally correct.

FATF Recommendations 24 and 25 establish global expectations for beneficial ownership transparency, while domestic or supranational law defines how institutions identify and verify beneficial owners in their jurisdiction. The EU AML Regulation contains detailed multi-layer ownership and control provisions. OFAC's 50 Percent Rule uses a different U.S. sanctions concept, including aggregation of ownership by blocked persons and specific treatment of indirect ownership. Other sanctions regimes have their own ownership and control tests.

A robust architecture therefore separates the relationship graph from the legal rules engine. The graph says who owns what, by how much, through which chain, according to what source and when. The rules engine applies the relevant regime. This makes policy updates safer: when law changes, the bank can change the interpretation without rebuilding historical relationship evidence.

Graph analytics should surface hypotheses, not verdicts

Network algorithms can rank or group entities, but their outputs should be phrased as hypotheses.

A community-detection algorithm may identify twelve customers as a tightly connected cluster. The investigator still needs to ask what the links mean. If the cluster is based on payroll payments from one employer, it may be entirely legitimate. If the cluster is based on shared devices, rapid transfers and common beneficiaries immediately after onboarding, it may be materially more concerning.

Likewise, a “central” node can be central because it is a legitimate service provider. Payment processors, marketplace operators, utility companies and large merchants naturally connect many customers. The system should therefore support suppression, categorisation or contextual treatment of known high-degree legitimate nodes so investigators are not overwhelmed by obvious infrastructure.

A useful alert explains the network feature that mattered: for example, “five newly opened personal accounts used two devices and sent 83% of received funds to the same two beneficiaries within six hours.” That is more actionable than “network risk score 87.”

From alert to investigation

Consider a network-driven alert on a corporate customer. The customer itself has no unusual transaction volume. The alert is generated because the company shares a director and registered office with three entities that have recently received suspicious payments.

The investigator should first validate resolution. Is the director really the same person? Is the address a corporate-services office used by hundreds of clients? Are the three related companies active in the same legitimate group? Only after those questions should the analyst interpret the network.

Next comes activity. Do the companies transact with one another? Are payments circular? Are counterparties consistent with stated business? Are funds rapidly transferred outside the group? Did the shared director have authority when the transactions occurred? Does beneficial ownership overlap, or is the director a professional nominee or service provider?

If suspicion remains, the case narrative should say what was observed and why the combination is concerning. It should not rely on a graph screenshot as the conclusion. A useful narrative might state that four companies with common control indicators were incorporated within a short period, used the same contact details, received funds from unrelated third parties and transferred most value to two common overseas beneficiaries with no clear economic rationale. Each statement should be traceable to evidence.

If local law requires an STR or SAR, the filing route follows that jurisdiction. If the concern relates to a sanctions ownership chain, sanctions specialists apply the relevant regime. If the network shows scam proceeds and active mule accounts, fraud operations may need to intervene before the AML investigation is complete. Graph analysis can support all three outcomes without merging their decision standards.

Collaborative analytics changes the legal boundary

Network value increases when institutions can see beyond their own customer base. FATF's work on data pooling and its July 2026 review of public-private partnerships shows why information sharing can improve detection of complex, cross-institution financial crime. The same work stresses legal basis, governance, data protection and secure exchange.

A bank should never assume that because two institutions can technically match customers, they may freely exchange raw customer data. Permitted sharing varies by jurisdiction, purpose and partnership model. Some arrangements share typologies. Some allow operational case information under specific legal gateways. Some use privacy-enhancing technology to compare or analyse data without centralising all raw records.

For architects, the important requirement is policy-aware data use. Every external relationship signal should identify its legal or contractual source and any restriction on onward use. Investigators need to know whether the information can be included in a customer communication, shared with another group entity or cited in a regulatory filing.

The evidence file should survive a later challenge

A defensible network case should allow a reviewer to reproduce the material path without rerunning the entire platform. That means preserving the relevant entity-resolution decisions, relationship evidence, graph snapshot or relationship list, model or rule version, alert rationale and analyst conclusions.

If a vendor label or external registry record is material, preserve the value and date accessed. If the graph used a machine-generated confidence score, preserve the score and model version. If an analyst overrode a match, preserve the reason. If the customer was restricted, document which evidence justified the action.

The question for quality assurance is straightforward: could an independent reviewer, months later, understand why the bank believed these entities were the same, why these relationships mattered and why the final outcome followed? If not, the graph may be visually impressive but operationally weak.

Advanced practice: architecture, requirements and validation

Network analytics often fails because the bank buys graph technology before agreeing what an entity, relationship or defensible match means. This section turns the chapter into delivery requirements that architects, business analysts, developers, testers, compliance owners and model-governance teams can use.

Reference architecture

A production capability normally contains several logical layers even if one vendor provides more than one of them.

The source layer contains customer master data, KYC and KYB records, accounts, payments, cards, merchant data, corporate ownership, screening data, devices, fraud signals, trade information and, where relevant, virtual-asset data. Each source should expose stable record identifiers and change timestamps. If a system cannot provide those, downstream lineage becomes fragile.

The standardisation layer prepares values for comparison while preserving originals. It can normalise names, parse addresses, map country codes, standardise company identifiers and handle transliteration. Every transformation should be versioned because a change to name-normalisation logic can alter match results across millions of records.

The resolution layer generates candidates, compares attributes and assigns canonical entity identities or relationship hypotheses. It should support both automatic decisions and human-review queues. Its output is not merely a master identifier; it should include match evidence, confidence, method, version and history.

The relationship layer stores verified and inferred edges with provenance and effective dates. Some banks use a graph database; others use relational or analytical platforms that expose graph functions. The technology choice matters less than whether the data model supports time, provenance, edge types and controlled history.

The analytics layer calculates network features, patterns and risk signals. It may contain rules, statistical models, graph algorithms or machine learning. These outputs should be treated as features or hypotheses, not legal facts.

The case and decision layer sends material evidence to alert and case management, captures investigator decisions and records downstream outcomes such as monitoring, KYC review, fraud action, sanctions escalation or SAR/STR consideration. Reviewed outcomes can feed controlled tuning, but only after data-quality and label-quality checks.

Requirements that are specific enough to build

A weak requirement says, “The solution shall identify related customers.” A buildable requirement describes how the relationship is established and what happens when confidence is uncertain.

For example: when two natural-person records share a verified government identifier but have materially inconsistent dates of birth, the service must not automatically merge them; it must route the candidate to an exception rule or review queue and retain both source values. The acceptance test then becomes clear.

Another example: when a network alert relies on an inferred relationship, the case payload must include the relationship type, endpoints, source records, confidence, observation date, algorithm or rule version and the features that materially contributed to the inference. This prevents case investigators from receiving an unexplained graph score.

A third requirement concerns history: when an ownership relationship changes, the platform must close the prior edge with an effective end date rather than overwrite it. Investigators must be able to reconstruct the network as of a specified date.

Business analysts should define similar requirements for manual overrides, split and merge correction, source precedence, relationship suppression, privacy restrictions, cross-border data access and evidence export.

Testing entity resolution requires ground truth

Resolution testing is difficult because production data often does not contain perfect truth. A bank should therefore build a labelled test set using verified examples, synthetic cases and adjudicated production samples.

Pairwise tests answer whether two records were correctly classified as same or different. Cluster tests are harder and often more valuable because one wrong merge can connect several records into one entity. Testers should check whether the final cluster contains all and only the records that belong together.

Metrics need context. Precision measures how often declared matches are correct; recall measures how many true matches were found. A sanctions-related use case may prioritise avoidance of false merges differently from a low-impact duplicate-cleanup use case. The bank should establish tolerances based on downstream harm rather than chasing one abstract accuracy target.

Negative tests are essential. They should include common names, twins, parent and child with similar names, reused phone numbers, shared family addresses, corporate service-provider addresses, multiple subsidiaries with near-identical names and unrelated companies with the same trading name in different jurisdictions. A model that performs well only on obvious positive matches is not safe for financial-crime decisions.

Multilingual and transliteration testing should use the populations the bank actually serves. Test sets should include missing fields, reordered names, accented characters, local scripts and legacy encodings. Where legally permissible, performance should be monitored for segments in which data quality or naming conventions differ materially.

Testing the graph, not only the matcher

Even perfect entity resolution can feed a poor network model. Edge tests should verify that ownership, control, payment, device and contact relationships are created from the correct sources and with the correct direction. An ownership edge from Entity A to Person B is not equivalent to Person B owning Entity A unless the relationship semantics and direction are explicit.

Temporal tests should change relationships over time and confirm that historic queries return the correct state. A director who resigned before the alert period should not appear as current. A sanctions ownership calculation should use the relationship facts effective at the relevant time and the correct legal rule version.

Graph-algorithm tests should use networks with known structures. Community detection should be challenged with legitimate dense networks such as payroll populations. Centrality should be tested against utilities, marketplaces and payment processors so high-volume infrastructure does not automatically dominate risk ranking. Path analysis should show the exact source-backed edges rather than only the number of hops.

Model and change governance

If probabilistic matching, machine learning or graph scoring falls within the bank's model-risk framework, it should be inventoried and independently challenged accordingly. Even where the technology is classified as a rules engine rather than a model, material changes still need governance because the customer impact can be significant.

Validation should examine conceptual soundness, data suitability, performance, stability, explainability and downstream use. A strong validation asks not only “does the matcher perform?” but “what happens when it is wrong?” That means tracing false merges into screening, risk rating, monitoring and case outcomes.

Changes to thresholds, features or source data should be backtested before release. The bank should compare the old and new entity population, identify newly merged and newly split clusters, measure alert impact and sample high-risk changes. Rollback capability matters because a bad resolution release can contaminate many downstream controls rapidly.

Vendor governance is equally important. If an external provider supplies entity links, corporate hierarchies, device intelligence or blockchain attribution, the bank should understand coverage, refresh frequency, confidence methodology, known limitations and change notification. Vendor output should not enter the graph as unquestioned truth.

Privacy and access-control tests

Network platforms can expose sensitive associations to users who never had access to all source systems. Role-based access must therefore be tested from the graph outward. A user who may see a customer name should not automatically gain access to restricted fraud intelligence, employee data or another jurisdiction's protected customer information because a graph edge connects them.

Testers should verify masking, purpose-based access, jurisdictional segmentation, export restrictions and case-sharing controls. Logs should show who viewed or expanded sensitive relationships. Data-retention rules should apply to derived relationships as well as source records where required.

Operational resilience

Resolution and graph services can become critical dependencies for onboarding, screening and monitoring. Failure modes need explicit design.

If the graph is unavailable, does payment screening fall back to direct-list matching? Does onboarding continue without relationship enrichment? Are alerts queued for later graph evaluation? Which services can safely operate in degraded mode, and which must stop? The answer depends on the control and applicable obligations, but it should be predetermined and tested.

Recovery also needs consistency. Replaying missed events must not create duplicate edges or lose effective dates. After a data-source outage, the bank should reconcile expected versus received records and assess whether risk decisions made during the outage require retrospective review.

A production-ready network capability is therefore not just an algorithm. It is a controlled chain from source data through resolution, relationships, analytics, human judgement and evidence preservation, with testing at every boundary.

Practice close: what to challenge before approving the control

A network-analytics implementation should be approved only when the bank can explain its identity decisions, relationship evidence and downstream use. The following questions are useful during requirements, design review, testing, model validation and operational QA.

For entity resolution, ask what evidence can automatically join records, what contradictions prevent a merge, which cases require human review and how a false merge is reversed. Confirm that source records remain traceable and that the same matching threshold is not reused blindly across customer deduplication, AML monitoring and sanctions decisions.

For relationships, ask whether each edge has a clear type, source, date and confidence. Verify that verified ownership is distinguishable from inferred association, that a payment edge cannot be mistaken for ownership, and that shared addresses, devices or contact details are not treated as proof of common control without corroboration.

For time, require the platform to preserve effective dates and observation dates. Test whether an investigator can reconstruct a customer's network on the date of an alert and whether a later ownership update incorrectly rewrites history.

For analytics, require the alert or case to explain which network features mattered. Test legitimate dense networks such as payroll, marketplaces, households and corporate-service providers. A useful control should reduce unexplained noise rather than simply create more connected alerts.

For legal interpretation, keep the graph factual. FATF beneficial-ownership standards, EU AML rules, OFAC ownership rules and other national frameworks are not interchangeable. The system should expose ownership and control evidence to the appropriate policy or legal decision layer rather than embedding one jurisdiction's threshold as global truth.

For privacy, confirm that every data source has an approved purpose, access model and retention rule. Test cross-border access, exports and derived relationships. A graph can accidentally expose information from a restricted source system to a much wider population unless access controls are applied at relationship and attribute level.

For operations, make sure investigators can start with a small alerting subgraph, expand relationships deliberately, view provenance and record why they accepted or rejected an inferred link. The case record should preserve the material evidence used for the decision rather than depend on a live graph that will change later.

Acceptance criteria for a production release

A strong release should demonstrate that known same-entity records resolve correctly, known different entities remain separate, ambiguous cases follow an approved review path, graph edges preserve provenance and time, and all five diagrammed control stages work end to end from source data to case evidence. Regression testing should compare cluster changes before and after a release and sample high-risk merges or splits.

The bank should also demonstrate failure behaviour. When a source feed is late, a vendor is unavailable or a graph service fails, the control should have a documented degraded mode and recovery process. Replayed events must not create duplicate relationships. Decisions made during an outage should be identifiable for retrospective assessment where required.

Knowledge check

Why is a shared device not enough to merge two customers? Because device use is a relationship signal, not a verified identity fact. Families, businesses and assisted channels can legitimately share devices.

Why store both effective date and observation date? Because the relationship may have changed before the bank learned about it; investigation and control-performance review may need both perspectives.

Why should sanctions ownership logic sit outside generic graph traversal? Because ownership and control consequences differ by regime. The graph should supply facts; the applicable legal framework supplies the legal result.

What is the most dangerous visual mistake in graph investigations? Showing weak and strong edges as if they have equal evidential value. The user must be able to distinguish verified, declared and inferred relationships.

What makes network analytics effective rather than merely sophisticated? It helps the bank identify material risk earlier, gives investigators explainable evidence, reduces avoidable noise, preserves customer fairness and produces outcomes that can be defended to independent reviewers.

Masterclass: when six customers become one investigation

This synthetic case shows how network evidence changes an investigation without allowing the graph to decide guilt. Names and figures are illustrative.

A retail bank receives separate transaction-monitoring alerts on six personal customers. Each account has been open for less than four months. Individually, the accounts do not cross the bank's highest-risk transaction thresholds. Each receives multiple credits from unrelated domestic customers and sends most of the value onward within the same day.

The first investigator sees an ordinary rapid-movement alert. The network view changes the question.

Entity resolution shows that the six customers are genuinely different natural persons: different verified identity documents, dates of birth and addresses. They must not be merged. Relationship analysis, however, identifies that four of the customers have logged in repeatedly from the same two devices. Five have paid the same beneficiary company. Three used the same telephone number at different stages of onboarding, although the number was later replaced. The beneficiary company is newly incorporated and its sole director is also an authorised signatory on a second company receiving payments from two of the accounts.

At this point the network is interesting, not conclusive. A shared device can have innocent explanations. A repeated beneficiary may be a legitimate merchant. Reused contact details can result from family assistance or an introducer.

The investigator expands the evidence carefully. Session history shows that the two shared devices accessed the accounts within minutes of one another from the same network location, often immediately before onward payments. The customers' declared occupations and addresses do not indicate a household or employer relationship. Customer contact confirms that two account holders say a “friend” helped them set up mobile banking and asked them to receive money temporarily. One customer reports being paid a small fee. The beneficiary company's registered address is a corporate-service provider, which is therefore treated as weak evidence. Its bank activity, however, shows immediate transfers to a virtual-asset service provider and to another recently incorporated company.

Corporate data reveals that the second company has a different director but shares an ultimate owner according to independently obtained KYC evidence. The bank does not infer this merely from the common service-provider address. It records the ownership source, effective date and verification status.

A sanctions-screening enrichment then produces a further issue: a minority shareholder several layers above the second company has a name similar to a designated person. The graph does not mark the entire network as sanctioned. Sanctions specialists first resolve the shareholder's identity and then apply the ownership and control rules of the relevant regime. The apparent name match is ultimately cleared as a different person. Keeping the sanctions decision separate prevents a behavioural AML case from being contaminated by a false legal conclusion.

The AML investigation continues on its own evidence. The combined pattern now shows recruited personal accounts, common operational control signals, rapid pass-through movement, common beneficiaries and conversion of value through a VASP. The bank's local policy threshold for suspicion is met after investigator review. The relevant accounts and entities are linked in one case, while each customer remains a separate resolved person. The institution follows local SAR/STR requirements and fraud teams assess whether any customers are victims or recruited mules requiring different treatment.

The case also produces control feedback. Investigators discover that the original account-level scenarios created six alerts with no automatic connection. The network feature would have allowed earlier prioritisation. Product teams add a controlled feature that identifies combinations of shared devices, first-month rapid movement and common beneficiaries. They do not create a rule that any shared device equals a mule network.

Three lessons matter. First, entity resolution and network analysis solve different problems: the six people were separate entities but part of one operational network. Second, the strongest case came from corroborating independent relationship types rather than trusting one shared identifier. Third, legal outcomes such as sanctions applicability remained separately governed even though the same graph supplied facts to both AML and sanctions teams.

That is the standard a bank should aim for: a network that makes hidden relationships visible while preserving the difference between evidence, inference and legal conclusion.

References and further reading

Global standards and beneficial ownership

Risk-based monitoring and data use

Jurisdiction-specific ownership and control context

Entity and relationship data

These sources provide standards, legal context and control principles. A bank should still use the current law, regulator guidance, sanctions regime, privacy requirements and internal policy applicable to the specific legal entity, customer and transaction before making a legal or reporting decision.