Fraud Blocker Senzing Entity Resolution Documentation » Match Data Pro

Senzing Entity Resolution Documentation

Resolving Entities with Senzing

Overview

Entity resolution answers a question ordinary matching cannot: which records, across all of your data sources, describe the same real-world person or organization? Senzing resolves those records into entities, gives each one an Entity ID, and tells you why it made every decision.

The difference from the Fuzzy Matching module is how much you configure. Fuzzy Matching gives you full control over definitions, criteria and thresholds. Senzing is hands-off: you map your columns to Senzing attributes, pick your data sources, and Senzing decides what resolves together using its own algorithms.

What you get back:

  • Entities grouping every record that describes the same person or organization, across all data sources.
  • A Match Key for each decision, showing which attributes drove it, for example +NAME+ADDRESS.
  • A Match Level, such as RESOLVED, saying how strongly the record belongs to the entity.
  • Relationships between entities that are not the same but are connected.
  • Counts by category you can act on: matched, ambiguous, possible, related and singletons.

Senzing sits in your project alongside every other module, so you can import, cleanse and parse first, resolve second, then export the result or pass it downstream.

Step 1: Import Datasources

Open Senzing from your project workflow, or from Senzing Entity Resolution then Mapping in the left menu. The screen opens on step 1 of 2.

Each data source appears as a card showing its name and Record Count. Tick the ones you want to resolve.

  • Add Data Source brings another data source into the Senzing module.
  • Select all ticks everything, and Remove selected clears your choices.
  • Click Next to move to mapping.

Select one data source to find duplicates inside it, or several to resolve records across them. Each must have finished importing first.

Step 2: Column Header Mapping

Senzing needs to know what each column means. Mapping tells it that your “Contact Name” column is a full name and your “DOB” column is a date of birth.

Matching Options

Choose Person to match people or Enterprise to match companies. Your choice changes which Senzing attributes are offered, so set it before you start mapping. Person is the default.

View Options

A dropdown gives you three ways to work through the board:

  • Tiles shows one card per column. Each card carries a Senzing: Ready badge when that column is mapped, the source field with its Senzing attribute in brackets such as Date of Birth (DATE_OF_BIRTH), and a Data Sources Mapped count such as 2/2.
  • Columns shows the Mapping List of Senzing attributes on the left with your data source columns alongside, one dropdown per row. Anything unmapped reads Not Mapped.
  • Rows is the most compact view of the same information.

Filters

Radio filters narrow the board, and a filter appears only when there are columns in that group:

FilterShows
AllEvery column
CompleteMapped to a Senzing attribute and present in every selected data source
PartialMapped to a Senzing attribute but missing from at least one data source
Other MappingsPresent in more than one data source but not yet mapped
Single MappingPresent in only one data source and not yet mapped

Each group is labelled with its count, such as Complete Mapping (11). Use the Filter by column name box to jump to a specific column.

Auto-suggest, Remember and Forget

Match Data Pro proposes mappings by matching your headers against a dictionary of known names plus any you have chosen to remember. After mapping a header by hand, choose Remember so it maps automatically next time. Forget removes it.

What to map

At minimum, map what identifies the entity: a name, then whatever else you have such as address, date of birth, phone, email or an identifier. The more identifying attributes you map, the better Senzing resolves. Unmapped columns are carried through but are not used for matching.

Running Senzing

Click Run Senzing, at the top right of the mapping step.

Senzing loads the mapped records, resolves them and writes a result you can review. Processing runs in the background, so you can leave the page and come back. Time depends on the record count and the number of data sources.

Senzing History

Senzing Entity Resolution then Senzing History lists every run under the heading Senzing Result, with Processed Data Sources, Started Date, Ended Date, Status and Actions.

Available actions:

  • View opens the results.
  • Log appears on a failed run and shows the error, so you can correct the setup and run again.
  • Cancel stops a run that is still going. Cancelling cannot be undone.
  • Delete removes a run and its results.

Reviewing the Results

Click View on a finished run, then choose how to look at it with the two radios at the top:

  • Individual Data Source shows what Senzing found within one data source. This is where you look for duplicates inside a single file.
  • Cross Data Source shows what it found between a pair of data sources, listed as a pair such as NEW PROSPECT RECORDS / MASTER DATA. This is where you look for the same entity appearing in two systems.

Pick the data source, or the pair, from the Data Sources dropdown. Total Records and Total Entities for that selection are shown on the right.

Category tabs

Each tab carries a live count:

TabMeaningAppears in
AllEvery entity in the selectionBoth views
MatchedRecords Senzing resolved together as one entityBoth
SingletonsRecords that resolved to an entity of exactly one record, meaning nothing matched themIndividual only
AmbiguousAn entity that could resolve to more than one entity, where those entities cannot resolve to each otherBoth
PossiblesTwo entities that share high-strength attributes but still cannot resolve togetherBoth
RelationshipEntities that are not the same but are connected, such as a shared addressBoth

Ambiguous and Possibles are worth understanding, because they are where the interesting edge cases live.

An ambiguous match happens when a record could resolve to more than one entity, and those entities cannot resolve to each other. Say you have a Patrick Smith and a Patricia Smith at the same address. If a Pat Smith then arrives at that same address, it could be either one. Since it cannot be known which, it is held apart as ambiguous to both.

A possible match happens when two entities share high-strength attributes, such as an identifier, yet still cannot be resolved together because of other differences. If the same Patrick and Patricia records carried the same driver’s licence number, they would be a possible match to each other. If they were married, one person’s ID may simply have been put on the other’s account by mistake.

The results table

ColumnContents
Entity IDThe identifier Senzing assigned to the resolved entity
Entity NameThe name Senzing chose to represent the entity
Match KeyWhich attributes drove the match, for example +NAME+ADDRESS
Match LevelHow the record relates to the entity, for example RESOLVED
RELATED_ENTITYThe other entity, on a relationship row
Data SourceWhich data source the record came from
FEATURESEvery attribute value Senzing used for that record, listed by attribute name

Records belonging to the same entity share an Entity ID and sit together. The first record acts as the anchor and has no Match Key. Each further record shows the Match Key and Match Level explaining how it joined.

Every column has a search box taking plain text or a regular expression. Show One Record Per Entity collapses each entity to a single row on screen, which turns the view into a deduplicated list. It changes the view only and does not affect what an export contains.

Exporting to a New Data Source

Results are not only for reviewing. Turn any category into a new data source in your project, ready to export or feed another module.

Open Senzing Export Task and set:

  • New Data Source Name for the result.
  • Data Source Options: Select All to include every data source in the run, or Custom Select to choose specific ones.
  • Export Options: choose one category, either Duplicated, Ambiguous, Possibles, Relationship, Singletons or All Data. One category per export task, so run the task again for another slice.

Click Export. Each request is listed in the Export Tasks table with its data sources, date, export type and new data source name.

Tips

  • Cleanse and parse first. Senzing resolves better on standardized data. Run Cleansing, the Parsers and Address Verification before you resolve.
  • Map identifiers wherever you have them. A date of birth or a government identifier does far more for accuracy than another name field.
  • Work the board from Complete outward. Get the columns shared by every data source right before dealing with the rest.
  • Use Singletons as a quality check. A large singleton count usually means a mapping is missing or a column was not standardized, rather than that the records are genuinely unique.
  • Read the Match Key. When a match looks wrong, the Match Key names the attributes that caused it, which usually points at the real problem in the data.
  • Review Ambiguous and Possibles by hand. These are the cases Senzing deliberately held apart, and they are where a human eye is worth most.
  • Export Duplicated for a cleanup list, Singletons for what is already unique. Together they account for your whole file.

Preguntas frecuentes

Fuzzy Matching is configurable: you define match definitions, criteria and how strict each one is. Senzing is hands-off: you map your columns and it decides what resolves, using its own algorithms. Use Fuzzy Matching when you want control, Senzing when you want entity resolution without tuning.

Person matches people. Enterprise matches companies. The choice changes which Senzing attributes are available when you map your columns.

No. Your imported data sources are untouched. Results live in the Senzing run, and anything you export becomes a new data source.

Yes. Select a single data source and use the Individual Data Source view to find duplicates within it.

It names the attributes that drove the decision, so +NAME+ADDRESS means name and address together resolved those records.

A record that ended up as an entity of one, meaning nothing else in your data resolved to it.

That record is the anchor of its entity. The Match Key appears on the records that joined it.

An ambiguous match is a record that could belong to more than one entity, where those entities cannot resolve to each other, such as a Pat Smith who could be either Patrick or Patricia Smith at the same address. A possible match is two entities that share a high-strength attribute such as an identifier but cannot resolve together because of other differences.

Yes. Senzing can be automated as part of Match Data Pro automation, so a configured project resolves on a schedule or on demand without anyone opening the module.

Run a Senzing Export Task for the category you want, which creates a new data source, then use Data Export or pass it to another module.

Comienza tu primer proyecto

Para comenzar, haga clic en el botón Nuevo proyecto desde el panel de control.