Fraud Blocker Data Cleansing: Parsers » Match Data Pro

Data Cleansing: Parsers

Parsing Addresses and Names

Overview

Parsers split a single free-text field into its parts. The Parsers tab has two parsers:

  • Address Parser – takes an address, or an address spread over several columns, and breaks it into house number, street, unit, city, state, and postal code, standardizing each part along the way.
  • Name Parser – takes a full name and breaks it into title, first, middle, last, and suffix, and adds the proper name and likely gender from a nickname lookup.

Parsed columns are exactly what fuzzy matching and deduplication need: matching on a clean street name and house number is far more reliable than matching on a whole address line. Both parsers leave your original columns unchanged and add new columns with a prefix you choose.

Address Parser

  1. Select Datasource.
  2. Select the columns that hold the address. You can pick any combination of Address Column, City Column, State/Province Column, and Postal Code/ZIP Column, and more than one column in each. Pick only an address column if everything is in one field, or all four if your data is already split. At least one column is required.
  3. Enter an Output Column Prefix. Every new column starts with this prefix, for example ParseAddr_road and ParseAddr_city when the prefix is ParseAddr.
  4. Click Save and Add to Task List.

Behind the scenes the parser joins the selected columns into one address, identifies each component, and then cleans it up:

  • Unit designators such as Apt, Suite, Ste, Unit, #, PMB, and PO Box are recognized even when the spacing is odd (“Ste101”, “#APT 4B”) and split into the unit type and the unit number.
  • Street names are standardized: street types become postal abbreviations (Street to st, Avenue to ave, Boulevard to blvd) and directions become single letters (North to n, Southwest to sw). The street is also split into its pre-direction, name, type, and post-direction.
  • State names are converted to their 2-letter codes (California to CA).
  • ZIP+4 codes are split into the 5-digit ZIP and the 4-digit add-on.
  • Common traps are handled: “Highway 101”, cities such as “Colorado Springs” or “Saint Paul”, and lettered streets such as “St D”.

Parsed values are written in lower case, which makes them ideal for matching.

What the Address Parser Produces

With the prefix ParseAddr, the new columns are:

ColumnContents
ParseAddr_house_numberThe building number, e.g. 805
ParseAddr_roadThe full standardized street, e.g. veterans blvd
ParseAddr_road_pre_directionDirection before the street name, e.g. n
ParseAddr_road_nameThe street name only
ParseAddr_road_suffixThe standardized street type, e.g. blvd
ParseAddr_road_post_directionDirection after the street name, e.g. sw
ParseAddr_unitThe unit designator and number as written, e.g. ste 320
ParseAddr_unit_numberThe unit number only, e.g. 320
ParseAddr_po_boxPO Box number when the address is a PO Box
ParseAddr_cityCiudad
ParseAddr_suburbNeighborhood or district, where present
ParseAddr_state2-letter state or province code
ParseAddr_postcodeThe 5-digit ZIP or postal code
ParseAddr_postcode_5The 5-digit ZIP
ParseAddr_postcode_4The ZIP+4 add-on when present

Columns that would be empty for every record are not added.

Name Parser

  1. Select Datasource.
  2. Select Name Columns – one full-name column, or several columns that together make up the name, for example First, Middle, and Last stored separately but inconsistently. Multiple columns are joined in the order selected before parsing.
  3. Enter a Column Prefix.
  4. Click Save and Add to Task List.

The parser understands titles (Dr, Mr, Ms), suffixes (Jr, Sr, III, PhD, Esq), hyphenated and multi-word last names, and names written “Last, First”. It also looks each first name up in a built-in nickname list: “Bill” produces the proper name “William”, and the lookup supplies a likely gender.

What the Name Parser Produces

With the prefix ParseName, the new columns are:

ColumnContents
ParseName_titleDr, Mr, Ms, and similar
ParseName_firstFirst name
ParseName_middleMiddle name or initial
ParseName_lastApellido
ParseName_last2Second last name, for double or compound surnames
ParseName_suffixJr, Sr, II, III, PhD, Esq, and similar
ParseName_properThe formal first name from the nickname lookup, e.g. William for Bill, Victor for Vick; the first name itself when there is no nickname match
ParseName_genderM or F from the nickname lookup, blank when unknown

Columns that would be empty for every record are not added.

Tips

  • Run the parsers before matching. Match on parsed house number, street name, and postcode, or on parsed first and last name, rather than on the raw line.
  • Give the parser everything you have. If city, state, and ZIP are already in separate columns, select them too. The parser uses them to resolve ambiguous street and city names.
  • Use the proper name for matching people. ParseName_proper lets “Bill Smith” and “William Smith” match.
  • Pick distinct prefixes when you run more than one parser on the same data source, so the columns don’t collide.
  • Address Verification for the last mile. Parsing standardizes what’s in your data. The Address Verification module checks it against Loqate’s reference data and corrects it.

Preguntas frecuentes

Addresses from anywhere in the world can be parsed into their components. Street-type, direction, and state standardization follow US and Canadian postal conventions.

No. They add new prefixed columns and leave the originals in place.

The record is kept and the parsed columns are left blank for it.

Yes. The comma form is recognized and the parts are placed in the right columns.

Match Data Pro includes a built-in nickname dictionary that maps common nicknames to proper names and a likely gender.

Comienza tu primer proyecto

Para comenzar, haga clic en el botón Nuevo proyecto desde el panel de control.