Maximizing Linkage in Address Data: Spatial, Exact, and Fuzzy Matching✩
Pub. online: 10 July 2026
Type: Computing In Data Science
Open Access
✩
Approved for Public Release, Distribution Unlimited. Public Release Case Number 25-2360.
Received
8 September 2025
8 September 2025
Accepted
29 May 2026
29 May 2026
Published
10 July 2026
10 July 2026
Abstract
High quality record linkages are critical for enriching survey data with alternative data sources. To enhance the American Community Survey (ACS) (US Census Bureau, 2025), we evaluated several options to match commercially available property data to housing unit records in the ACS and the Census Master Address File (MAF) (US Census Bureau, 2022) data, with the ultimate goal of supplementing data collected in the survey with the commercial property data. Techniques to match address data come in two flavors: spatial matching and address matching. Spatial matching is done by overlaying commercial boundary shape files on the lat-long coordinates on the MAF to associate Census housing unit records to commercial property parcel records. This method is useful because it does not require matching of text fields, but performs poorly when parcels include many housing units (e.g., large apartment buildings). Address matching, or entity resolution at the address level, links records across the two sources based on the content of various address fields, and offers the possibility of disambiguating multiple matches and matching objects that cannot be successfully assigned a unique match through spatial matching. Several approaches are available for address matching, including rule-based, deterministic, fuzzy, and probabilistic methods. In our article we summarize literature comparing various combinations of spatial and address matching techniques to illustrate the trade-off between linkage rates and linkage quality. We also consider hybrid solutions that leverage spatial matching as well as several types of address matching to maximize high quality linkages for our research and discuss future directions such as incorporating probabilistic matching as an added step.
Supplementary material
Supplementary MaterialZip file contains the original Powerpoint presentation, FCSM 2024 Matching Presentation Draft 10102024 Clean_CB.pptx and Word document describing files and matching process, JDS_matching_data_and_steps.docx.
References
Aldridge RW, Shaji K, Hayward AC, Abubakar I (2015). Accuracy of probabilistic linkage using the enhanced matching system for public health and epidemiological studies. PLoS ONE, 10(8): e0136179. https://doi.org/10.1371/journal.pone.0136179
Binette O, Steorts RC (2022). (almost) all of entity resolution. Science Advances, 8(12): eabi8021. https://doi.org/10.1126/sciadv.abi8021
Cortes TR, Silveira IHd, Junger WL (2021). Improving geocoding matching rates of structured addresses in Rio de Janeiro, Brazil. Cadernos de Saúde Pública, 37: e00039321. https://doi.org/10.1590/0102-311x00039321
Hagger-Johnson G, Harron K, Aldridge R, Fu B, Setakis E, …, Gilbert R (2017). Combining deterministic and probabilistic matching to reduce data linkage errors in hospital administrative data. International Journal of Population Data Science 1(1): 296. https://doi.org/10.23889/ijpds.v1i1.316
Prindle J, Suthar H, Putnam-Hornstein E (2023). An open-source probabilistic record linkage process for records with family-level information: Simulation study and applied analysis. PLoS ONE, 18(10): e0291581. https://doi.org/10.1371/journal.pone.0291581
Tyris J, Dwyer G, Parikh K, Gourishankar A, Patel S (2024). Geocoding and geospatial analysis: Transforming addresses to understand communities and health. Hospital Pediatrics, 14(6): e292–e297. https://doi.org/10.1542/hpeds.2023-007383
Zhu Y, Matsuyama Y, Ohashi Y, Setoguchi S (2015). When to conduct probabilistic linkage vs. deterministic linkage? A simulation study. Journal of Biomedical Informatics, 56: 80–86. https://doi.org/10.1016/j.jbi.2015.05.012