Description
Provides tokenizers and string similarity measures for Python. It helps data cleaning, entity matching, deduplication, search, and record-linkage workflows compare text values at scale.
Similarity scores are heuristics, not proof of identity. Tune thresholds and review matches before merging records or making user-impacting decisions.