I'm using the pandas-dedupe library to sift through a 51,540 record dataframe with 43 columns.
It's taking about an hour to load the file and then cluster it after providing the active learning inputs.
I've tried changing the sample_size, but overall I must be missing something, since this is quite a small dataset comparatively to what I'm seeing other users throw at it on these forums.
Is it too many columns specified for dedupe? I checked other parameters to include, but didn't come up with anything I'd further include.
if __name__ == "__main__":
df = pd.read_csv(R"filepath.txt", sep="\t", encoding="ISO-8859-1")
dedupe = pdd.dedupe_dataframe(
df,
['fname', 'lname', 'company', 'email'],
sample_size=0.05
)
dedupe.to_csv(R"filepath.txt")