python pandas_dedupe loads and clusters file slowly

Viewed 32

I'm using the pandas-dedupe library to sift through a 51,540 record dataframe with 43 columns.

It's taking about an hour to load the file and then cluster it after providing the active learning inputs.

I've tried changing the sample_size, but overall I must be missing something, since this is quite a small dataset comparatively to what I'm seeing other users throw at it on these forums.

Is it too many columns specified for dedupe? I checked other parameters to include, but didn't come up with anything I'd further include.

if __name__ == "__main__":
    df = pd.read_csv(R"filepath.txt", sep="\t", encoding="ISO-8859-1")

    dedupe = pdd.dedupe_dataframe(
        df,
        ['fname', 'lname', 'company', 'email'],
        sample_size=0.05
    )
    dedupe.to_csv(R"filepath.txt")
0 Answers
Related