I'm trying to run a dedupe with pandas_dedupe but when running it it gives me this error "AttributeError: 'float' object has no attribute 'keys'"
import pandas as pd
from pandas_dedupe import dedupe_dataframe
df = pd.read_csv('/home/biminsal/Documentos/sampleDataDos.csv', sep=";")
dfinal = dedupe_dataframe(df, ['id', 'nombre_uno','apellido_uno'], canonicalize=True, sample_size=1)
Data sample:
id nombre_uno apellido_uno
0 a001 Karlo Perez
1 a002 Carlos Perez
2 a003 Juan Gomez
3 b001 Carlos Perez
4 b002 Juan Gomez
Error:
AttributeError Traceback (most recent call last)
...
File ~/.local/lib/python3.9/site-packages/dedupe/labeler.py:203, in BlockLearner._sample_indices(self, sample_size)
201 sample_ids = ((keys[i][0], keys[i][1]) for i in sample_indices)
202 else:
--> 203 sample_ids = weight.keys()
205 return sample_ids
AttributeError: 'float' object has no attribute 'keys'
I have tried to find information and have applied several possible fixes, but the same problem remains. Thanks so much for any help.