H2OFrame() in Python is adding additional duplicate rows to the Pandas DataFrame- Bug?

Viewed 1293

When converting a Pandas dataframe to a H2O frame using the h2o.H2OFrame() function an error is occurring.

Additional rows are being created in the H2o Frame. When I looked into this, it appears the new rows are duplicates of other rows. Depending on the data size the number of duplicate rows added varies, but typically around 2-10.

Code:

train_h2o = h2o.H2OFrame(python_obj=train_df_complete)

print(train_df_complete.shape[0])
print(train_h2o.nrow)

Output:

3871998
3872000

As you can see here, 2 additional rows have being added. When studied closer there are now 2 rows per user for 2 of the users. I.e. 2 rows have being duplicated.

This appears to be a major bug, does anyone have experience of this problem and is there a way to fix it?

Thanks

3 Answers

I had the same issue, assume your "train_h2o" does not have duplicates, just identify the index of the duplicates in dataframe and remove it. Unfortunately, the h2o Dataframe has limited functionality.

temp_df = train_h2o.as_data_frame()
train_h2o = train_h2o.drop(list(temp_df[temp_df.duplicated()].index), axis=0)

In case your dataset can contain other duplicate rows that do not come from this H2O bug, the proposed solution will drop also those rows. If you want to make sure that you remove only the additional rows added by H2O, this solution might help you out:

temp_df = train_df_complete.copy()
temp_df['__temp_id__'] = np.arange(len(temp_df))
train_h2o = H2OFrame(temp_df)
train_h2o.drop_duplicates(columns=['__temp_id__'], keep='first')
train_h2o = train_h2o.drop('__temp_id__', axis=1)

What I'm doing here is creating a temporary column that I'll then use as ID in order to drop only the duplicates that have been generated by H2OFrame. Once the duplicates have been remove I drop the temporary column. It might not the most elegant way, but it works.

I had the same issue with a specific dataset. Reset index on the base data frame worked for me.

import h2o

train_df_complete = train_df_complete.reset_index()
train_h2o = h2o.H2OFrame(train_df_complete)

I am using h2o 3.30.1.3.

Related