How to perform conditional cross-join in pandas (without performing a cross-join and then later filtering)

Viewed 50

I have 2 dataframes in pandas. These dataframes contain the shared columns that I want to compare in non-equality manner (which rows have column values that are within 5 of each other).

The end goal would be to compare each row of each dataframe (a cross-join) and keep all the rows which meet the difference criteria I have specific above.

If I did a cross-join, I think I might be running out of memory due to the size of the 2 dataframes therefore I wouldn't want to perform the cross-join and then work on the result of that. However, the result of the operation would have very few rows in it, that is most of the time the criteria would be false and we would not keep the pair of rows.

Beyond iterating each over pair of rows with a double for statement, how could this be achieved in pandas. Could it be done in a vectorized manner?

0 Answers
Related