This is an example of a problem I need to solve. I have 2 tables And I have to find a substring(dff2) in a string(dff). And if substring in a string exists add this row to the output dataframe. The problem is that I wrote the code using loops and it works too slow for thousands of rows. How could I rewrite the code using pandas methods?
data = {'long string column': ['aaaabbbbccccdddd', 'bbbbccccddddeeee','ccccddddeeeeffff','ddddeeeeffffgggg'],}
dff = pd.DataFrame(data, columns = ['long string column'])
data2 = {'substring_column': ['aaaa', 'bbbb','cccc','dddd'], 'status': ['best', 'good','bad','worst'],}
dff2 = pd.DataFrame(data2, columns = ['substring_column','status'])
df_output = pd.DataFrame()
df_output_small = pd.DataFrame()
for i in range(len(dff2)):
df_output_small = dff[(dff['long string column'].str.contains(dff2['substring_column'].iloc[i]))]
df_output_small['status'] = dff2['status'].iloc[i]
df_output = df_output.append(df_output_small, ignore_index=True)
df_output