I have dataframe which I read with dask. If I read original dataframe it looks like this:
df=dd.read_csv('table/short_table.csv')
df.compute()
>>> Unnamed: 0 val1 val2 crop val_int val_3 val_4 y
0 0 -18.23 0.34 soybeans 14.0 0.32 0.11 0.55
1 1 -18.92 0.33 oranges 14.0 0.12 0.87 0.62
2 2 -18.11 0.22 cotton 14.0 0.44 0.12 0.43
...
I have tried to take only subset of that dataframe,but for some reason it takes wrong columns :
df=df[['crop','val_int','val1','val2','y']]
df.compute()
>>> crop val_int val1 val2 y
-18.23 14.0 soybenas 0.55 0.34
-18.92 14.0 oranges 0.62 0.33
-18.11 14.0 cotton 0.43 0.22
I have reset the kernel and every time when I run it it takes wrong values , sometimes in different order (like it mismtach different columns). I could not find any reason for this phenomenon, why would it change the values? is there any way to get the correct columns with correct values?