dask switch columns values when select subset of existing dataframe

Viewed 43

I have dataframe which I read with dask. If I read original dataframe it looks like this:

df=dd.read_csv('table/short_table.csv')

df.compute()

>>> Unnamed: 0     val1     val2   crop        val_int   val_3   val_4   y
0      0          -18.23    0.34   soybeans     14.0     0.32     0.11   0.55
1      1          -18.92    0.33   oranges      14.0     0.12     0.87   0.62
2      2          -18.11    0.22   cotton       14.0     0.44     0.12   0.43
...

I have tried to take only subset of that dataframe,but for some reason it takes wrong columns :

df=df[['crop','val_int','val1','val2','y']]
df.compute()

 
>>> crop     val_int  val1       val2    y
    -18.23    14.0    soybenas   0.55   0.34
    -18.92    14.0    oranges    0.62   0.33
    -18.11    14.0    cotton     0.43   0.22

I have reset the kernel and every time when I run it it takes wrong values , sometimes in different order (like it mismtach different columns). I could not find any reason for this phenomenon, why would it change the values? is there any way to get the correct columns with correct values?

0 Answers
Related