Running the same script on the same pandas data produces very slightly different dataframes floating-point values

Viewed 118

I am executing a script that I have run before on the same data. The dataframe I get is only so slightly different than the previous one (in the 10th decimal point or so). For example:

  • in some column (and row) the old dataframe contains the price 5673391.88.
  • In the same column and same row of the new dataframe the value seems to be exactly the same (5673391.88).
  • However if I subtract the two columns I get a difference of -9.445123e-10.

This is of course the case for the entire column, not just the particular row. How can that be? Please note that I cannot confirm same environment (pandas or Python version) between the two script runs. Can it be one of these two reasons? Something else?

1 Answers

One possible reason: Pandas 1.2.0 which was released back in 26 Dec 2020, they have highlighted this issue:

Change in default floating precision for read_csv and read_table

the methods read_csv() and read_table() could read floating point numbers slightly incorrectly with respect to the last bit in precision.

Before this version floating_precision="high" has always been available to avoid this issue.

But, within this version the default is now floating_precision=None to make precision more acurate. It won't have any impact on performance.

Related