With PySpark I want to store following date under the TimestampType in Spark: 0001-01-01T00:00:00Z. It fails when trying to access the data (i.e. with DataFrame.head()) with the following error:
Year 0 is out of range
Which is completely baffling as I am not using year 0.
I did a bit digging into the code and found PySpark is doing the following:
return datetime.datetime.fromtimestamp(ts // 1000000).replace(microsecond=ts % 1000000)
I did some tests and it seems the issues with timestamps starts for 0001-01-01T23:59:59Z so it seems it cannot deal with "the first day of year 1".
I have tested it with the below:
d = datetime(1,1,1,23,59,59)
print(d.timestamp())
What's the expected solution for the problem? It feels like a bug in PySpark? (Or even in Python as I found this: https://bugs.python.org/issue31212, but I have issues not only with "datetime.min" (I could imagine root cause is the same?)