Why does dateutil.parser('May 10, 2019') return inconsistent year values when there is/isn't a space after the comma?

Viewed 89
>>> from dateutil import parser
>>> parser.parse('May 10,2019')
datetime.datetime(2020, 5, 10, 0, 0)
>>> parser.parse('May 10, 2019')
datetime.datetime(2019, 5, 10, 0, 0)

Notice the space or no space after the comma.

It seems to be parsing a two digit year when there is no space after the comma, and a four digit year when there is a space after the comma.

Is this expected?

The versions I have:

$ pip show python-dateutil Name: python-dateutil Version: 2.8.0

$ python3 Python 3.6.9 (default, Apr 18 2020, 01:56:04)

1 Answers

This probably won't be much help, but should at least provide some extra information.

It's not parsing a two-digit year in one case and a four-digit one in the other, it's actually defaulting to the current year in the case without a space, where it can't parse the year for some reason.

>>> from dateutil import parser
>>> parser.parse("August 06, 1881")
datetime.datetime(1881, 8, 6, 0, 0)
>>> parser.parse("August 06,1881")
datetime.datetime(2020, 8, 6, 0, 0)

This issue has since been opened on Github https://github.com/dateutil/dateutil/issues/939 and seems to be related to the fact that the comma can be used as a separator within times (say something like 23,5 seconds). It also apparently used to work: https://github.com/dateutil/dateutil/issues/1075 So there's hope for a fix, but it will involve diving into the code.

A band-aid fix of applying .replace(",", ", ") to the string might work in the meantime, but it's certainly not the easiest thing to read.

This could potentially be useful too, but the Github issue is probably the best: https://dateutil.readthedocs.io/en/stable/parser.html#dateutil.parser.parserinfo.JUMP

Related