So basically I have a simple program where I'm reading a tab delimited file (it's big around 1m rows) and I'm getting an error like the example below. In the 3rd row because of a space after Jan Verheijen doesn't register the tab after and it doesn't put the date value into the next column. This happens to some random rows in the files.
Can it be fixed with the code or does it have to be fixed directly in the tsv files?
Thank you so much.
This is the code I'm using.
df = pd.read_csv(data, sep='\t',header=None, index_col=None,encoding= 'unicode_escape')
| Language | Translated By | Date | Version |
|---|---|---|---|
| Portuguese | Paulo Guzmán | 25/04/2019 | 2.41 |
| Czech | Čampulka Jiří | 11/06/2014 | 1.96 |
| Dutch | Jan Verheijen 17/07/2022 2.60 | ||
| French | skorpix38 | 28/02/2018 | 2.39 |