I want to read a plaintext file using pandas. I have entries without delimiters and with different widths like this:
59967Y98Doe John 6211100004545SO20140314- 00024278
N0546664SCHMIDT-PETER 7441100008300AW20140314- 00023643
G4894jmhTAKLONSKY-JUERGEN 4211100005000TB20140315 00023882
34875738PODESBERG-SCHUMPERTS6211100003671SO20140315 00024622
- 1-8 is a string.
- 9-28 is a string.
- 29-31 is numeric.
- 32-34 is numeric.
- 35-41 is numeric.
- 42-43 is a string.
- 44-51 is a date (yyyyMMdd).
- 52 is minus or a blank
- Rest is a currency amount without a decimal point (the last 2 digits is always after the decimal point). For example: - 00024278 = -242.78 €
I know there is pd.read_fwf
There is an argument width. I could do this:
pd.read_fwf(StringIO(txt), widths=[8], header="Peronal Nr.")
But how could I read my file with different columns widths?