I have multiple .txt from the USPS which need to be parsed.
The .txt file is like this:
C COPYRIGHT(C) 09-22 USPS 001 D91801Z221103711AFC049 HOWARD ST 00000005280000000528EELMHURST GUEST HOME B33983398B 050108CA03727 Z20050D91801Z221103713AFC027E MAIN ST 00000006300000000630EBROAD SOLUTIONS B40644064B 050108CA03727 Z20050D91801Z221103716AFC006E MAIN ST 00000010010000001001ODR FORTNER/FORTIER PROF. B41994199B 050108CA03727 Z20050D91801Z221103721AFC050N 4TH ST 00000001280000000128EMY LADIES GUEST HOME B34903490B 050108CA03727 Z20050D91801Z221103723AFC044N GARFIELD AVE 00000006240000000624BSCHILLINGS B14681468B 050108CA03727 Z20050D91801Z221103724AFC044N GARFIELD AVE 00000007120000000712BNAVARRO CONST CO B14971497B 050108CA03727 Z20050D91801Z221103725AFC044N GARFIELD AVE 00000004120000000412EHOLY TRINITY CHURCH B24982498B 050108CA03727 Z20050D91801Z221103726AFC044N GARFIELD AVE 00000004200000000420ERUSSELL MAIORANA DDS B24972497B 050108CA03727 Z20050D91801Z221103729AFC049N OLIVE AVE 00000000210000000021OBETHANY CHURCH B33863386B 050108CA03727 Z20050D91801Z221103731AFC061S 1ST ST 00000001110000000111OALHAMBRA CITY HALL B37963796B 050108CA03727 Z20050D91801Z221103734AFC036S ALMANSOR ST 00000008400000000840EFIRST LUTHERAN CHURCH B45994599B 050108CA03727 Z20050D91801Z221103735AFC023S ATLANTIC BLVD 00000002140000000214EATHERTON BAPTIST HOMES B32983298B 050108CA03727
I have tried reading the .txt as:
df = read.table("918.txt", sep = "", header = F)
This method kind of works but all the records are read in one row (only one observation with many columns). How can I have a new row whenever a condition is met. The information after Z should be a new row. Each row or record contains 182 characters (spaces included).
In a sense it should read like so:
Z20050D91801Z221103725AFC044N GARFIELD AVE 00000004120000000412EHOLY TRINITY CHURCH B24982498B 050108CA03727
Z20050D91801Z221103726AFC044N GARFIELD AVE 00000004200000000420ERUSSELL MAIORANA DDS B24972497B 050108CA03727
Z20050D91801Z221103729AFC049N OLIVE AVE 00000000210000000021OBETHANY CHURCH B33863386B 050108CA03727
Also, is there a way to parse certain fields even more? For example, the first portion of each address starts with Z followed by a bunch of characters. I'd like to separate those further.
Thank you all for your help!