I have an xml file which contains data in a single line where same string is repeated multiple times in it.
I am looking to identify the position of nth occurrence of a string in that file so that i can split single file into multiple files based on that position so that it will be easy for processing.
sample data in file:
<id = 1><\id><id = 2><\id><id = 3><\id><id = 4><\id><id = 5><\id><id = 6><\id><id = 7><\id><id = 8><\id><id = 9><\id><id = 10><\id><id = 11><\id>
So i want to split the file based on the id tag. for eg i want to look for position of 5th occurrence of the id tag and need to split the file into 3 files totally
Output:
file_1:
<id = 1><\id><id = 2><\id><id = 3><\id><id = 4><\id>
file_2:
<id = 5><\id><id = 6><\id><id = 7><\id><id = 8><\id>
file_3:
<id = 9><\id><id = 10><\id><id = 11><\id>
I tried splitting the one line into multiple lines with a simple sed
sed 's/></>\n</g' $file > data.txt
Later with a simple grep i identified the line number and started splitting based on the line number. This is working for smaller files but some file are in GB's (10-20) which is causing issues.
Could you help me if there is any easy way to get the position of the nth occurrence of a string in file so that i can split single file into multiple files based on the string position.