how to get the position of nth occurrence of a string in a file

Viewed 84

I have an xml file which contains data in a single line where same string is repeated multiple times in it.

I am looking to identify the position of nth occurrence of a string in that file so that i can split single file into multiple files based on that position so that it will be easy for processing.

sample data in file:

<id = 1><\id><id = 2><\id><id = 3><\id><id = 4><\id><id = 5><\id><id = 6><\id><id = 7><\id><id = 8><\id><id = 9><\id><id = 10><\id><id = 11><\id>

So i want to split the file based on the id tag. for eg i want to look for position of 5th occurrence of the id tag and need to split the file into 3 files totally

Output:

file_1:

<id = 1><\id><id = 2><\id><id = 3><\id><id = 4><\id>

file_2:

<id = 5><\id><id = 6><\id><id = 7><\id><id = 8><\id>

file_3:

<id = 9><\id><id = 10><\id><id = 11><\id>

I tried splitting the one line into multiple lines with a simple sed sed 's/></>\n</g' $file > data.txt Later with a simple grep i identified the line number and started splitting based on the line number. This is working for smaller files but some file are in GB's (10-20) which is causing issues.

Could you help me if there is any easy way to get the position of the nth occurrence of a string in file so that i can split single file into multiple files based on the string position.

2 Answers

This might work for you (GNU sed):

sed 's/<\\id>/&\n/4;P;D' file | sed -ne '1~3w file1' -e '2~3w file2' -e '3~3w file3'

In the first sed invocation, split each line into three following the fourth <\id> (BTW should this be </id>?).

Pipe the result to a second sed invocation.

In the second sed invocation, send the first line modulo three to file1, the second line modulo three to file2 and the third line modulo three to file3.

Alternative using split instead of the second invocation of sed:

sed 's/<\\id>/&\n/4;P;D' file | split -a 1 --nume=1 -dn r/3 - file

suggesting an gawk script (standard awk in most Linux machines) that can do all the splittings as well:

Parameters:

f : filename prefix

m : number of id elements in file

gawk script:

gawk 'BEGIN{RS="<\\\\id>"}{a=a$0 RT}NR%m==0{print a > f NR;a=""}END{if (a)print a> f (NR-(NR%m-1)) }' f="FileName_" m=5 input.txt

Sample run on given example: f=file_ m=4

gawk 'BEGIN{RS="<\\\\id>"}{a=a$0 RT}NR%m==0{print a > f NR;a=""}END{if (a)print a> f (NR-(NR%m-1)) }' f="file_" m=4 input.xml

file_1

<id = 1><\id><id = 2><\id><id = 3><\id><id = 4><\id>

file_5

<id = 5><\id><id = 6><\id><id = 7><\id><id = 8><\id>

file_9

<id = 9><\id><id = 10><\id><id = 11><\id>

gawk script explanation

BEGIN{RS="<\\\\id>"}    # set awk record seperator to <\id>
{output = output $0 RT} # accumulate each record in output variable
(NR % chunkSize) == 0 { # when read chunkSize of records
  # save output variable into file. named: filePrefix appended with record cout
  print output > filePrefix NR; 
  output = "";          # reset output variable
}
END { # after processing the last record
  if (output) { # there is existing output
    lastChunk = NR-((NR % chunkSize) - 1); # compute the file begin ID
    print output > filePrefix lastChunk; # write output to last file
  }
}
Related