remove new line \n in large file (10GB)

Viewed 156

I have large file 1.txt containing:

User: Test1
Password: P@sawFia1_f

User: Test2
Password: C99vijJiDB9fo@K!!1

I'm using sed -i '/\nPassword/ s///g' 1.txt for remove new line with Password: but it's not removing it. Why? The final output needs to be:

User: Test1;P@sawFia1_f

User: Test2;C99vijJiDB9fo@K!!1
6 Answers

Assuming the lines are paired like that, you can use the following:

perl -pe'
   s/^User:.*\K\n/;/;
   s/^Password:\s*//;
' file.in >file.out

(It can be used as-is or placed all on one line.)

Assumptions:

  • every User: line is followed by a Password: line
  • the actual password value does not contain white space
  • each User/password combo is followed by a blank line
  • all other lines in the file are ignored/discarded (otherwise OP should update the sample input to show how other lines of data are to be processed)

One awk approach:

$ awk '/^User:/ {printf "%s",$0} /^Password:/ {printf ";%s\n\n",$2}' 1.txt
User: Test1;P@sawFia1_f

User: Test2;C99vijJiDB9fo@K!!1

Once OP confirms the script works as needed, and assuming OP wants to overwrite the original file, and assuming OP is running GNU awk, OP can add the -i inplace flag to have 1.txt overwritten, eg:

awk -i inplace '/^User:/ { printf "%s", $0 } /^Password:/ { printf ";%s\n\n",$2}' 1.txt

Using any awk, given your provided sample input/output all you'd need is:

$ awk -v RS= '{print $1, $2 ";" $4}' file1.txt
User: Test1;P@sawFia1_f
User: Test2;C99vijJiDB9fo@K!!1

or if you really do need a blank line between each output line:

$ awk -v RS= -v ORS='\n\n' '{print $1, $2 ";" $4}' file1.txt
User: Test1;P@sawFia1_f

User: Test2;C99vijJiDB9fo@K!!1

If that's not all you need then please edit your question to include more truly representative sample input/output including cases that the above doesn't work for.

Assuming the shown structure, of User and Password lines followed by an empty line

perl -i.bak -00 -wpe's/\nPassword:\s*/;/' file

Reads the file in paragraphs (by -00 switch), so applying the regex to each pair of lines in a string.

The -i.bak changes the input file "in-place" but also keeps a backup (file.bak). If you don't want a backup just remove .bak part, once it's all well tested.


Or, process line by line

perl -i.bak -wnlE'/^Password:\s*(.*)/ ? say "$u;$1" : /^User/ ? $u=$_ : say' file

This works with, and reprints, any other lines as well.

If there is only an empty line in between, and which needn't be retained, it simplifies to

perl -i.bak -wnlE'/^Password:\s*(.*)/ ? say "$u;$1" : ($u=$_)' file

With your shown samples, please try following awk code, written and tested in GNU awk.

awk -v RS='(^|\n)User:[^\n]*\nPassword:[^\n]*' '
RT{
  sub(/^\n/,"",RT)
  sub(/\n/,";",RT)
  print RT
}
' Input_file

Explanation: Using GNU awk, setting RS(record separator) to (^|\n)User:[^\n]*\nPassword:[^\n]*(explained further in post). In main section of awk checking if RT is NOT NULL then substituting starting new line with NULL in it and then substituting new line with ;, finally printing its value as per required output.

NOTE: Above will print the output on terminal, once you are happy with results you can use GNU awk's -i inplace option, change awk to awk -i inplace in above code.

One liner form of above code:

awk -v RS='(^|\n)User:[^\n]*\nPassword:[^\n]*' 'RT{sub(/^\n/,"",RT);sub(/\n/,";",RT);print RT}' Input_file

I'm using sed -i '/\nPassword/ s///g' 1.txt for remove new line with Password: but it's not removing it. Why?

You are misunderstanding how GNU sed works. In basic usage it does apply changes to each line, latter understand as characters between start of file or newline and end of file or newline, therefore such line does not contain newline. Your task require knowning 2 lines of input before procuring 1 line of output. This can be done exploiting GNU sed feature dubbed hold space following way, let file.txt content be

User: Test1
Password: P@sawFia1_f

User: Test2
Password: C99vijJiDB9fo@K!!1

then

sed -e '/^User/{h;d}' -e '/^Password/{H;g;s/\nPassword: /;/}' file.txt

gives output

User: Test1;P@sawFia1_f

User: Test2;C99vijJiDB9fo@K!!1

Explanation:

  • for line starting with User, save current line into hold (h) and go to next line (d)
  • for line starting with Password append newline and current line to hold (H) then set current line content to that of hold (g) then replace newline followed by Password followed by : followed by space using semicolon.

Disclaimer: this solution assumes that every line starting with User is always followed by line starting with Password and every line starting with Password is preceded by line starting with User.

(tested in GNU sed 4.2.2)

Related