How can I remove non-breaking spaces from a text file in bash?

Viewed 689

I have a csv file with text and numbers.

If a number is bigger than 1000, formatted like this: 1 000, so it has a space as thousand separator, but it is not space. I tried to sed it, and it worked where real space was, but not in this format.

It is also not TAB, I removed all the TABs with "expand -t 1".

The following is a line that demonstrates the issue:

x17_Provident_GDN_REMARKETING_provident.hu_listák;Display_Hálózat;Szeged;2021-03-09;Kedd;Mobil;HUF;1 736;9;130.83;0.00

In penultimate row, in column 8: 1 736 is the problem.

And running this: grep -E -m 1 -e '[;]1[^;]+736[;]' <yourfile.csv | hexdump -C

gives:

00000000  78 31 37 5f 50 72 6f 76  69 64 65 6e 74 5f 47 44  |x17_Provident_GD|
00000010  4e 5f 52 45 4d 41 52 4b  45 54 49 4e 47 5f 70 72  |N_REMARKETING_pr|
00000020  6f 76 69 64 65 6e 74 2e  68 75 5f 6c 69 73 74 c3  |ovident.hu_list.|
00000030  a1 6b 3b 44 69 73 70 6c  61 79 5f 48 c3 a1 6c c3  |.k;Display_H..l.|
00000040  b3 7a 61 74 3b 53 7a 65  67 65 64 3b 32 30 32 31  |.zat;Szeged;2021|
00000050  2d 30 33 2d 30 39 3b 4b  65 64 64 3b 4d 6f 62 69  |-03-09;Kedd;Mobi|
00000060  6c 3b 48 55 46 3b 31 c2  a0 37 33 36 3b 39 3b 31  |l;HUF;1..736;9;1|
00000070  33 30 2e 38 33 3b 30 2e  30 30 0a                 |30.83;0.00.|
0000007b
3 Answers

It's a 2 byte, UTF-8 encoded non breaking space - c2 a0.

You can use perl to safely remove it.

perl -pe 's/\xc2\xa0//g' dirty.csv > clean.csv

After we know it is No break space, I simply sed it on mac with entry method:

opt+space
cat test4.csv | sed 's/ //g'

Similar to perl, you can use GNU sed with LC_ALL=C:

LC_ALL=C sed 's/\xc2\xa0//g'
Related