Confusing character encoding in R

Viewed 157

I have two dataframes which I am unable to share fully due to the data being confidential. I am supposed to merge them using the LABEL variable which exists in both datasets and contains some Unicode characters such as č, ž and so forth. However, the merging process created more rows than expected and on further inspection, I have found out that in the first dataframe, the values that contain Unicode characters are transcribed literally (e.g. you can see the label VŽ in the dataframe), whereas in the second dataframe, the label is shown via its Unicode code, so instead of VŽ you would see V\u008e. I used the stri_enc_mark function on both dataframes, here is the code and output for dataframe 1:

stri_enc_mark(unique(data1$Label)) %>% cbind(unique(data1$Label))

Output:

      .           
 [1,] "ASCII" "ZD"
 [2,] "ASCII" "RI"
 [3,] "ASCII" "PU"
 [4,] "ASCII" "ZG"
 [5,] "ASCII" "DU"
 [6,] NA      NA  
 [7,] "ASCII" "KR"
 [8,] "ASCII" "DA"
 [9,] "ASCII" "MA"
[10,] "ASCII" "ST"
[11,] "UTF-8" "VŽ"
[12,] "ASCII" "KA"
[13,] "ASCII" "SB"
[14,] "ASCII" "BM"
[15,] "ASCII" "VT"
[16,] "ASCII" "BJ"
[17,] "ASCII" "DJ"
[18,] "ASCII" "OS"
[19,] "ASCII" "SK"
[20,] "ASCII" "GS"
[21,] "UTF-8" "PŽ"
[22,] "UTF-8" "ŠI"
[23,] "UTF-8" "KŽ"
[24,] "ASCII" "Vk"
[25,] "UTF-8" "ŽU"
[26,] "ASCII" "KC"
[27,] "ASCII" "DE"
[28,] "ASCII" "NA"
[29,] "UTF-8" "ČK"
[30,] "ASCII" "KT"
[31,] "ASCII" "IM"
[32,] "ASCII" "VU"
[33,] "ASCII" "NG"
[34,] "ASCII" "VK"
[35,] "ASCII" "OG"
[36,] "ASCII" "SL"

And for dataframe 2:

stri_enc_mark(unique(data2$Label)) %>% cbind(unique(data2$Label))

Output:

      .                
 [1,] "ASCII" "BJ"     
 [2,] "ASCII" "BM"     
 [3,] "UTF-8" "\xc8K"  
 [4,] "ASCII" "DA"     
 [5,] "ASCII" "DE"     
 [6,] "ASCII" "DJ"     
 [7,] "ASCII" "DU"     
 [8,] "ASCII" "GS"     
 [9,] "ASCII" "IM"     
[10,] "ASCII" "KA"     
[11,] "ASCII" "KC"     
[12,] "ASCII" "KR"     
[13,] "ASCII" "KT"     
[14,] "UTF-8" "K\u008e"
[15,] "ASCII" "MA"     
[16,] "ASCII" "NA"     
[17,] "ASCII" "NG"     
[18,] "ASCII" "OG"     
[19,] "ASCII" "OS"     
[20,] "ASCII" "PU"     
[21,] "UTF-8" "P\u008e"
[22,] "ASCII" "RI"     
[23,] "ASCII" "SB"     
[24,] "ASCII" "SK"     
[25,] "ASCII" "ST"     
[26,] "UTF-8" "\u008aI"
[27,] "ASCII" "VK"     
[28,] "ASCII" "VU"     
[29,] "UTF-8" "V\u008e"
[30,] "ASCII" "ZD"     
[31,] "ASCII" "ZG"     
[32,] "UTF-8" "\u008eU"
[33,] "ASCII" "VT"  

So as far as I can see, both the "literal" labels and the labels with Unicode code are encoded as UTF-8, which I find surprising because if that is the case, I cannot comprehend why is one dataframe showing VŽ and the other one V\u008e.

I want to convert the coded labels to literal labels, I have tried to the following:

data2 %>%
  mutate(Label = recode(Label, "\xc8K" = "ČK",
                             "K\u008e" = "KŽ",
                             "P\u008e" = "PŽ",
                             "\u008aI" = "ŠI",
                             "V\u008e" = "VŽ",
                             "\u008eU" = "ŽU"))

But this doesn't succeed and I get the following warnings:

Warning messages:
1: unable to translate 'K<U+008E>' to native encoding 
2: unable to translate 'P<U+008E>' to native encoding 
3: unable to translate '<U+008A>I' to native encoding 
4: unable to translate 'V<U+008E>' to native encoding 
5: unable to translate '<U+008E>U' to native encoding 

So, how can I recode the values properly?

0 Answers
Related