\xe8 matching \xf1 in str_detect() and str_replace_all()

Viewed 171

I want to process text files including some characters shown in hexadecimals on R. When I tried to convert those back into more readable characters, I encountered some unexpected (to me) behaviours of stringr functions. Specifically, \xe8 apparently matches \xf1:

> library("tidyverse")
> str <- "ni\xf1a"
> str_detect(str, "\xe8")
[1] TRUE

This is inconvenient when I want to convert \xe8 into è and \xf1 into ñ in the same files:

> str %>% 
+   str_replace_all("\xe8", "è") %>% 
+   str_replace_all("\xf1", "ñ")
[1] "nièa" # I expect niña

Interestingly, gsub() works as I expect:

> str %>% 
+   gsub("\xe8", "è", .) %>% 
+   gsub("\xf1", "ñ", .)
[1] "niña"
  1. Why does \xe8 match \xf1 in str_detect() and str_replace_all()? Is there a way to avoid it?
  2. Why is the behaviour different between stringr functions and gsub()?

Update

Here is part of the output of devtools::session_info():

> devtools::session_info()
─ Session info ──────────────────────────────────────────────────────────────────
 setting  value                       
 version  R version 4.0.2 (2020-06-22)
 os       macOS Catalina 10.15.7      
 system   x86_64, darwin17.0          
 ui       RStudio                     
 language (EN)                        
 collate  en_GB.UTF-8                 
 ctype    en_GB.UTF-8                 
 tz       Europe/London               
 date     2020-09-30                  

─ Packages ──────────────────────────────────────────────────────────────────────
 package     * version date       lib source     
...
stringr     * 1.4.0   2019-02-10 [1] CRAN (R 4.0.2)
...
0 Answers
Related