How to modify/remove R rows that don't fit the regex pattern of the column?

Viewed 66

Here is an example of my current column, as well as my desired replacement.

Times <- c("12h00","16h30","Afternoon","15h00","14h20","7h30","06h00")
           
Output: ["12","16",NA,"15","14","7","6"]

I'm using a messy dataset right now, but I only want the column to contain the hours of each time. The vast majority are in the "##h##" format (07h30).

I assumed str_replace_all(Time, pattern, replacement) would work in this scenario, but am having doubts. I assume this "^\\d{2}h\\d{2}$" would be the appropriate code. What is the easiest way to nullify data that does not fit the column pattern?

My end goal is to create a histogram of 24 bins for each hour of the day, each time is an occurrence of a shark attack.

What do you think?

EDIT: there are a few with the #h## format as in "7h30", while I hope replace that with a plain "7" it isn't 100% necessary due to how few there are.

1 Answers

You can use

library(stringr)
Times <- c("12h00","16h30","Afternoon","15h00","14h20","7h30","06h00")
str_extract(Times, '[1-9]\\d*(?=h)')
## => [1] "12" "16" NA   "15" "14" "7"  "6" 

The pattern will extract

  • [1-9] - a non-zero digit
  • \d* - zero or more digits
  • (?=h) - that are immediately followed with h.

See the regex demo and the R demo.

Related