To select a list of column names starting with a specific string, one can use the starts_with() function in dplyr. To illustrate, we'll select the two columns that start with the string Sepal, as in Sepal.Length and Sepal.Width.
library(dplyr)
select(iris,starts_with("Sepal")) %>% head()
...and the output:
> select(iris,starts_with("Sepal")) %>% head()
Sepal.Length Sepal.Width
1 5.1 3.5
2 4.9 3.0
3 4.7 3.2
4 4.6 3.1
5 5.0 3.6
6 5.4 3.9
>
We can do the same thing in Base R with grepl() and a regular expression.
# base R version
head(iris[,grepl("^Sepal",names(iris))])
...and the output:
> head(iris[,grepl("^Sepal",names(iris))])
Sepal.Length Sepal.Width
1 5.1 3.5
2 4.9 3.0
3 4.7 3.2
4 4.6 3.1
5 5.0 3.6
6 5.4 3.9
>
Also note that if one is using read.csv() to create a data frame in R, it converts any occurrences of * in column headings to ..
# confirm that * is converted to . in read.csv()
textFile <- 'v*1,v*2
1,2
3,4
5,6'
data <- read.csv(text = textFile,header = TRUE)
# see how illegal column name * is converted to .
names(data)
...and the output:
> names(data)
[1] "v.1" "v.2"
>