I'm looking for a fast way to read in a single column from tab-separated text that lives as a character vector in memory.
I'm using a file format specific to my field that roughly resembles a compressed tsv file. It is fast and easy to read in a subset of lines from such file, but unfeasible to read in data directly with read.table(), data.table::fread() or readr::read_tsv() due to memory limitations (and the rows I need aren't known a priori).
So, I end up with a character vector in memory, with an element for every line, but with the tab separators still in there. I'm a bit puzzled on how to quickly extract a specific column from this text. In the example below, what is the fastest way to extract the third column? There aren't any 'surprises' in the text such as comments or quoted names, but in my real case the columns don't have a fixed width. The fastest method I've found so far is to use the readr::read_tsv() function.
library(readr)
set.seed(0)
# About 88Mb of memory
n_examples <- 1e6
text <- paste(
as.character(as.hexmode(sample(n_examples))),
as.character(as.hexmode(sample(n_examples))),
as.character(as.hexmode(sample(n_examples))),
as.character(as.hexmode(sample(n_examples))),
sep = "\t"
)
fun_read.table <- function(x, i) {
read.table(
text = x, sep = "\t",
colClasses = c("character", "character", "character", "character")
)[[i]]
}
fun_read_tsv <- function(x, i) {
read_tsv(file = I(x), col_select = all_of(i),
col_types = "cccc", col_names = LETTERS[1:4])[[1]]
}
bm <- bench::mark(
fun_read.table(text, 3),
fun_read_tsv(text, 3),
min_iterations = 5
)
#> Warning: Some expressions had a GC in every iteration; so filtering is disabled.
print(bm)
#> # A tibble: 2 x 13
#> expression min median `itr/sec` mem_alloc `gc/sec` n_itr
#> <bch:expr> <bch:tm> <bch:tm> <dbl> <bch:byt> <dbl> <int>
#> 1 fun_read.table(text, 3) 1.34s 1.57s 0.619 93.2MB 1.11 5
#> 2 fun_read_tsv(text, 3) 879.36ms 903.72ms 1.10 35.1MB 0.219 5
#> # ... with 6 more variables: n_gc <dbl>, total_time <bch:tm>, result <list>,
#> # memory <list>, time <list>, gc <list>
Below are some alternatives that I've tried, but weren't faster than read_tsv(). The data.table::fread() method was surprisingly slow due to it writing the input text to a temporary file first. I haven't managed to figure out a regex-based method to capture the third column, so I don't know if that would be faster.
library(data.table)
#> Warning: package 'data.table' was built under R version 4.1.1
fun_tstrsplit <- function(x, i) {
tstrsplit(x, "\t", keep = i)[[1]]
}
fun_fread <- function(x, i) {
fread(
text = x, sep = "\t",
colClasses = c("character", "character", "character", "character"),
select = i
)[[1]]
}
fun_scan <- function(x, i) {
ncols <- lengths(regmatches(x[[1]], gregexpr("\t", x[[1]]))) + 1
scan(
text = x, sep = "\t", what = "", quiet = TRUE
)[seq_along(x) %% ncols == i]
}
Created on 2021-10-13 by the reprex package (v2.0.1)