I am trying to understand how to extract the top ten rows from a table based on the values in one numeric column, but only from rows that meet a condition applied to a second numeric column.
First about the data. I have a table listing, for a few thousand human genes, the difference in expression to a control (log_fold_change) and a p-value for this difference (p_value). The table looks something like this:
log_fold_change p_value
APOD 1.7388209 0.4820801
S100B -1.1514299 0.5995658
CD63 0.6066951 0.4935413
PMEL -1.4977796 0.1862176
MT2A -0.9311173 0.8273733
S100A6 -0.4555436 0.6684667
TIMP1 -1.9464387 0.7942399
VIM -0.4704482 0.1079436
PAEP 1.4787634 0.7237109
CSTB -0.6386040 0.4112744
The data can be re-created using these commands (creates a table with data for n fictional genes):
n <- 50
log_fold_change <- runif(n, -2.0, 2.0)
p_value <- runif(n, 0, 1.0)
df <- data.frame(log_fold_change, p_value)
rownames(df) <- stringi::stri_paste(stringi::stri_rand_strings(n, 3, '[A-Z]'),stringi::stri_rand_strings(n, 1, '[1-9]'))
I have created a column for the labels (df$label <- NA), in which I plan to transfer the gene names that I want to label when I plot my graph. Which genes, you ask? I wish to, amongst genes with a positive log_fold_change, extract the ten genes with the smallest p_value.
I have already found a way to extract and label the 10 genes with the smallest p_value:
df$label[with(df, rank(p_value)) %in% c(1:10)] <- rownames(df)[with(df, rank(p_value)) %in% c(1:10)]
Now, how do I enforce the condition df$log_fold_change > 0 so that my ten genes with the smallest p_value are only selected from the genes with a positive log_fold_change? Any help would be really appreciated!