Ordering entries in dataframe based on a defined function in R

Viewed 37

I have a dataframe contains probability values like this

#   Spot 1 Spot 2
# 1  0.140   0.12
# 2  0.220   0.50
# 3  0.154   0.40
# 4  0.300   0.12
# 5  0.220   0.60
# 6  0.400   0.23
# 7  0.550   0.40
# 8  0.600   0.56

The rownames are actually cell names, for example, 1 means cell type 1, etc. I make a function to categorise whether the cell type is present in a spot or not. The cell type is present if the probability value >= threshold t. So, I make a function to identify the cell type present in a spot or not like this

solution_matrix <- function(probability_matrix, threshold, n=0) {
  #parameter n is the number of cell type ranking we want. For example, n=3 means we choose the highest ranking of 3 cell types to be present in a spot
  probability_matrix <- probability_matrix
  for (i in 1:nrow(probability_matrix)) {
    for (j in 1:ncol(probability_matrix)) {
      if (probability_matrix[i, j] >= threshold) {
        probability_matrix[i, j] <- probability_matrix[i, j]
      } else {
        probability_matrix[i, j] <- 0
      }
    }
  }
  
  probability_matrix <- probability_matrix >= threshold  
  
  solution_matrix <- probability_matrix
  for (i in 1:nrow(solution_matrix)) {
    for (j in 1:ncol(solution_matrix)) {
      if (solution_matrix[i, j] == TRUE) {
        solution_matrix[i, j] <- i - 1 #We minus 1 since the cluster numbers start from 0 to 9 (there are 10 clusters)
      } else {
        solution_matrix[i, j] <- NA
      }
    }
  }
  for (j in 1:ncol(solution_matrix)) {
    solution_matrix[, j] <- sort(as.numeric(solution_matrix[, j]), na.last=TRUE) #This is to make the NA entries to be filled after the number entries
  }
  
  solution_matrix <- solution_matrix
  colnames(solution_matrix) <- gsub("\\.", "-", colnames(solution_matrix)) #We need to rename the spot IDs in estimated_solution from using . to - to match it to true_solution obtained from the synthetic_data 
  
  return(list(probability_matrix, solution_matrix))
}#solution_matrix#
solution_matrix(mydata, 0.2, 3)

It will result output like this

[[1]]
Spot 1 Spot 2
1  FALSE  FALSE
2   TRUE   TRUE
3  FALSE   TRUE
4   TRUE  FALSE
5   TRUE   TRUE
6   TRUE   TRUE
7   TRUE   TRUE
8   TRUE   TRUE

[[2]]
Spot 1 Spot 2
1      1      1
2      3      2
3      4      4
4      5      5
5      6      6
6      7      7
7     NA     NA
8     NA     NA

I want to order the cell types present in the spot based on their probability values (from the high values should be rank 1, rank 2, and so on). However, I do not know how to rank those cell types while keeping a record of what cell type is present in what spot. In this output, I have no idea which cell types have a stronger presence in a spot. The variable n should be included in the function to show how many best cell types presences we want. For example, in this output, we have 6 cell types in spot 1. If I set n = 3 then what I want is to have only the best 3 cell types, instead of 6. Any idea, please?

data

mydata <- structure(list(`Spot 1` = c(0.14, 0.22, 0.154, 0.3, 0.22, 0.4, 
0.55, 0.6), `Spot 2` = c(0.12, 0.5, 0.4, 0.12, 0.6, 0.23, 0.4, 
0.56)), class = "data.frame", row.names = c(NA, 8L))
0 Answers
Related