Finding peaks within a data frame in R

Viewed 46

I am trying to analyse some biological data within R. I have a data frame that contains a window for positions on a dna sequence that I want to analyse. For example, 237-1437. I have a count of files which contains the position and a count. For each window, I want to analyse each position in the count file and look for significant peaks in counts. Does anyone know how to do this?

The count file looks like this and is within a data frame labelled df2:

V1    V2    V3   
gene  1     6
gene  2     0
gene  3     0
gene  4     10
....

The data frame which contains the window looks like this and is labelled df:

seqnames    start    end    strand   window_end   
gene        65       1237   +        1437
gene        1262     2134   +        2334
gene        2178     4511   +        4711

I want the output to produce a list of significant peaks.

1 Answers

When you say "finding peaks", statistically that means finding outliers in the data, or finding minimum and maximum numbers to help you investigate further and analyse these peak values.

Using Statistical Summary:

If you're interested in a specific column, let's say from your data frame df column V3, then in base R you can do the following:

summary(df$V3)

This would result in six statistical values in your data: minimum value, first quantile, median, mean, third quantile, and maximum value. Also, you can store the values in a vector and use the values for further analyses by using the index of each value in the summary.

Visualisation of the above along with outliers: In addition to printing these values, you can plot them in R using boxplot function; this would show you the outliers or the peaks with circles.

boxplot(df$V3)

Demo:

#generating df with additional random data to be able to plot and show outliers:
df = data.frame(V1 = rep("gene", 10), 
                V2 = 1:10, 
                V3 = c(6,0,0,10,50, 20, 5, 7, 9, 100))

df

Result:

     V1 V2  v3
1  gene  1   6
2  gene  2   0
3  gene  3   0
4  gene  4  10
5  gene  5  50
6  gene  6  20
7  gene  7   5
8  gene  8   7
9  gene  9   9
10 gene 10 100

The statistical summary:

summary(df$V3)

Result:

   Min. 1st Qu.  Median    Mean 3rd Qu.    Max. 
   0.00    5.25    8.00   20.70   17.50  100.00 

The boxplot:

boxplot(df$v3, ylab = "V3", main = "Boxplot")

Resultant plot:

enter image description here

Related