My boss asked me to plot a matrix of DNA nucleotides using the pdf graphics function in R. I have a bit of code I'm working with but I can't figure it out and have spent way too much time trying! I know there are likely other methods/packages out there to visualize this genetic data, and I am absolutely interested in hearing them, but I also need to do this the way it was assigned to me.
I have sequence data in R like so:
> head(b)
Sequence X236 X237 X238 X239 X240 X241 X242 X244 X246 X247 X248 X249 X250 X251 X252 X253 X254 X255 X256 X257 X258 X259
1 L19088.1 G G G G G A G A C C A A G A T G G C C G A A
2 chr1_43580199_43586187 · · · · · · · · · · · · · · · · · · · · g g
There are a total of 1040 rows and 483 columns, with character possibilities A, a, G, g, T, t, C, c, mid-dot, or X.
I want to color the different characters and plot them in a way that is similar to a heatmap. The dots and the Xs don't need to be colored. The code I am working with so far is:
pdf(
sprintf(
"%s/L1.pdf",
out_dir),
width = 8.5, height = 11 )
par(omi = rep(0.5,4))
par(mai = rep(0.5,4))
par(bg = "#eeeeee")
plot( NULL,
xlim = c(1,100), ylim = c(1,140),
xlab = NA, ylab = NA,
xaxt = "n", yaxt = "n",
bty = "n", asp = 1 )
plot_width <- 100
w <- plot_width / 600
genome_colors <- list()
genome_colors[["A"]] <- "#ea0064"
genome_colors[["a"]] <- "#ea0064"
genome_colors[["C"]] <- "#008a3f"
genome_colors[["c"]] <- "#008a3f"
genome_colors[["G"]] <- "#116eff"
genome_colors[["g"]] <- "#116eff"
genome_colors[["T"]] <- "#cf00dc"
genome_colors[["t"]] <- "#cf00dc"
I <- nrow(b)
J <- ncol(b)
for ( i in 1:I ){
for ( j in i:J ){
# plot nucleotide as rectangle with color and text label, something like:
# plot nucleotides with genome_colors
# rect( (j-1)*w, top-(i-1)*w, j*w, top-i*w, col = color, border = NA )
}
# text( (j+1)*w, top-(i-1)*w, labels = i, cex = 0.05, col = "#dddddd" )
}
dev.off()
If anyone can help me with the plotting loop or point me in a helpful direction I will be so thankful!