Regressing out or Removing age as confounding factor from experimental result

Viewed 1262

I have obtained cycle threshold values (CT values) for some genes for diseased and healthy samples. The healthy samples were younger than the diseased. I want to check if the age (exact age values) are impacting the CT values. And if so, I want to obtain an adjusted CT value matrix in which the gene values are not affected by age.

I have checked various sources for confounding variable adjustment, but they all deal with categorical confounding factors (like batch effect). I can't get how to do it for age.

I have done the following:

modcombat = model.matrix(~1, data=data.frame(data_val))
modcancer = model.matrix(~Age, data=data.frame(data_val))
combat_edata = ComBat(dat=t(data_val), batch=Age, mod=modcombat, par.prior=TRUE, prior.plots=FALSE)

pValuesComBat = f.pvalue(combat_edata,mod,mod0)
qValuesComBat = p.adjust(pValuesComBat,method="BH")

data_val is the gene expression/CT values matrix. Age is the age vector for all the samples. For some genes the p-value is significant. So how to correctly modify those gene values so as to remove the age effect?

I tried linear regression as well (upon checking some blogs):

lm1 = lm(data_val[1,] ~ Age) #1 indicates first gene. Did this for all genes
cor.test(lm1$residuals, Age)

The blog suggested checking p-val of correlation of residuals and confounding factors. I don't get why to test correlation of residuals with age. And how to apply a correction to CT values using regression?

Please guide if what I have done is correct. In case it's incorrect, kindly tell me how to obtain data_val with no age effect.

1 Answers

There are many methods to solve this:-

Basic statistical approach

A very basic method to incorporate the effect of Age parameter in the data and make the final dataset age agnostic is:

Do centring and scaling of your data based on Age. By this I mean group your data by age and then take out the mean of each group and then standardise your data based on these groups using this mean.

For standardising you can use two methods:

1) z-score normalisation : In this you can change each data point to as (x-mean(x))/standard-dev(x)); by using group-mean and group-standard deviation.

2) mean normalization: In this you simply subtract groupmean from every observation.

3) min-max normalisation: This is a modification to z-score normalisation, in this in place of standard deviation you can use min or max of the group, ie (x-mean(x))/min(x)) or (x-mean(x))/max(x)).

On to more complex statistics:

You can get the importance of all the features/columns in your dataset using some algorithms like PCA(principle component analysis) (https://en.wikipedia.org/wiki/Principal_component_analysis), though it is generally used as a dimensionality reduction algorithm, still it can be used to get the variance in the whole data set and also get the importance of features.

Below is a simple example explaining it:

I have plotted the importance using the biplot and graph, using the decathlon dataset from factoextra package:

library("factoextra")
data(decathlon2)

colnames(data)

data<-decathlon2[,1:10] # taking only 10 variables/columns for easyness

res.pca <- prcomp(data, scale = TRUE)
#fviz_eig(res.pca)
fviz_pca_var(res.pca,
             col.var = "contrib", # Color by contributions to the PC
             gradient.cols = c("#00AFBB", "#E7B800", "#FC4E07"),
             repel = TRUE     # Avoid text overlapping
)


hep.PC.cor = prcomp(data, scale=TRUE)

biplot(hep.PC.cor)

output

[1] "X100m"        "Long.jump"    "Shot.put"     "High.jump"    "X400m"        "X110m.hurdle"
 [7] "Discus"       "Pole.vault"   "Javeline"     "X1500m"

enter image description here

enter image description here

On these similar lines you can use PCA on your data to get the importance of the age parameter in your data.

I hope this helps, if I find more such methods I will share.

Related