I am a biologist. I want to copy an approach that I read in a paper: "To permit associations with death rates to be investigated independently from weight, residuals for death rates were calculated by subtracting the predicted from observed values".
I have a set of death rates (that range from about 0.1 to 0.5), a set of body weights (that range from about 2 to 80), and I want to calculate residuals for the death rates after accounting for body weight.
I wrote this code:
import scipy
from scipy import stats
import sys
# This reads in the weight and mortality data to two lists.
Weight = []
Mortality = []
for line in open(sys.argv[1]):
line = line.strip().split()
Weight.append(float(line[-2]))
Mortality.append(float(line[-1]))
# This calculates the regression equation.
slope, intercept, r_value, p_value, std_err = scipy.stats.linregress(Mortality,Weight)
# This calculates the predicted value for each observed value
obs_values = Mortality
pred_values = []
for i in obs_values:
pred_i = float(i) * float(slope) + float(intercept)
pred_values.append(pred_i)
# This prints the residual for each pair of observations
for obs_v,pred_v in zip(obs_values,pred_values):
Residual = str(obs_v - pred_v)
print Residual
My question is, when I run this code, some of my residuals seem quite large:
> Sample1 839.710240214 > Sample2 325.787250084 > Sample3 -41.3006000084 > Sample4 -70.6676280159 > Sample5 267.05319407 > Sample6 399.204820103 > Sample7 560.723474144 > Sample8 766.292670196 > Sample9 267.05319407 > Sample10 2.7499420027
I'm wondering, do these results seem "normal"/ should they be "standardised" in some way/ did I do something wrong to obtain residuals for mortality rates after accounting for weight?
I would appreciate simple "plain english" answers with possibly code snippets if it were possible, as I am not a statistics expert!
Many thanks