Linear Regression Residuals - Should I "standardise" the results and how to do this

Viewed 3805

I am a biologist. I want to copy an approach that I read in a paper: "To permit associations with death rates to be investigated independently from weight, residuals for death rates were calculated by subtracting the predicted from observed values".

I have a set of death rates (that range from about 0.1 to 0.5), a set of body weights (that range from about 2 to 80), and I want to calculate residuals for the death rates after accounting for body weight.

I wrote this code:

import scipy
from scipy import stats
import sys


# This reads in the weight and mortality data to two lists. 
Weight = []
Mortality = []
for line in open(sys.argv[1]):
        line = line.strip().split()
        Weight.append(float(line[-2]))
        Mortality.append(float(line[-1]))

# This calculates the regression equation.
slope, intercept, r_value, p_value, std_err = scipy.stats.linregress(Mortality,Weight)

# This calculates the predicted value for each observed value
obs_values = Mortality
pred_values = []
for i in obs_values:
    pred_i = float(i) * float(slope) + float(intercept)
    pred_values.append(pred_i)

# This prints the residual for each pair of observations
for obs_v,pred_v in zip(obs_values,pred_values):
    Residual = str(obs_v - pred_v)
    print Residual

My question is, when I run this code, some of my residuals seem quite large:

> Sample1 839.710240214 
> Sample2 325.787250084 
> Sample3 -41.3006000084
> Sample4 -70.6676280159
> Sample5 267.05319407
> Sample6 399.204820103
> Sample7 560.723474144
> Sample8 766.292670196
> Sample9 267.05319407
> Sample10 2.7499420027

I'm wondering, do these results seem "normal"/ should they be "standardised" in some way/ did I do something wrong to obtain residuals for mortality rates after accounting for weight?

I would appreciate simple "plain english" answers with possibly code snippets if it were possible, as I am not a statistics expert!

Many thanks

2 Answers
Related