I have a joint probability table for which I would like to calculate basic statistics, e.g.: E(x), standard deviation, covariance, correllation coefficient, etc.
I'm new to R and I've been able to figure out how to do all of these things manually, however I'd like learn a more efficient way of doing these calculations.
These are two soccer teams joint probability distribution of scoring 0-3 goals, hence the row and column names are numbers such that the margin sums multiplied by the column/row header, give the expected value of x or y.
#fill matrix by rows
bivdist <- matrix( c(.14, .11, .09, .1, .12, .1, .05, .02, .09, .07, .04, .01, .03, .02, .01, 0), nrow=4, ncol=4, byrow = TRUE)
#Give names to the rows and columns
dimnames(bivdist) <- list( c("0", "1", "2", "3" ), ("0", "1", "2", "3"))
This is how I implemented a portion of the tasks, for example:
#get the margin sums
msumx<-marginSums(bivdist, 1);msumx;
#calculate expected value
EVx<-msumx[1]*0+msumx[2]*1+msumx[3]*2+msumx[4]*3; EVx;
#calculate variance
varx<-(0-EVx)^2*msumx[1]+(1-EVx)^2*msumx[2]+(2-EVx)^2*msumx[3]+(3-EVx)^2*msumx[4];varx;
#calculate standard deviation
sigmax<-sqrt(varx);sigmax;
To get Covariance of XY, I found a way that's a bit
cleaner:
X <- 0:3; Y <- 0:3;
joint_pmf<-matrix(c(.14, .11, .09, .1, .12, .1, .05, .02, .09, .07, .04, .01,
.03, .02, .01, 0), nrow=4, ncol=4, byrow = TRUE)
mu_X <- rowSums(joint_pmf) %*% X;
mu_Y<- colSums(joint_pmf) %*% Y;
cov_XY <- X %*% joint_pmf %*% Y - mu_X * mu_Y;cov_XY;
corXY<-cov_XY/(sigmax*sigmay);corXY;
What's the more efficient, cleaner way of doing all of the above?