I'm doing some practice exercises with respect to regression analysis in R. One of the questions asks me to perform some simple analysis using a linear regression function. To document the intermediate steps, I added the values to a preloaded data set which is what the following screen grab is:
The columns GPA and ACT_Score came preloaded with the data set. The following is the code I used to add the fitted_GPA column:
> GPA_lm = lm(formula = GPA ~ ACT_Score, data = ch_1_exer_119_GPA)
> fitted_GPA = coef(GPA_lm)[[1]] + coef(GPA_lm)[[2]]*ch_1_exer_119_GPA[,2] # created vector of fitted values
> ch_1_exer_119_GPA$fitted_GPA = fitted_GPA # added column of fitted values to data frame
So now when I got to examine the type of my new column compared to one of the preloaded columns I have the following observation
> typeof(ch_1_exer_119_GPA$fitted_GPA) #added column to data frame
[1] "list"
> typeof(ch_1_exer_119_GPA$GPA) #preloaded column to data frame
[1] "double"
This came up when I was entering the name of one of the created columns for another calculation and noticed that the icon in front of the variable was not a "purple tag" like the variables that came loaded with the data set, but instead had the "data frame" icon.
This didn't have a direct effect on any of the simple calculations I did, but I can envision something like this presenting a problem in the future when I'm dealing with more complex scenarios. So I'd like to get an understanding as to what it is that I did to create this and how to rectify it?
Thank you in advance.
EDIT: As requested from r2evans the following output:
> dput(head(ch_1_exer_119_GPA,15))
structure(list(GPA = c(3.897, 3.885, 3.778, 2.54, 3.028, 3.865,
2.962, 3.961, 0.5, 3.178, 3.31, 3.538, 3.083, 3.013, 3.245),
ACT_Score = c(21, 14, 28, 22, 21, 31, 32, 27, 29, 26, 24,
30, 24, 24, 33), fitted_GPA = structure(list(ACT_Score = c(2.92941895227791,
2.65762906394109, 3.20120884061472, 2.96824607918317, 2.92941895227791,
3.3176902213305, 3.35651734823576, 3.16238171370946, 3.24003596751998,
3.1235545868042, 3.04590033299369, 3.27886309442524, 3.04590033299369,
3.04590033299369, 3.39534447514102)), class = "data.frame", row.names = c(NA,
-15L)), residuals_GPA = structure(list(ACT_Score = c(0.967581047722093,
1.22737093605891, 0.576791159385276, -0.428246079183166,
0.0985810477220932, 0.547309778669498, -0.394517348235762,
0.798618286290536, -2.74003596751998, 0.0544454131957952,
0.264099667006314, 0.259136905574757, 0.0370996670063146,
-0.0329003329936857, -0.150344475141022)), class = "data.frame", row.names = c(NA,
-15L))), row.names = c(NA, -15L), class = c("tbl_df", "tbl",
"data.frame"))
