I am running a Logit Model on data I found in Kaggle.
https://www.kaggle.com/datasets/leonardopena/top50spotify2019
My goal is to predict which songs will be an international hit (TRUE). The model seems to predicting the songs that not going to be an international hit (FALSE).
Could someone shed some light on why the model is predicting FALSE instead of TRUE? Appreciate all help.
structure(list(bpm = c(105L, 170L, 120L, 87L, 129L, 125L),
nrgy = c(72L,
71L, 42L, 38L, 71L, 94L), dnce = c(72L, 74L, 75L, 72L, 58L,
74L), dB = c(-7L, -4L, -8L, -8L, -8L, -1L), hit = c(TRUE,
TRUE, TRUE, FALSE, TRUE, FALSE)), row.names = c(8L, 80L,
15L, 361L, 42L, 185L), class = "data.frame")
dfTop50 <- read.csv("SpotifyTop50country_prepared.csv",
row.names = 1, stringsAsFactors = FALSE)
train <- 0.7
nCases <- nrow(dfTop50)
set.seed(123)
trainCases <- sample(1:nCases, floor(train*nCases))
dfTop50Train <- dfTop50[ trainCases ,]
dfTop50Test <- dfTop50[ -trainCases ,]
mdlA <- hit ~ bpm + nrgy + dnce + dB
str(mdlA)
rsltLogit <- glm(mdlA, data = dfTop50Train, family =
binomial("logit"))
predLogit <- predict(rsltLogit, dfTop50Test, type =
"response")
head(cbind(Observed = dfTop50Test$hit, Predicted =
predLogit))
predLogit <- factor(as.numeric(predLogit > 0.5),
levels = c(0,1),
labels=c("FALSE","TRUE"))
accLogit <- mean(predLogit == dfTop50Test$hit)
describe(accLogit)
tblLog <- table(Predicted = predLogit,
Observed = dfTop50Test$hit)
View(tblLog)