AttributeError: 'list' object has no attribute 'dtype' when training a DNN

Viewed 425

I realise that my problem might is potentially very simple, and that a solution would be fairly straightforward, but I just cant seem to fix it.

Summary: I am trying to train a DNN to be able to predict what map is being picked or banned in a CS:GO map selection procedure. My data consists of 7 dummy variables which is used as target variables, and 4 predictor variables. I keep getting the error: AttributeError: 'list' object has no attribute 'dtype'. I dont understand why, because the training dataset im using is classified as a data frame, but the error implies that I'm feeding the network a list. I know that dataframes are technically also lists in R, but it should still be able to work.

I have reproduced the error with the following R-script, in which I construct a tiny representative dataset and include the small data cleaning/feature engineering I do in my main script. My comments are marked by a #.


library(tidyverse)
library(caret)
library(dplyr)
library(plyr)
library(rjson)
library(jsonlite)
library(keras)
library(tensorflow)
library(fastDummies)

# Here I construct the data set.

order <- c(1,2,3,4,5,6,7,1,2,3,4,5,6,7,1,2,3,4,5,6,7)
teamId <- c(11111, 22222, 11111, 22222, 11111, 22222, NA, 33333, 44444, 33333, 44444, 33333, 44444, NA, 55555, 66666, 55555, 66666, 55555, 66666, NA)
pickOrBan <- c("ban", "ban", "pick", "pick","ban", "ban", "pick", "ban", "ban", "pick", "pick","ban", "ban", "pick", "ban", "ban", "pick", "pick","ban", "ban", "pick")
mapName <- c("Mirage", "Nuke", "vertigo", "Dust2", "Train", "Overpass", "Inferno", "Mirage", "Nuke", "vertigo", "Dust2", "Train", "Overpass", "Inferno", "Mirage", "Nuke", "vertigo", "Dust2", "Train", "Ancient", "Inferno")
opponentId <- c(22222, 11111, 22222, 11111, 22222, 11111, NA, 44444, 33333, 44444, 33333, 44444, 33333, NA, 66666, 55555, 66666, 55555, 66666, 55555, NA)
fixtureId <- c(123, 123, 123, 123, 123, 123, 123, 456, 456, 456, 456, 456, 456, 456, 789, 789, 789, 789, 789, 789, 789)

df <- data.frame(order, teamId, pickOrBan, mapName, opponentId, fixtureId)

# Here I manually one hot encode the "mapName" variable, creating 7 dummy variables. 
# mapName originally had 8 possible values, to avoid multicollinearity I create 7 dummy variables.

df$miragemap <- ifelse(df$mapName == 'Mirage', 1, 0)
df$nukemap <- ifelse(df$mapName == 'Nuke', 1, 0)
df$vertigomap <- ifelse(df$mapName == 'Vertigo', 1, 0)
df$dust2map <- ifelse(df$mapName == 'Dust2', 1, 0)
df$trainmap <- ifelse(df$mapName == 'Train', 1, 0)
df$overpassmap <- ifelse(df$mapName == 'Overpass', 1, 0)
df$infernomap <- ifelse(df$mapName == 'Inferno', 1, 0)

# Replacing NA's in "teamID" to "decider". This is common terminology for the third map that will be played, if the first two maps will be 1-1.

df$teamId <- replace_na(df$teamId, 0)

# Changing opponentId such that the decider of teamId wont have any opponent, but itself.

df <- within(df, opponentId[teamId == 0] <- 0)

# Renaming column pick or ban column #

names(df)[names(df) == "pickOrBan"] <- "pick1ban0"

# See variable types and making sure everything is correct type

str(df)
df$teamId <- factor(df$teamId)
df$pick1ban0 <- factor(df$pick1ban0)
df$opponentId <- factor(df$opponentId)

# Removing unused variables

df <- subset(df, select = -c(mapName, fixtureId))

# Creating dummy variables using package: fastDummies

df <- dummy_columns(df, remove_first_dummy = TRUE)

# Relocating a variable

df <- df %>% relocate(pick1ban0_pick, .before = order)

# Creating my neural network and running it

set.seed(111)

# Sample size

ind <- sample(2, nrow(df), replace=TRUE, prob=c(0.75,0.25))

smp_size <- floor(0.75 * nrow(df))

train_ind <- sample(seq_len(nrow(df)), size = smp_size)

# Split data into training and testing

df_trainx <- df[train_ind, c(1:5, 13:24)]
df_testx <- df[-train_ind, c(1:5, 13:24)]
df_trainy <- df[train_ind, 6:12]
df_testy <- df[-train_ind, 6:12]

# Create model

model <- keras_model_sequential()

# Adding layers
model %>%
  layer_dense(units = 36, activation = 'relu', 
              input_shape = c(17)) %>%
  layer_dense(units = 36, activation = 'relu') %>%
  layer_dense(units = 18, activation = 'relu') %>%
  layer_dense(units = 7, activation = 'softmax')

# Compile model
model %>% compile(
  loss = 'categorical_crossentropy',
  optimizer = 'adam',
  metrics = 'accuracy'
)

# run model
history <- model %>% fit(
  df_trainx,
  df_trainy,
  epochs = 10,
  batch_size = 128
)
1 Answers

I have had the same issue, and seemingly it was due to feeding the target variable as factor, and not numeric. At least, the error did not show up after I coerced the target to numeri

Related