update:
Turned out it's caused by different classes of variables.
Many thanks to @r2evans, who solved this issue by converting interger64 to numeric when reading the data. His method is effective, but what's worth studying further is his problem-solving logic.
I deleted the data for confidentiality reasons.
Below is the previous question
I plotted histograms of all numeric clomuns in my data table.
head(dt) %>%
keep(is.numeric) %>%
gather() %>% na.omit() %>%
ggplot(aes(value)) +
facet_wrap(~ key, scales = "free") +
geom_histogram()
I chose head() as the data table is too large.
then I had this error:
Error in if (length(unique(intervals)) > 1 & any(diff(scale(intervals)) < : missing value where TRUE/FALSE needed
Then I let
eg <- head(dt)
write.csv2(head(dt), "eg.csv")
and saved eg here on github.
then
eg <- fread("https://raw.githubusercontent.com/Deborah-Jia/Complete_Analysis_da2/main/eg.csv")
eg %>%
keep(is.numeric) %>%
gather() %>% na.omit() %>%
ggplot(aes(value)) +
facet_wrap(~ key, scales = "free") +
geom_histogram()
I got those right histograms!
What happened when I saved the data and read it again? Or is there a way to fix dt?
PS: dt was also created from saving csv and reading from fread. when I use
eg <- head(dt, 10000)
and save it on github, read again. same error happened.
Is it because my dt is too long (3 million rows) and had some wrong rows?
