I am about to train a 5 million rows of data containing 7 categorical variables (string), but soon will train a 31 million rows of data. I am wondering what the maximum number of worker nodes we can use in a cluster, because even if I type something like: 2,000,000, it doesn't show any indication of an error.
Another question would be, what would be the best way to determine how many worker nodes needed?
Thank you in advance!