be proud to get in touch with my first question on here :) I'm on my way to a self aducated data scientist and everyday brings another challenge. My challenge today:
I followed an example where someone used to bin the values via pandas 'cut' function, but this doesn't get rid of the data's high skewness:
(-0.512, 102.466] 838
(102.466, 204.932] 33
(204.932, 307.398] 17
(409.863, 512.329] 3
(307.398, 409.863] 0
So I binned them with pandas 'qcut' to get mostly even sized bins (no "visual" skewness for the distribution of bins 'value_counts'):
(7.854, 10.5] 184
(21.679, 39.688] 180
(-0.001, 7.854] 179
(39.688, 512.329] 176
(10.5, 21.679] 172
My intuition tells me, that that's not the way to handle skewed data after I read about data transformations via 'log' or 'sqrt' f.e.
Am I comparing apples with oranges here? :)