I am very new to Python and Pandas, and am trying to use it for a statistical analysis of a very large dataset (10 million cases) because the other options (SPSS and R) are unable to handle the dataset on the authorized hardware.
In this analysis, I need to search a range of columns (30 to be exact) row-wise to extract individual strings (about 200 are possible, not sure how many will actually be present in the dataset) and then create a categorical variable for each string.
The data looks like this
Dx1 Dx2 Dx3 etc...
001 234 456
231 001 444
245 777 001
What is needed is
Dx1 Dx2 Dx3 Var001 Var234 Var456 Var231 etc..
001 234 456 True True True False
231 001 444 True False False True
245 777 001 True False False False
Any thoughts on how to do this?
df.dtypes shows
AGE int64
DISPUNIFORM int64
DRG int64
DRGVER int64
Readmit_30D int64
DXCCS1 int64
DXCCS2 int64
DXCCS3 int64
DXCCS4 int64
...on to DXCCS30